Job Definitions
A Job Definition is a reusable template that describes an automated workflow. Each definition is translated into an Apache Airflow DAG and executed on Kubernetes. Creating Job Definitions requires the Administrator or Builder role.
Navigate to Workspace → Operations → Job Definitions and click + to create a new definition.
General Fields
| Field | Description |
|---|---|
| Name (required) | A unique name for the job. Must start with a letter. Allowed characters: letters, numbers, spaces, and hyphens. |
| Description | Optional description of the job's purpose. |
| Job Type (required) | The operational category of the job (see Job Types below). |
| Schedule (Cron) | Optional CRON expression for automatic execution (e.g., 0 2 * * * for every day at 2:00 AM). Leave empty for manual-only execution. |
| DAG Timeout | Maximum runtime for the whole job. Enter a number and choose the unit (minutes or hours). Leave empty to use the default of 4 hours. If the job exceeds this limit, it is stopped and marked as failed. |
| DAG Default Args | Optional JSON object with Airflow default arguments applied to every task, such as retries (number of automatic retry attempts) and retry_delay (wait time between attempts, e.g. "5m"). Leave empty to use the defaults. |
| Enabled | If disabled, the job will not execute, even if triggered manually. |
| Locked | If locked, the definition cannot be edited or deleted. |
Job Types
Job types fall into two categories: custom task types for your own code, and system types used by Lakehousecat's built-in automation.
Custom Task Types
These types are for workflows you build and package yourself. Select the type that matches the minimum required privilege level for the task.
| Type | Required Role | Description |
|---|---|---|
USER_TASK | Builder or Admin | Custom workflow running with user-level permissions. Suitable for data exports, reporting, and non-privileged automation. |
BUILDER_TASK | Builder or Admin | Custom workflow running with builder-level permissions. Suitable for data pipeline tasks and model operations that require Builder access. |
ADMIN_TASK | Admin only | Custom workflow running with admin-level permissions. Has the widest system access — use only when administrative operations are required. |
Custom tasks require a ZIP artefact containing your code and a run.sh entry point. The code can be written in any language or use any CLI tool available in the worker image.
System Built-in Types
These types are used by Job Definitions that Lakehousecat creates automatically when you enable data loading or semantic operations on a data source or model. They run in the background without manual configuration and are pre-deployed as locked definitions.
| Type | Triggered by | Description |
|---|---|---|
FULL_LOAD_DATASOURCE | Data source operations | Full reload of all data from a connected datasource into ClickHouse |
INCREMENTAL_LOAD_DATASOURCE | Data source operations | Loads only new or updated records since the last run |
CLEAR_DATASOURCE | Data source operations | Removes all loaded data for a datasource |
CREATE_SEMANTIC_DATASOURCE | Semantic extraction | Creates the semantic layer (schema, descriptions) for a datasource |
UPDATE_SEMANTIC_DATASOURCE | Semantic extraction | Updates the semantic layer after datasource changes |
DELETE_SEMANTIC_DATASOURCE | Datasource deletion | Removes the semantic layer for a datasource |
CREATE_SEMANTIC_MODEL | Model configuration | Creates the semantic layer for a custom model |
UPDATE_SEMANTIC_MODEL | Model configuration | Updates the semantic model layer |
DELETE_SEMANTIC_MODEL | Model deletion | Removes the semantic model layer |
BACKUP_POSTGRES | Scheduled backup | Creates a backup of the PostgreSQL database |
BACKUP_CLICKHOUSE | Scheduled backup | Creates a backup of the ClickHouse database |
System built-in job definitions are locked and cannot be edited or deleted. They appear in the Operations list for visibility and monitoring only.
All system built-in types except BACKUP_POSTGRES and BACKUP_CLICKHOUSE are templates: Lakehousecat clones them per datasource or model when you enable an operation, and only the clone is actually executed. The built-in template itself has no meaningful target and cannot be triggered directly — see Triggering and Monitoring.
Tasks
Every Job Definition must contain at least one task. Each task runs as an isolated Kubernetes Pod.
Task Fields
| Field | Description |
|---|---|
| Task ID (required) | A unique identifier for the task within this definition. Used when defining dependencies. |
| Worker Type (required) | Determines the execution context: admin, builder, user, service, or semantic. Controls the base image and Kubernetes service account. |
| CPU Request / Limit | Kubernetes CPU resources for the pod (e.g., 100m request, 500m limit). |
| Memory Request / Limit | Kubernetes memory resources for the pod (e.g., 128Mi request, 512Mi limit). |
| Task Timeout | Maximum runtime for this individual task. Enter a number and choose the unit (minutes or hours). Leave empty to use the default of 60 minutes. A task that exceeds this limit is stopped and marked as failed, independent of the overall DAG Timeout. |
| Task Dependencies | Other tasks within this definition that must complete before this task starts. Creates the DAG execution order. |
| Advanced Config | JSON object with env_vars (key-value environment variables) and secrets_ref (list of Kubernetes secret names to mount in the pod). |
| ZIP Artefact (required) | A .zip file containing the executable code for this task. Must include a run.sh entry point. |
ZIP Artefact
The ZIP artefact is the executable payload of a task. When the task runs, the ZIP is extracted inside the Kubernetes pod and run.sh is called automatically.
Minimal artefact structure:
my-task.zip
├── run.sh ← entry point, always required
├── main.py ← your code
└── requirements.txt
run.sh is the only required file. It receives environment variables injected by Airflow and any variables you defined in the Advanced Config. A typical entry point:
#!/bin/bash
set -e
python main.py
If the ZIP contains both main.py and requirements.txt, the task installs the packages listed in requirements.txt from the wheels/ folder inside the ZIP and runs python main.py itself. run.sh is not called in that case. See Installing Python Packages.
Your code can:
- Run Python scripts or modules (
python main.py,python -m mymodule) - Call CLI tools (
dbt run,spark-submit,curl) - Interact with Lakehousecat APIs or external services using environment variables for credentials
Accessing environment variables in Python:
import os
datasource_id = os.environ["DATASOURCE_ID"]
api_key = os.environ["LHC_API_KEY"]
Each task requires its own ZIP. When editing a definition, the existing ZIP can be downloaded, modified, and re-uploaded.
What the Worker Image Contains
Worker images are deliberately small. They are based on Debian with Python 3.12 and contain no third-party Python packages. Every worker image includes bash, the Python standard library, pip, curl, unzip, the MinIO client mc, the AWS CLI and the lhc CLI. The admin worker additionally includes the PostgreSQL 18 client tools (psql, pg_dump, pg_restore), kubectl, rclone and age.
The root filesystem of a task pod is read-only. /tmp, the home directory and the task's working directory are writable.
Installing Python Packages
Install the packages your code needs when the task runs. pip install --user installs into /tmp, so it works despite the read-only root filesystem. By default, pip installs only from files shipped with the task and does not contact PyPI. There are three ways to provide packages:
| Way | How | Network access |
|---|---|---|
| Ship wheels in the ZIP (recommended) | Run pip download -r requirements.txt -d wheels/ --only-binary=:all: --platform manylinux_2_28_x86_64 --python-version 3.12 locally (use manylinux_2_28_aarch64 for ARM clusters) and add the wheels/ folder to the ZIP. The task installs them offline. | None |
| Install from PyPI | Set egress_profile: external-https on the task and install in run.sh with env -u PIP_NO_INDEX -u PIP_FIND_LINKS pip install --user <package>. The two variables are what keeps pip offline by default. | Outbound HTTPS |
| Store once, reuse | A provisioning task with egress_profile: external-https downloads the wheels, packs them with tar and stores the archive with lhc files upload. Later tasks fetch it with lhc files download <file-id> and install offline. This is the pattern the vendor driver job definitions use. | Only for the provisioning task |
Only binary wheels can be installed. The images contain no compiler, so source distributions with C extensions do not build. Installed packages count towards the task's ephemeral storage.
Task Dependencies
Tasks can depend on other tasks within the same Job Definition. A task will only start after all its upstream tasks have completed successfully. This allows sequential and branching workflows.
When more than one task is defined, an Execution Flow Preview is shown below the task list — it visualizes the full dependency graph before saving.
Job Definition Map
In edit mode, the Map tab next to Tasks shows the structure of the saved definition as an interactive graph: its tasks, the dependencies between them, and the start and end points of the workflow. It is the quickest way to understand a larger definition with many branching tasks.
- The map shows the saved state. Define tasks in the Tasks tab, save, then open Map; it is brought up to date each time you open the tab.
- It shows structure only — task order and dependencies. Secrets and detailed task configuration stay in the Tasks tab.
- Everyone who can read the definition can see its map: the owner, administrators, and users it is shared with.
- If the definition has no map yet (for example because it could not be parsed), the tab shows a "No Job Definition Map available" message.
Use the legend to filter the graph. The map complements the Execution Flow Preview in the Tasks tab, which previews the dependencies of unsaved edits.
Triggering and Monitoring
In edit mode, two additional actions are available:
| Action | Description |
|---|---|
| Trigger Job | Manually triggers an immediate execution of the job, independent of any schedule |
| View Runs | Navigates to the Job Runs tab filtered to this definition |
For built-in template definitions (see note above), Trigger Job is disabled and shows the label "Template — not directly triggerable" with a tooltip explaining that the operation must be triggered via the associated datasource or model operation instead. This applies only to the template itself — clones created for a specific datasource or model remain triggerable as normal.
When edits take effect
After you edit a Job Definition and save, changes are picked up on the next run:
- CPU / Memory resources always apply to the next run automatically — no waiting is required, even if you trigger immediately after saving.
- DAG Timeout, Task Timeout, retries, and retry delay need a short moment to propagate to the execution engine. If you trigger a job right after changing one of these values, Lakehousecat briefly waits (up to about a minute) until the change has been applied, so a run never starts with an outdated timeout. If it has not propagated in time, the trigger is rejected with the message "Job definition was recently updated and Airflow has not yet picked up the change. Please retry the trigger shortly." — simply trigger again a few seconds later. Scheduled runs are unaffected.
CRON Schedule Reference
| Expression | Meaning |
|---|---|
0 * * * * | Every hour |
0 2 * * * | Every day at 2:00 AM |
0 2 * * 1 | Every Monday at 2:00 AM |
0 0 1 * * | First day of every month at midnight |
*/15 * * * * | Every 15 minutes |
Advanced Config Reference
The Advanced Config field accepts a JSON object with the following keys:
{
"env_vars": {
"MY_VAR": "my_value"
},
"secrets_ref": [
"my-kubernetes-secret-name"
]
}
env_vars— merged with system-generated environment variables in the podsecrets_ref— list of Kubernetes secret names; each secret is mounted as environment variables or a volume depending on its configuration