Job Definitions
A Job Definition is a reusable template that describes an automated workflow. Each definition is translated into an Apache Airflow DAG and executed on Kubernetes. Creating Job Definitions requires the Administrator or Builder role.
Navigate to Workspace → Operations → Job Definitions and click + to create a new definition.
General Fields
| Field | Description |
|---|---|
| Name (required) | A unique name for the job. Must start with a letter. Allowed characters: letters, numbers, spaces, and hyphens. |
| Description | Optional description of the job's purpose. |
| Job Type (required) | The operational category of the job (see Job Types below). |
| Schedule (Cron) | Optional CRON expression for automatic execution (e.g., 0 2 * * * for every day at 2:00 AM). Leave empty for manual-only execution. |
| Enabled | If disabled, the job will not execute, even if triggered manually. |
| Locked | If locked, the definition cannot be edited or deleted. |
Job Types
Job types fall into two categories: custom task types for your own code, and system types used by Lakehousecat's built-in automation.
Custom Task Types
These types are for workflows you build and package yourself. Select the type that matches the minimum required privilege level for the task.
| Type | Required Role | Description |
|---|---|---|
USER_TASK | Builder or Admin | Custom workflow running with user-level permissions. Suitable for data exports, reporting, and non-privileged automation. |
BUILDER_TASK | Builder or Admin | Custom workflow running with builder-level permissions. Suitable for data pipeline tasks and model operations that require Builder access. |
ADMIN_TASK | Admin only | Custom workflow running with admin-level permissions. Has the widest system access — use only when administrative operations are required. |
Custom tasks require a ZIP artefact containing your code and a run.sh entry point. The code can be written in any language or use any CLI tool available in the worker image.
System Built-in Types
These types are used by Job Definitions that Lakehousecat creates automatically when you enable data loading or semantic operations on a data source or model. They run in the background without manual configuration and are pre-deployed as locked definitions.
| Type | Triggered by | Description |
|---|---|---|
FULL_LOAD_DATASOURCE | Data source operations | Full reload of all data from a connected datasource into ClickHouse |
INCREMENTAL_LOAD_DATASOURCE | Data source operations | Loads only new or updated records since the last run |
CLEAR_DATASOURCE | Data source operations | Removes all loaded data for a datasource |
CREATE_SEMANTIC_DATASOURCE | Semantic extraction | Creates the semantic layer (schema, descriptions) for a datasource |
UPDATE_SEMANTIC_DATASOURCE | Semantic extraction | Updates the semantic layer after datasource changes |
DELETE_SEMANTIC_DATASOURCE | Datasource deletion | Removes the semantic layer for a datasource |
CREATE_SEMANTIC_MODEL | Model configuration | Creates the semantic layer for a custom model |
UPDATE_SEMANTIC_MODEL | Model configuration | Updates the semantic model layer |
DELETE_SEMANTIC_MODEL | Model deletion | Removes the semantic model layer |
BACKUP_POSTGRES | Scheduled backup | Creates a backup of the PostgreSQL database |
BACKUP_CLICKHOUSE | Scheduled backup | Creates a backup of the ClickHouse database |
System built-in job definitions are locked and cannot be edited or deleted. They appear in the Operations list for visibility and monitoring only.
Tasks
Every Job Definition must contain at least one task. Each task runs as an isolated Kubernetes Pod.
Task Fields
| Field | Description |
|---|---|
| Task ID (required) | A unique identifier for the task within this definition. Used when defining dependencies. |
| Worker Type (required) | Determines the execution context: admin, builder, user, service, or semantic. Controls the base image and Kubernetes service account. |
| CPU Request / Limit | Kubernetes CPU resources for the pod (e.g., 100m request, 500m limit). |
| Memory Request / Limit | Kubernetes memory resources for the pod (e.g., 128Mi request, 512Mi limit). |
| Task Dependencies | Other tasks within this definition that must complete before this task starts. Creates the DAG execution order. |
| Advanced Config | JSON object with env_vars (key-value environment variables) and secrets_ref (list of Kubernetes secret names to mount in the pod). |
| ZIP Artefact (required) | A .zip file containing the executable code for this task. Must include a run.sh entry point. |
ZIP Artefact
The ZIP artefact is the executable payload of a task. When the task runs, the ZIP is extracted inside the Kubernetes pod and run.sh is called automatically.
Minimal artefact structure:
my-task.zip
├── run.sh ← entry point, always required
├── main.py ← your code
└── requirements.txt
run.sh is the only required file. It receives environment variables injected by Airflow and any variables you defined in the Advanced Config. A typical entry point:
#!/bin/bash
set -e
pip install -r requirements.txt --quiet
python main.py
Your code can:
- Run Python scripts or modules (
python main.py,python -m mymodule) - Call CLI tools (
dbt run,spark-submit,curl) - Interact with Lakehousecat APIs or external services using environment variables for credentials
Accessing environment variables in Python:
import os
datasource_id = os.environ["DATASOURCE_ID"]
api_key = os.environ["LHC_API_KEY"]
Each task requires its own ZIP. When editing a definition, the existing ZIP can be downloaded, modified, and re-uploaded.
Task Dependencies
Tasks can depend on other tasks within the same Job Definition. A task will only start after all its upstream tasks have completed successfully. This allows sequential and branching workflows.
When more than one task is defined, an Execution Flow Preview is shown below the task list — it visualizes the full dependency graph before saving.
Triggering and Monitoring
In edit mode, two additional actions are available:
| Action | Description |
|---|---|
| Trigger Job | Manually triggers an immediate execution of the job, independent of any schedule |
| View Runs | Navigates to the Job Runs tab filtered to this definition |
CRON Schedule Reference
| Expression | Meaning |
|---|---|
0 * * * * | Every hour |
0 2 * * * | Every day at 2:00 AM |
0 2 * * 1 | Every Monday at 2:00 AM |
0 0 1 * * | First day of every month at midnight |
*/15 * * * * | Every 15 minutes |
Advanced Config Reference
The Advanced Config field accepts a JSON object with the following keys:
{
"env_vars": {
"MY_VAR": "my_value"
},
"secrets_ref": [
"my-kubernetes-secret-name"
]
}
env_vars— merged with system-generated environment variables in the podsecrets_ref— list of Kubernetes secret names; each secret is mounted as environment variables or a volume depending on its configuration