Skip to main content
Version: Next

Job Definitions

A Job Definition is a reusable template that describes an automated workflow. Each definition is translated into an Apache Airflow DAG and executed on Kubernetes. Creating Job Definitions requires the Administrator or Builder role.

Navigate to Workspace → Operations → Job Definitions and click + to create a new definition.

General Fields​

FieldDescription
Name (required)A unique name for the job. Must start with a letter. Allowed characters: letters, numbers, spaces, and hyphens.
DescriptionOptional description of the job's purpose.
Job Type (required)The operational category of the job (see Job Types below).
Schedule (Cron)Optional CRON expression for automatic execution (e.g., 0 2 * * * for every day at 2:00 AM). Leave empty for manual-only execution.
DAG TimeoutMaximum runtime for the whole job. Enter a number and choose the unit (minutes or hours). Leave empty to use the default of 4 hours. If the job exceeds this limit, it is stopped and marked as failed.
DAG Default ArgsOptional JSON object with Airflow default arguments applied to every task, such as retries (number of automatic retry attempts) and retry_delay (wait time between attempts, e.g. "5m"). Leave empty to use the defaults.
EnabledIf disabled, the job will not execute, even if triggered manually.
LockedIf locked, the definition cannot be edited or deleted.

Job Types​

Job types fall into two categories: custom task types for your own code, and system types used by Lakehousecat's built-in automation.

Custom Task Types​

These types are for workflows you build and package yourself. Select the type that matches the minimum required privilege level for the task.

TypeRequired RoleDescription
USER_TASKBuilder or AdminCustom workflow running with user-level permissions. Suitable for data exports, reporting, and non-privileged automation.
BUILDER_TASKBuilder or AdminCustom workflow running with builder-level permissions. Suitable for data pipeline tasks and model operations that require Builder access.
ADMIN_TASKAdmin onlyCustom workflow running with admin-level permissions. Has the widest system access — use only when administrative operations are required.

Custom tasks require a ZIP artefact containing your code and a run.sh entry point. The code can be written in any language or use any CLI tool available in the worker image.

System Built-in Types​

These types are used by Job Definitions that Lakehousecat creates automatically when you enable data loading or semantic operations on a data source or model. They run in the background without manual configuration and are pre-deployed as locked definitions.

TypeTriggered byDescription
FULL_LOAD_DATASOURCEData source operationsFull reload of all data from a connected datasource into ClickHouse
INCREMENTAL_LOAD_DATASOURCEData source operationsLoads only new or updated records since the last run
CLEAR_DATASOURCEData source operationsRemoves all loaded data for a datasource
CREATE_SEMANTIC_DATASOURCESemantic extractionCreates the semantic layer (schema, descriptions) for a datasource
UPDATE_SEMANTIC_DATASOURCESemantic extractionUpdates the semantic layer after datasource changes
DELETE_SEMANTIC_DATASOURCEDatasource deletionRemoves the semantic layer for a datasource
CREATE_SEMANTIC_MODELModel configurationCreates the semantic layer for a custom model
UPDATE_SEMANTIC_MODELModel configurationUpdates the semantic model layer
DELETE_SEMANTIC_MODELModel deletionRemoves the semantic model layer
BACKUP_POSTGRESScheduled backupCreates a backup of the PostgreSQL database
BACKUP_CLICKHOUSEScheduled backupCreates a backup of the ClickHouse database
note

System built-in job definitions are locked and cannot be edited or deleted. They appear in the Operations list for visibility and monitoring only.

All system built-in types except BACKUP_POSTGRES and BACKUP_CLICKHOUSE are templates: Lakehousecat clones them per datasource or model when you enable an operation, and only the clone is actually executed. The built-in template itself has no meaningful target and cannot be triggered directly — see Triggering and Monitoring.

Tasks​

Every Job Definition must contain at least one task. Each task runs as an isolated Kubernetes Pod.

Task Fields​

FieldDescription
Task ID (required)A unique identifier for the task within this definition. Used when defining dependencies.
Worker Type (required)Determines the execution context: admin, builder, user, service, or semantic. Controls the base image and Kubernetes service account.
CPU Request / LimitKubernetes CPU resources for the pod (e.g., 100m request, 500m limit).
Memory Request / LimitKubernetes memory resources for the pod (e.g., 128Mi request, 512Mi limit).
Task TimeoutMaximum runtime for this individual task. Enter a number and choose the unit (minutes or hours). Leave empty to use the default of 60 minutes. A task that exceeds this limit is stopped and marked as failed, independent of the overall DAG Timeout.
Task DependenciesOther tasks within this definition that must complete before this task starts. Creates the DAG execution order.
Advanced ConfigJSON object with env_vars (key-value environment variables) and secrets_ref (list of Kubernetes secret names to mount in the pod).
ZIP Artefact (required)A .zip file containing the executable code for this task. Must include a run.sh entry point.

ZIP Artefact​

The ZIP artefact is the executable payload of a task. When the task runs, the ZIP is extracted inside the Kubernetes pod and run.sh is called automatically.

Minimal artefact structure:

my-task.zip
├── run.sh ← entry point, always required
├── main.py ← your code
└── requirements.txt

run.sh is the only required file. It receives environment variables injected by Airflow and any variables you defined in the Advanced Config. A typical entry point:

#!/bin/bash
set -e
python main.py
note

If the ZIP contains both main.py and requirements.txt, the task installs the packages listed in requirements.txt from the wheels/ folder inside the ZIP and runs python main.py itself. run.sh is not called in that case. See Installing Python Packages.

Your code can:

  • Run Python scripts or modules (python main.py, python -m mymodule)
  • Call CLI tools (dbt run, spark-submit, curl)
  • Interact with Lakehousecat APIs or external services using environment variables for credentials

Accessing environment variables in Python:

import os

datasource_id = os.environ["DATASOURCE_ID"]
api_key = os.environ["LHC_API_KEY"]

Each task requires its own ZIP. When editing a definition, the existing ZIP can be downloaded, modified, and re-uploaded.

What the Worker Image Contains​

Worker images are deliberately small. They are based on Debian with Python 3.12 and contain no third-party Python packages. Every worker image includes bash, the Python standard library, pip, curl, unzip, the MinIO client mc, the AWS CLI and the lhc CLI. The admin worker additionally includes the PostgreSQL 18 client tools (psql, pg_dump, pg_restore), kubectl, rclone and age.

The root filesystem of a task pod is read-only. /tmp, the home directory and the task's working directory are writable.

Installing Python Packages​

Install the packages your code needs when the task runs. pip install --user installs into /tmp, so it works despite the read-only root filesystem. By default, pip installs only from files shipped with the task and does not contact PyPI. There are three ways to provide packages:

WayHowNetwork access
Ship wheels in the ZIP (recommended)Run pip download -r requirements.txt -d wheels/ --only-binary=:all: --platform manylinux_2_28_x86_64 --python-version 3.12 locally (use manylinux_2_28_aarch64 for ARM clusters) and add the wheels/ folder to the ZIP. The task installs them offline.None
Install from PyPISet egress_profile: external-https on the task and install in run.sh with env -u PIP_NO_INDEX -u PIP_FIND_LINKS pip install --user <package>. The two variables are what keeps pip offline by default.Outbound HTTPS
Store once, reuseA provisioning task with egress_profile: external-https downloads the wheels, packs them with tar and stores the archive with lhc files upload. Later tasks fetch it with lhc files download <file-id> and install offline. This is the pattern the vendor driver job definitions use.Only for the provisioning task

Only binary wheels can be installed. The images contain no compiler, so source distributions with C extensions do not build. Installed packages count towards the task's ephemeral storage.

Task Dependencies​

Tasks can depend on other tasks within the same Job Definition. A task will only start after all its upstream tasks have completed successfully. This allows sequential and branching workflows.

When more than one task is defined, an Execution Flow Preview is shown below the task list — it visualizes the full dependency graph before saving.

Job Definition Map​

In edit mode, the Map tab next to Tasks shows the structure of the saved definition as an interactive graph: its tasks, the dependencies between them, and the start and end points of the workflow. It is the quickest way to understand a larger definition with many branching tasks.

  • The map shows the saved state. Define tasks in the Tasks tab, save, then open Map; it is brought up to date each time you open the tab.
  • It shows structure only — task order and dependencies. Secrets and detailed task configuration stay in the Tasks tab.
  • Everyone who can read the definition can see its map: the owner, administrators, and users it is shared with.
  • If the definition has no map yet (for example because it could not be parsed), the tab shows a "No Job Definition Map available" message.

Use the legend to filter the graph. The map complements the Execution Flow Preview in the Tasks tab, which previews the dependencies of unsaved edits.

Triggering and Monitoring​

In edit mode, two additional actions are available:

ActionDescription
Trigger JobManually triggers an immediate execution of the job, independent of any schedule
View RunsNavigates to the Job Runs tab filtered to this definition

For built-in template definitions (see note above), Trigger Job is disabled and shows the label "Template — not directly triggerable" with a tooltip explaining that the operation must be triggered via the associated datasource or model operation instead. This applies only to the template itself — clones created for a specific datasource or model remain triggerable as normal.

When edits take effect​

After you edit a Job Definition and save, changes are picked up on the next run:

  • CPU / Memory resources always apply to the next run automatically — no waiting is required, even if you trigger immediately after saving.
  • DAG Timeout, Task Timeout, retries, and retry delay need a short moment to propagate to the execution engine. If you trigger a job right after changing one of these values, Lakehousecat briefly waits (up to about a minute) until the change has been applied, so a run never starts with an outdated timeout. If it has not propagated in time, the trigger is rejected with the message "Job definition was recently updated and Airflow has not yet picked up the change. Please retry the trigger shortly." — simply trigger again a few seconds later. Scheduled runs are unaffected.

CRON Schedule Reference​

ExpressionMeaning
0 * * * *Every hour
0 2 * * *Every day at 2:00 AM
0 2 * * 1Every Monday at 2:00 AM
0 0 1 * *First day of every month at midnight
*/15 * * * *Every 15 minutes

Advanced Config Reference​

The Advanced Config field accepts a JSON object with the following keys:

{
"env_vars": {
"MY_VAR": "my_value"
},
"secrets_ref": [
"my-kubernetes-secret-name"
]
}
  • env_vars — merged with system-generated environment variables in the pod
  • secrets_ref — list of Kubernetes secret names; each secret is mounted as environment variables or a volume depending on its configuration