Skip to main content
Version: 0.0.41

Job Definitions

A Job Definition is a reusable template that describes an automated workflow. Each definition is translated into an Apache Airflow DAG and executed on Kubernetes. Creating Job Definitions requires the Administrator or Builder role.

Navigate to Workspace → Operations → Job Definitions and click + to create a new definition.

General Fields​

FieldDescription
Name (required)A unique name for the job. Must start with a letter. Allowed characters: letters, numbers, spaces, and hyphens.
DescriptionOptional description of the job's purpose.
Job Type (required)The operational category of the job (see Job Types below).
Schedule (Cron)Optional CRON expression for automatic execution (e.g., 0 2 * * * for every day at 2:00 AM). Leave empty for manual-only execution.
EnabledIf disabled, the job will not execute, even if triggered manually.
LockedIf locked, the definition cannot be edited or deleted.

Job Types​

Job types fall into two categories: custom task types for your own code, and system types used by Lakehousecat's built-in automation.

Custom Task Types​

These types are for workflows you build and package yourself. Select the type that matches the minimum required privilege level for the task.

TypeRequired RoleDescription
USER_TASKBuilder or AdminCustom workflow running with user-level permissions. Suitable for data exports, reporting, and non-privileged automation.
BUILDER_TASKBuilder or AdminCustom workflow running with builder-level permissions. Suitable for data pipeline tasks and model operations that require Builder access.
ADMIN_TASKAdmin onlyCustom workflow running with admin-level permissions. Has the widest system access — use only when administrative operations are required.

Custom tasks require a ZIP artefact containing your code and a run.sh entry point. The code can be written in any language or use any CLI tool available in the worker image.

System Built-in Types​

These types are used by Job Definitions that Lakehousecat creates automatically when you enable data loading or semantic operations on a data source or model. They run in the background without manual configuration and are pre-deployed as locked definitions.

TypeTriggered byDescription
FULL_LOAD_DATASOURCEData source operationsFull reload of all data from a connected datasource into ClickHouse
INCREMENTAL_LOAD_DATASOURCEData source operationsLoads only new or updated records since the last run
CLEAR_DATASOURCEData source operationsRemoves all loaded data for a datasource
CREATE_SEMANTIC_DATASOURCESemantic extractionCreates the semantic layer (schema, descriptions) for a datasource
UPDATE_SEMANTIC_DATASOURCESemantic extractionUpdates the semantic layer after datasource changes
DELETE_SEMANTIC_DATASOURCEDatasource deletionRemoves the semantic layer for a datasource
CREATE_SEMANTIC_MODELModel configurationCreates the semantic layer for a custom model
UPDATE_SEMANTIC_MODELModel configurationUpdates the semantic model layer
DELETE_SEMANTIC_MODELModel deletionRemoves the semantic model layer
BACKUP_POSTGRESScheduled backupCreates a backup of the PostgreSQL database
BACKUP_CLICKHOUSEScheduled backupCreates a backup of the ClickHouse database
note

System built-in job definitions are locked and cannot be edited or deleted. They appear in the Operations list for visibility and monitoring only.

Tasks​

Every Job Definition must contain at least one task. Each task runs as an isolated Kubernetes Pod.

Task Fields​

FieldDescription
Task ID (required)A unique identifier for the task within this definition. Used when defining dependencies.
Worker Type (required)Determines the execution context: admin, builder, user, service, or semantic. Controls the base image and Kubernetes service account.
CPU Request / LimitKubernetes CPU resources for the pod (e.g., 100m request, 500m limit).
Memory Request / LimitKubernetes memory resources for the pod (e.g., 128Mi request, 512Mi limit).
Task DependenciesOther tasks within this definition that must complete before this task starts. Creates the DAG execution order.
Advanced ConfigJSON object with env_vars (key-value environment variables) and secrets_ref (list of Kubernetes secret names to mount in the pod).
ZIP Artefact (required)A .zip file containing the executable code for this task. Must include a run.sh entry point.

ZIP Artefact​

The ZIP artefact is the executable payload of a task. When the task runs, the ZIP is extracted inside the Kubernetes pod and run.sh is called automatically.

Minimal artefact structure:

my-task.zip
├── run.sh ← entry point, always required
├── main.py ← your code
└── requirements.txt

run.sh is the only required file. It receives environment variables injected by Airflow and any variables you defined in the Advanced Config. A typical entry point:

#!/bin/bash
set -e
pip install -r requirements.txt --quiet
python main.py

Your code can:

  • Run Python scripts or modules (python main.py, python -m mymodule)
  • Call CLI tools (dbt run, spark-submit, curl)
  • Interact with Lakehousecat APIs or external services using environment variables for credentials

Accessing environment variables in Python:

import os

datasource_id = os.environ["DATASOURCE_ID"]
api_key = os.environ["LHC_API_KEY"]

Each task requires its own ZIP. When editing a definition, the existing ZIP can be downloaded, modified, and re-uploaded.

Task Dependencies​

Tasks can depend on other tasks within the same Job Definition. A task will only start after all its upstream tasks have completed successfully. This allows sequential and branching workflows.

When more than one task is defined, an Execution Flow Preview is shown below the task list — it visualizes the full dependency graph before saving.

Triggering and Monitoring​

In edit mode, two additional actions are available:

ActionDescription
Trigger JobManually triggers an immediate execution of the job, independent of any schedule
View RunsNavigates to the Job Runs tab filtered to this definition

CRON Schedule Reference​

ExpressionMeaning
0 * * * *Every hour
0 2 * * *Every day at 2:00 AM
0 2 * * 1Every Monday at 2:00 AM
0 0 1 * *First day of every month at midnight
*/15 * * * *Every 15 minutes

Advanced Config Reference​

The Advanced Config field accepts a JSON object with the following keys:

{
"env_vars": {
"MY_VAR": "my_value"
},
"secrets_ref": [
"my-kubernetes-secret-name"
]
}
  • env_vars — merged with system-generated environment variables in the pod
  • secrets_ref — list of Kubernetes secret names; each secret is mounted as environment variables or a volume depending on its configuration