Operations
Operations covers the automated processes that act on data — loading, semantic extraction, and custom workflow jobs. For system-level management (users, backup, scaling, upgrade), see Administration.
Operations is the automation engine of Lakehousecat. It provides a structured way to define, schedule, and monitor recurring and on-demand workflows. Internally, Operations is built on Apache Airflow — every Job Definition you create is automatically translated into an Airflow DAG and executed on Kubernetes.
Core Concepts
Operations is organized around two entities:
| Entity | Description |
|---|---|
| Job Definition | A template that describes a workflow: its tasks, dependencies, schedule, and resource requirements |
| Job Run | A single execution instance of a Job Definition, created either by a schedule or triggered manually |
Architecture
Each task within a Job Definition runs as an isolated Kubernetes Pod using the KubernetesPodOperator. This gives each task full resource isolation and allows fine-grained control over CPU, memory, environment variables, and Kubernetes secrets.
Tasks are packaged as ZIP artefacts containing the executable code. The entry point is always a run.sh script. The code can be written in Python, call CLI tools, or perform any other operation supported by the selected worker image.
Dependencies can be defined between tasks within the same Job Definition. This creates a directed acyclic graph (DAG) that Airflow uses to determine execution order. Lakehousecat reflects the full Job Definition to Airflow automatically on save.
Access Control
Creating and managing Job Definitions requires the Administrator or Builder role. End users cannot create or modify Job Definitions.
The execution context of each task is determined by the selected Worker Type:
| Worker Type | Description |
|---|---|
admin | Runs with admin-level service account and base image |
builder | Runs with builder-level service account |
user | Runs with user-level service account |
service | Dedicated service worker |
semantic | Worker for semantic layer operations |
Built-in Job Definitions
Lakehousecat ships with a set of built-in Job Definitions that cover the most common automation needs. They are pre-deployed and locked against deletion:
| Job Type | Description |
|---|---|
| Full Load Datasource | Performs a full reload of a connected datasource |
| Incremental Load Datasource | Loads only new or updated records since the last run |
| Clear Datasource | Clears all data from a datasource |
| Create / Update / Delete Semantic Datasource | Semantic layer lifecycle operations for datasources |
| Create / Update / Delete Semantic Model | Semantic layer lifecycle operations for models |
| Backup PostgreSQL | Creates a backup of the PostgreSQL database |
| Backup ClickHouse | Creates a backup of the ClickHouse database |
Operations Overview
The Operations Overview tab displays a live statistical summary of all Job Definitions and Job Runs.
Job Definitions
- Total Definitions — total number of job definitions across all types
- Active / Inactive — how many definitions are currently enabled or disabled
- Scheduled — how many definitions have a CRON schedule configured
- Definitions by Type — breakdown by job type with a visual progress bar
Job Runs
- Total Runs — total number of job runs executed
- Running — currently active runs
- Successful / Failed — outcome breakdown
- Recent (7 days) — runs executed in the last 7 days, including success and failure counts
- Runs by Status — breakdown by status with visual progress bars
- Timeline — oldest and newest run dates, average runs per day