Skip to main content
Version: 0.0.42

Operations

Operations covers the automated processes that act on data — loading, semantic extraction, and custom workflow jobs. For system-level management (users, backup, scaling, upgrade), see Administration.

Operations is the automation engine of Lakehousecat. It provides a structured way to define, schedule, and monitor recurring and on-demand workflows. Internally, Operations is built on Apache Airflow — every Job Definition you create is automatically translated into an Airflow DAG and executed on Kubernetes.

Core Concepts​

Operations is organized around two entities:

EntityDescription
Job DefinitionA template that describes a workflow: its tasks, dependencies, schedule, and resource requirements
Job RunA single execution instance of a Job Definition, created either by a schedule or triggered manually

Architecture​

Each task within a Job Definition runs as an isolated Kubernetes Pod using the KubernetesPodOperator. This gives each task full resource isolation and allows fine-grained control over CPU, memory, environment variables, and Kubernetes secrets.

Tasks are packaged as ZIP artefacts containing the executable code. The entry point is always a run.sh script. The code can be written in Python, call CLI tools, or perform any other operation supported by the selected worker image.

Dependencies can be defined between tasks within the same Job Definition. This creates a directed acyclic graph (DAG) that Airflow uses to determine execution order. Lakehousecat reflects the full Job Definition to Airflow automatically on save.

Access Control​

Creating and managing Job Definitions requires the Administrator or Builder role. End users cannot create or modify Job Definitions.

The execution context of each task is determined by the selected Worker Type:

Worker TypeDescription
adminRuns with admin-level service account and base image
builderRuns with builder-level service account
userRuns with user-level service account
serviceDedicated service worker
semanticWorker for semantic layer operations

Built-in Job Definitions​

Lakehousecat ships with a set of built-in Job Definitions that cover the most common automation needs. They are pre-deployed and locked against deletion:

Job TypeDescription
Full Load DatasourcePerforms a full reload of a connected datasource
Incremental Load DatasourceLoads only new or updated records since the last run
Clear DatasourceClears all data from a datasource
Create / Update / Delete Semantic DatasourceSemantic layer lifecycle operations for datasources
Create / Update / Delete Semantic ModelSemantic layer lifecycle operations for models
Backup PostgreSQLCreates a backup of the PostgreSQL database
Backup ClickHouseCreates a backup of the ClickHouse database

Operations Overview​

The Operations Overview tab displays a live statistical summary of all Job Definitions and Job Runs.

Job Definitions​

  • Total Definitions — total number of job definitions across all types
  • Active / Inactive — how many definitions are currently enabled or disabled
  • Scheduled — how many definitions have a CRON schedule configured
  • Definitions by Type — breakdown by job type with a visual progress bar

Job Runs​

  • Total Runs — total number of job runs executed
  • Running — currently active runs
  • Successful / Failed — outcome breakdown
  • Recent (7 days) — runs executed in the last 7 days, including success and failure counts
  • Runs by Status — breakdown by status with visual progress bars
  • Timeline — oldest and newest run dates, average runs per day