Skip to main content
Version: Next

Operations

The Operations section is the workflow automation layer of Lakehousecat. Here, administrators and builders can define, schedule, and monitor automated jobs — from data loads and semantic extractions to custom administrative tasks and data synchronization workflows.

Operations is built on an Apache Airflow workflow engine. Job scripts are written in Shell (run.sh) or Python, and tasks can be allocated to specific worker types with configurable Kubernetes resources.


Prerequisites​

TaskRequired Role
Create and manage Job DefinitionsAdmin, Builder
View and monitor Job RunsAdmin, Builder

  1. Open the workspace.
  2. In the left sidebar, click Operations.

The Operations section contains two areas:

  • Job Definitions — Define and configure automated jobs
  • Job Runs — Monitor execution history and status of jobs

Job Definitions​

A Job Definition is a named, reusable workflow configuration. It specifies what work should be done, how often, and by which worker types.

Job Definition List​

The Job Definitions list displays all job definitions you own, organized in three sections:

SectionDescription
My DefinitionsJob definitions you have created
Shared by MeJob definitions you have shared
Shared with MeJob definitions others have shared with you

List Features​

  • List View / Cards View — Toggle between display modes
  • Search — Filter by name
  • Tag Filter — Filter by assigned color tag
  • Job Type Filter — Filter by job type

Actions​

Hover over a job definition entry to access:

ActionDescription
Edit (pencil icon)Opens the job definition editor
TriggerManually triggers the job immediately
CloneCreates a copy of the job definition
ShareOpens the sharing dialog
DeletePermanently removes the job definition

Creating a Job Definition​

  1. On the Job Definitions page, click the + (Plus) button.
  2. Fill in the configuration form.

Basic Settings​

FieldRequiredDescription
NameYesA unique, descriptive name for the job. Use consistent naming conventions for clarity.
DescriptionNoAn explanation of what the job does and why it exists.
Job TypeYesThe category of work this job performs (see Job Types below).
Schedule (Cron)NoA cron expression defining when the job should run automatically (e.g., 0 2 * * * for daily at 2am). Leave empty for manual-only execution.
Enabled—Toggle to activate or deactivate the job. Disabled jobs will not run on schedule.
Locked—Marks the job definition as read-only to prevent unintended edits.

Job Types​

TypeDescription
Admin TaskGeneral administrative operations (e.g., cleanup, system maintenance)
Incremental LoadLoads only new or changed data from a data source since the last run
Full LoadReloads all data from a data source from scratch
Semantic ProcessTriggers semantic extraction or layer construction for a data source

Defining Tasks Within a Job​

Each Job Definition must contain at least one task. Tasks are the individual execution units within a job. They run in order and can depend on each other.

Task Settings​

FieldDescription
Task IDAutomatically generated unique identifier for the task
Worker TypeThe execution context: Admin, Builder, Semantic, or User. Controls which permissions the task runs with.
Execution ScriptThe entry point script (run.sh) — a shell script containing all commands to execute.
Network AccessOutbound network access for this task. Internal only (default) reaches Lakehousecat services and object storage only. External HTTPS (port 443) additionally allows HTTPS downloads, for example a vendor driver. External (any port) allows any outbound port, for systems you operate yourself. Choose a wider profile only for the tasks that need it; tasks with external access are marked with a badge in the task graph.
ResourcesKubernetes resource allocation for the task (CPU, memory). Leave at defaults unless specific resources are required.
Environment VariablesKey-value pairs injected into the task environment at runtime.
SecretsSensitive values (API keys, passwords) stored securely and injected as environment variables.

Writing the Execution Script​

The main script must be named run.sh and must be executable. It serves as the entry point for the task.

Example — simple data refresh:

#!/bin/bash
set -e

echo "Starting data refresh..."
python /workspace/scripts/refresh_data.py
echo "Done."

Example — semantic extraction trigger:

#!/bin/bash
set -e

python /workspace/scripts/trigger_semantic_extraction.py --datasource-id "$DATASOURCE_ID"

Bringing Custom Data Into the Product (Staging)​

If your data lives behind a source Lakehousecat cannot connect to natively — a proprietary driver, an internal legacy system, an SAP export — you can write a Builder-worker task that fetches the data yourself and writes it into a dedicated staging database, lhc_staging. From there, connect a regular PostgreSQL data source to lhc_staging and the semantic layer picks it up like any other source.

This is the sanctioned handoff point between your own code and the product: your task never needs, and is never given, credentials to anything beyond that one staging database.

To use it, add a Secret to your Builder task with:

FieldValue
Secret namelhc-postgresql-staging-credentials-secret
Key in secretuser-password
Environment variableA name of your choice, e.g. STAGING_DB_PASSWORD

A Builder task can only reference this secret's user-password key (read-write, scoped to lhc_staging only) — not admin-password, and not as a volume mount. Connect with database lhc_staging, user lhc_staging_user, and the password from the environment variable you configured.


Vendor Driver Definitions​

Some databases need a native driver that Lakehousecat is not allowed to ship. For Microsoft SQL Server, Azure Synapse Analytics and IBM DB2, Lakehousecat provides reference job definitions that fetch the driver from the vendor once and then load through it. The reference definitions are published in the lhc-jobs repository of the lakehousecat GitHub organisation. See Vendor Drivers for the steps.

Scheduling with Cron​

Job schedules use standard cron syntax:

┌─ minute (0-59)
│ ┌─ hour (0-23)
│ │ ┌─ day of month (1-31)
│ │ │ ┌─ month (1-12)
│ │ │ │ ┌─ day of week (0-7, 0=Sunday)
│ │ │ │ │
* * * * *
ExampleMeaning
0 2 * * *Every day at 2:00 AM
0 */6 * * *Every 6 hours
0 9 * * 1Every Monday at 9:00 AM
0 0 1 * *First day of every month at midnight

Manually Triggering a Job​

To run a job outside its schedule:

  1. Find the job definition in the list.
  2. Click the Trigger button (or use More (⋯) → Trigger).

The job starts immediately and a new entry appears in Job Runs.


Job Runs​

The Job Runs tab shows the execution history for all job definitions.

Job Run List​

Each entry shows:

  • Job Name — The job definition that was executed
  • Status — Current state: running, success, failed, queued
  • Start Time — When the run began
  • Duration — How long the run took
  • Trigger — Whether the run was scheduled or manually triggered

Viewing Run Details​

Click on a job run entry to open the detail view:

  • Logs — Full execution output for each task in the run
  • Task Status — Individual task success or failure states
  • Error Messages — Details of any failures, useful for debugging

Sharing Job Definitions​

To share a job definition:

  1. Click More (⋯) on the job definition.
  2. Select Share.
  3. In the sharing dialog, select users or groups.
  4. Confirm to apply.

Shared job definitions appear in Shared by Me for the owner and Shared with Me for recipients. Recipients can view and trigger the job, but cannot modify it unless explicitly granted edit access.

To stop sharing:

  1. Go to the Shared by Me section.
  2. Click More (⋯) on the shared definition.
  3. Select Unshare.

Best Practices​

  • Test before scheduling: Use the manual trigger to verify a job runs correctly before enabling its cron schedule.
  • Use descriptive names: Names like daily-sales-incremental-load or weekly-semantic-refresh-crm are self-documenting.
  • Set appropriate worker types: Use the Semantic worker type for semantic extraction tasks and Admin for system-level operations to ensure correct permission scoping.
  • Monitor after deployment: Check Job Runs regularly after enabling a new scheduled job to catch failures early.
  • Use secrets for credentials: Never hardcode API keys or passwords in scripts. Use the Secrets field in the task configuration instead.