Operations
The Operations section is the workflow automation layer of Lakehousecat. Here, administrators and builders can define, schedule, and monitor automated jobs — from data loads and semantic extractions to custom administrative tasks and data synchronization workflows.
Operations is built on an Apache Airflow workflow engine. Job scripts are written in Shell (run.sh) or Python, and tasks can be allocated to specific worker types with configurable Kubernetes resources.
Prerequisites
| Task | Required Role |
|---|---|
| Create and manage Job Definitions | Admin, Builder |
| View and monitor Job Runs | Admin, Builder |
Navigating to Operations
- Open the workspace.
- In the left sidebar, click Operations.
The Operations section contains two areas:
- Job Definitions — Define and configure automated jobs
- Job Runs — Monitor execution history and status of jobs
Job Definitions
A Job Definition is a named, reusable workflow configuration. It specifies what work should be done, how often, and by which worker types.
Job Definition List
The Job Definitions list displays all job definitions you own, organized in three sections:
| Section | Description |
|---|---|
| My Definitions | Job definitions you have created |
| Shared by Me | Job definitions you have shared |
| Shared with Me | Job definitions others have shared with you |
List Features
- List View / Cards View — Toggle between display modes
- Search — Filter by name
- Tag Filter — Filter by assigned color tag
- Job Type Filter — Filter by job type
Actions
Hover over a job definition entry to access:
| Action | Description |
|---|---|
| Edit (pencil icon) | Opens the job definition editor |
| Trigger | Manually triggers the job immediately |
| Clone | Creates a copy of the job definition |
| Share | Opens the sharing dialog |
| Delete | Permanently removes the job definition |
Creating a Job Definition
- On the Job Definitions page, click the + (Plus) button.
- Fill in the configuration form.
Basic Settings
| Field | Required | Description |
|---|---|---|
| Name | Yes | A unique, descriptive name for the job. Use consistent naming conventions for clarity. |
| Description | No | An explanation of what the job does and why it exists. |
| Job Type | Yes | The category of work this job performs (see Job Types below). |
| Schedule (Cron) | No | A cron expression defining when the job should run automatically (e.g., 0 2 * * * for daily at 2am). Leave empty for manual-only execution. |
| Enabled | — | Toggle to activate or deactivate the job. Disabled jobs will not run on schedule. |
| Locked | — | Marks the job definition as read-only to prevent unintended edits. |
Job Types
| Type | Description |
|---|---|
| Admin Task | General administrative operations (e.g., cleanup, system maintenance) |
| Incremental Load | Loads only new or changed data from a data source since the last run |
| Full Load | Reloads all data from a data source from scratch |
| Semantic Process | Triggers semantic extraction or layer construction for a data source |
Defining Tasks Within a Job
Each Job Definition must contain at least one task. Tasks are the individual execution units within a job. They run in order and can depend on each other.
Task Settings
| Field | Description |
|---|---|
| Task ID | Automatically generated unique identifier for the task |
| Worker Type | The execution context: Admin, Builder, Semantic, or User. Controls which permissions the task runs with. |
| Execution Script | The entry point script (run.sh) — a shell script containing all commands to execute. |
| Resources | Kubernetes resource allocation for the task (CPU, memory). Leave at defaults unless specific resources are required. |
| Environment Variables | Key-value pairs injected into the task environment at runtime. |
| Secrets | Sensitive values (API keys, passwords) stored securely and injected as environment variables. |
Writing the Execution Script
The main script must be named run.sh and must be executable. It serves as the entry point for the task.
Example — simple data refresh:
#!/bin/bash
set -e
echo "Starting data refresh..."
python /workspace/scripts/refresh_data.py
echo "Done."
Example — semantic extraction trigger:
#!/bin/bash
set -e
python /workspace/scripts/trigger_semantic_extraction.py --datasource-id "$DATASOURCE_ID"
Scheduling with Cron
Job schedules use standard cron syntax:
┌─ minute (0-59)
│ ┌─ hour (0-23)
│ │ ┌─ day of month (1-31)
│ │ │ ┌─ month (1-12)
│ │ │ │ ┌─ day of week (0-7, 0=Sunday)
│ │ │ │ │
* * * * *
| Example | Meaning |
|---|---|
0 2 * * * | Every day at 2:00 AM |
0 */6 * * * | Every 6 hours |
0 9 * * 1 | Every Monday at 9:00 AM |
0 0 1 * * | First day of every month at midnight |
Manually Triggering a Job
To run a job outside its schedule:
- Find the job definition in the list.
- Click the Trigger button (or use More (⋯) → Trigger).
The job starts immediately and a new entry appears in Job Runs.
Job Runs
The Job Runs tab shows the execution history for all job definitions.
Job Run List
Each entry shows:
- Job Name — The job definition that was executed
- Status — Current state:
running,success,failed,queued - Start Time — When the run began
- Duration — How long the run took
- Trigger — Whether the run was scheduled or manually triggered
Viewing Run Details
Click on a job run entry to open the detail view:
- Logs — Full execution output for each task in the run
- Task Status — Individual task success or failure states
- Error Messages — Details of any failures, useful for debugging
Sharing Job Definitions
To share a job definition:
- Click More (⋯) on the job definition.
- Select Share.
- In the sharing dialog, select users or groups.
- Confirm to apply.
Shared job definitions appear in Shared by Me for the owner and Shared with Me for recipients. Recipients can view and trigger the job, but cannot modify it unless explicitly granted edit access.
To stop sharing:
- Go to the Shared by Me section.
- Click More (⋯) on the shared definition.
- Select Unshare.
Best Practices
- Test before scheduling: Use the manual trigger to verify a job runs correctly before enabling its cron schedule.
- Use descriptive names: Names like
daily-sales-incremental-loadorweekly-semantic-refresh-crmare self-documenting. - Set appropriate worker types: Use the
Semanticworker type for semantic extraction tasks andAdminfor system-level operations to ensure correct permission scoping. - Monitor after deployment: Check Job Runs regularly after enabling a new scheduled job to catch failures early.
- Use secrets for credentials: Never hardcode API keys or passwords in scripts. Use the Secrets field in the task configuration instead.