Operations
The Operations tab controls which automated jobs are enabled for this data source. Each operation toggle creates a dedicated Job Definition in the Operations area — but does not start or schedule the job automatically.
You need Admin or Builder role to manage data source operations.
Available Operations
| Operation | Description |
|---|---|
| Full Data Load | Enables a job to perform a complete reload of all data from the source into ClickHouse. If the data source is linked to a Custom Model, the existing semantic layer and all dependent charts are preserved. |
| Incremental Data Load | Enables a job to load only new or changed data since the last run. The semantic layer and all dependent charts are always preserved. |
| Create Semantic Layer | Enables a job to build the initial semantic layer for this data source |
| Update Semantic Layer | Enables a job to update the existing semantic layer after schema or data changes |
| Delete Semantic Layer | Enables a job to completely remove the semantic layer for this data source |
| Clear Datasource | Enables a job to remove all loaded data for this data source from ClickHouse |
How Enabling an Operation Works
Toggling an operation to enabled does the following:
- A Job Definition is created in the Operations area for that operation type.
- Once the Job Definition is ready, a link — Open Job Definition — appears directly below the toggle.
- Click the link to navigate to the Job Definition and configure when and how it should run.
Enabling an operation does not start the job and does not set a schedule. It only creates the Job Definition. The Builder or Administrator must explicitly open the Job Definition and decide:
- Whether to run it manually (one-time trigger)
- Whether to set a cron schedule for recurring execution
The timing and frequency of data loads or semantic extraction jobs depends on the specific data source and use case. This is a deliberate decision by the modeler — not an automatic process.
Configuring the Schedule
After enabling an operation and clicking Open Job Definition:
- The Job Definition editor opens in the Operations section.
- Set the Schedule (Cron) field to define when the job should run automatically.
- Leave the schedule empty if you want to trigger the job manually only.
- Click Trigger Job to run it immediately for the first time.
For detailed instructions on cron schedules, task configuration, and manual triggering, see the Operations guide.
Choosing the Right Schedule
There is no universal schedule — the right interval depends on your data source and business requirements:
| Scenario | Suggested approach |
|---|---|
| Source data changes daily | Incremental Load daily, e.g. 0 2 * * * |
| Source schema rarely changes | Update Semantic Layer weekly or on-demand |
| Initial setup | Run Full Load and Create Semantic Layer manually once |
| High-frequency source data | Incremental Load multiple times per day |
| Static or infrequently updated source | Manual trigger only, no schedule |
Data Source Status After Operations
The data source list and detail view show the current state of each data source. After operations complete, the status reflects what has been done:
| Status | What it means |
|---|---|
| Created | Data source connection is configured but no data has been loaded yet |
| Loaded | A Full Load or Incremental Load has completed — data is available in ClickHouse |
| Semantic Layer Created | Semantic extraction has completed — the data source is trained and ready to be used in a Custom Model |
| Validated | The data source has been reviewed and confirmed as production-ready |
A data source must reach Semantic Layer Created status before it can be effectively used as a building block in a Custom Model. Only trained data sources produce meaningful analytical results.
How Incremental Load Works
When you trigger an Incremental Data Load, the system loads only new or changed rows from the source into ClickHouse using merge keys to identify which rows to update.
Merge Key Detection
Incremental loading requires merge keys — the columns that uniquely identify a row (typically a primary key like id). If merge keys are not yet configured for a data source, the system detects them automatically on the first incremental load attempt:
- Auto-detection (G3): The system inspects the source schema via SQLAlchemy reflection to find primary keys and unique constraints. An LLM agent refines results for tables where heuristic detection is ambiguous.
- Merge load: If merge keys are found, the load proceeds as an incremental merge — new rows are inserted and changed rows are replaced. Rows deleted in the source are not removed; see Incremental Load is insert/update-only below.
- Gentle Full fallback (G4): If no merge keys can be detected (for example, a ClickHouse source has no relational primary key metadata), the system falls back to a Gentle Full Load automatically.
Incremental Load is insert/update-only
An Incremental Load never removes a row from ClickHouse. If a row is deleted in the source database, it stays in the target table indefinitely — counts and sums over that table remain permanently too high, and charts and materialized star views keep showing values that no longer exist at the source.
This is a structural property of incremental loading, not a configuration you can switch on:
- An incremental read selects only rows where the cursor column is greater than the last recorded value. A
DELETEleaves no row behind that could satisfy that condition, so the deletion never reaches the loader at all. - Detecting the deletion would require the source itself to record it — a soft-delete column such as
deleted_at, or a change-data-capture stream. A source that hard-deletes its rows carries no such signal.
What to do instead:
| Situation | Recommended handling |
|---|---|
| Source only ever inserts and updates (append-style, event or log tables) | Incremental Load is sufficient |
| Source deletes rows occasionally | Schedule a periodic Full Load to reconcile — for example nightly Incremental, weekly Full |
| Source deletes rows frequently and stale rows are unacceptable | Use Full Load for that data source, or model the deletions as soft deletes in the source |
An update is only picked up if it moves the cursor column. A source that updates a row without touching its updated_at (or equivalent) is invisible to the incremental reader for exactly the same reason a delete is — keep the cursor column maintained, for example via a trigger.
Semantic Layer is Always Preserved
Both incremental merge loads and the Gentle Full fallback preserve the semantic layer:
- Table objects, view UUIDs, and all dependent chart references remain intact
- Star views and materialized views are refreshed after the load completes — they are never dropped
- The data source stays in Semantic Layer Created status throughout
If auto-detection finds merge keys, they are persisted to the data source configuration automatically. You can inspect or adjust them in the data source settings. Subsequent incremental loads will use the detected keys directly without re-running detection.
ClickHouse does not expose primary key metadata via standard SQLAlchemy reflection. Lakehousecat reads merge keys from ClickHouse system tables (sorting_key / primary_key) and verifies uniqueness on the source. If no unique key is found — for example ORDER BY tuple() or AggregatingMergeTree — the system falls back to Gentle Full Load (G4). The connection user needs SELECT, SHOW, and OPTIMIZE on source tables (not SELECT alone). See ClickHouse datasource prerequisites.
When to Re-run Semantic Extraction
Semantic extraction is resource-intensive and should not be run routinely. Re-run it only when the data structure changes — not when data values change:
| Change | Action needed |
|---|---|
| New tables or columns added to the source | Re-run Update Semantic Layer |
| Existing columns renamed or removed | Re-run Update Semantic Layer |
| Data values updated (new rows, changed values) | Re-run data load only — no semantic re-extraction needed |
| Schema unchanged, data quality improved | Optionally re-run extraction; not required if structure is the same |
If only data values change (which is the typical case in production), a new data load is sufficient. Semantic extraction does not need to be repeated.
Job History
After jobs have been triggered, the Job Runs area in the Operations section shows execution history, status, and logs. Use this to verify that loads and extractions completed successfully and to diagnose failures.
Navigate to Operations → Job Runs or use the View Runs button in the Job Definition editor.