Skip to main content
Version: 0.0.42

Operations

The Operations tab controls which automated jobs are enabled for this data source. Each operation toggle creates a dedicated Job Definition in the Operations area — but does not start or schedule the job automatically.

You need Admin or Builder role to manage data source operations.


Available Operations​

OperationDescription
Full Data LoadEnables a job to perform a complete reload of all data from the source into ClickHouse. If the data source is linked to a Custom Model, the existing semantic layer and all dependent charts are preserved.
Incremental Data LoadEnables a job to load only new or changed data since the last run. The semantic layer and all dependent charts are always preserved.
Create Semantic LayerEnables a job to build the initial semantic layer for this data source
Update Semantic LayerEnables a job to update the existing semantic layer after schema or data changes
Delete Semantic LayerEnables a job to completely remove the semantic layer for this data source
Clear DatasourceEnables a job to remove all loaded data for this data source from ClickHouse

How Enabling an Operation Works​

Toggling an operation to enabled does the following:

  1. A Job Definition is created in the Operations area for that operation type.
  2. Once the Job Definition is ready, a link — Open Job Definition — appears directly below the toggle.
  3. Click the link to navigate to the Job Definition and configure when and how it should run.
important

Enabling an operation does not start the job and does not set a schedule. It only creates the Job Definition. The Builder or Administrator must explicitly open the Job Definition and decide:

  • Whether to run it manually (one-time trigger)
  • Whether to set a cron schedule for recurring execution

The timing and frequency of data loads or semantic extraction jobs depends on the specific data source and use case. This is a deliberate decision by the modeler — not an automatic process.


Configuring the Schedule​

After enabling an operation and clicking Open Job Definition:

  1. The Job Definition editor opens in the Operations section.
  2. Set the Schedule (Cron) field to define when the job should run automatically.
  3. Leave the schedule empty if you want to trigger the job manually only.
  4. Click Trigger Job to run it immediately for the first time.

For detailed instructions on cron schedules, task configuration, and manual triggering, see the Operations guide.


Choosing the Right Schedule​

There is no universal schedule — the right interval depends on your data source and business requirements:

ScenarioSuggested approach
Source data changes dailyIncremental Load daily, e.g. 0 2 * * *
Source schema rarely changesUpdate Semantic Layer weekly or on-demand
Initial setupRun Full Load and Create Semantic Layer manually once
High-frequency source dataIncremental Load multiple times per day
Static or infrequently updated sourceManual trigger only, no schedule

Data Source Status After Operations​

The data source list and detail view show the current state of each data source. After operations complete, the status reflects what has been done:

StatusWhat it means
CreatedData source connection is configured but no data has been loaded yet
LoadedA Full Load or Incremental Load has completed — data is available in ClickHouse
Semantic Layer CreatedSemantic extraction has completed — the data source is trained and ready to be used in a Custom Model
ValidatedThe data source has been reviewed and confirmed as production-ready

A data source must reach Semantic Layer Created status before it can be effectively used as a building block in a Custom Model. Only trained data sources produce meaningful analytical results.


How Incremental Load Works​

When you trigger an Incremental Data Load, the system loads only new or changed rows from the source into ClickHouse using merge keys to identify which rows to update.

Merge Key Detection​

Incremental loading requires merge keys — the columns that uniquely identify a row (typically a primary key like id). If merge keys are not yet configured for a data source, the system detects them automatically on the first incremental load attempt:

  1. Auto-detection (G3): The system inspects the source schema via SQLAlchemy reflection to find primary keys and unique constraints. An LLM agent refines results for tables where heuristic detection is ambiguous.
  2. Merge load: If merge keys are found, the load proceeds as an incremental merge — new rows are inserted and changed rows are replaced. Rows deleted in the source are not removed; see Incremental Load is insert/update-only below.
  3. Gentle Full fallback (G4): If no merge keys can be detected (for example, a ClickHouse source has no relational primary key metadata), the system falls back to a Gentle Full Load automatically.

Incremental Load is insert/update-only​

An Incremental Load never removes a row from ClickHouse. If a row is deleted in the source database, it stays in the target table indefinitely — counts and sums over that table remain permanently too high, and charts and materialized star views keep showing values that no longer exist at the source.

This is a structural property of incremental loading, not a configuration you can switch on:

  • An incremental read selects only rows where the cursor column is greater than the last recorded value. A DELETE leaves no row behind that could satisfy that condition, so the deletion never reaches the loader at all.
  • Detecting the deletion would require the source itself to record it — a soft-delete column such as deleted_at, or a change-data-capture stream. A source that hard-deletes its rows carries no such signal.

What to do instead:

SituationRecommended handling
Source only ever inserts and updates (append-style, event or log tables)Incremental Load is sufficient
Source deletes rows occasionallySchedule a periodic Full Load to reconcile — for example nightly Incremental, weekly Full
Source deletes rows frequently and stale rows are unacceptableUse Full Load for that data source, or model the deletions as soft deletes in the source

An update is only picked up if it moves the cursor column. A source that updates a row without touching its updated_at (or equivalent) is invisible to the incremental reader for exactly the same reason a delete is — keep the cursor column maintained, for example via a trigger.

Semantic Layer is Always Preserved​

Both incremental merge loads and the Gentle Full fallback preserve the semantic layer:

  • Table objects, view UUIDs, and all dependent chart references remain intact
  • Star views and materialized views are refreshed after the load completes — they are never dropped
  • The data source stays in Semantic Layer Created status throughout
tip

If auto-detection finds merge keys, they are persisted to the data source configuration automatically. You can inspect or adjust them in the data source settings. Subsequent incremental loads will use the detected keys directly without re-running detection.

ClickHouse as a source

ClickHouse does not expose primary key metadata via standard SQLAlchemy reflection. Lakehousecat reads merge keys from ClickHouse system tables (sorting_key / primary_key) and verifies uniqueness on the source. If no unique key is found — for example ORDER BY tuple() or AggregatingMergeTree — the system falls back to Gentle Full Load (G4). The connection user needs SELECT, SHOW, and OPTIMIZE on source tables (not SELECT alone). See ClickHouse datasource prerequisites.


When to Re-run Semantic Extraction​

Semantic extraction is resource-intensive and should not be run routinely. Re-run it only when the data structure changes — not when data values change:

ChangeAction needed
New tables or columns added to the sourceRe-run Update Semantic Layer
Existing columns renamed or removedRe-run Update Semantic Layer
Data values updated (new rows, changed values)Re-run data load only — no semantic re-extraction needed
Schema unchanged, data quality improvedOptionally re-run extraction; not required if structure is the same

If only data values change (which is the typical case in production), a new data load is sufficient. Semantic extraction does not need to be repeated.


Job History​

After jobs have been triggered, the Job Runs area in the Operations section shows execution history, status, and logs. Use this to verify that loads and extractions completed successfully and to diagnose failures.

Navigate to Operations → Job Runs or use the View Runs button in the Job Definition editor.