Skip to main content
Version: Next

lhc semantic

Semantic extraction, hierarchies, and descriptions management.

Extraction Commands​

CommandDescription
lhc semantic create datasource <datasource-id>Start semantic extraction for a datasource
lhc semantic create model <model-id>Start semantic extraction for a model
lhc semantic update datasource <datasource-id>Re-run extraction for a datasource
lhc semantic update model <model-id>Re-run extraction for a model
lhc semantic delete datasource <datasource-id>Delete extraction data for a datasource (--confirm if used by a custom model)
lhc semantic delete model <model-id>Delete extraction data for a model (--confirm if it has dependent charts/dashboards)
lhc semantic clear datasource <datasource-id>Permanently delete all ClickHouse data for a datasource (requires --confirm)
lhc semantic load <datasource-id>Trigger data load for a datasource
lhc semantic schema <datasource-id>Get schema for a datasource
lhc semantic validate <datasource-id>Validate datasource connection and configuration
lhc semantic views <datasource-id>Get semantic views for a datasource
lhc semantic materialization <model-id>Show whether a model's star views are materialized; --enable switches them to refreshable materialized views (faster charts, refresh cost at load time), --disable back to plain views

ID types:

  • <datasource-id> — numeric datasource ID (e.g. 1220). Use lhc datasources list to find it.
  • <model-id> — the model's string identifier (e.g. my-sales-model), the same ID used in lhc models get <model-id>.

Async Commands​

create, update, and load are async (HTTP 202). The CLI prints the job ID on success:

started (job: datasource_load_42_a1b2c3d4)

Tracking progress​

create/update with --wait polls the job and prints one line per phase change — not a repeated line per poll interval. Progress is a phase counter (phase_index/phase_count), not an estimated percentage, plus a sub-counter for phases with a natural per-table loop (e.g. descriptions, star-view analysis):

job semantic_extraction_2_f9b5765d: status=running progress=1/7  Scanning schema
job semantic_extraction_2_f9b5765d: status=running progress=3/7 Building descriptions (35 of 35 tables)
job semantic_extraction_2_f9b5765d: status=completed progress=7/7 Semantic extraction completed successfully

Use this to tell a slow run from a stuck one: the job status also carries an updated_at timestamp of the last progress change, updated on every processed table within phases that have one, not just on phase transitions.

To check status without blocking, use the non-polling equivalents:

lhc semantic jobs get <job-id>
lhc semantic jobs cancel <job-id>

semantic jobs inspects these Redis-backed extraction/update jobs specifically — it is distinct from lhc jobs status <run-id>, which queries scheduled Airflow job runs (see Jobs).

Confirm Gate on Delete​

delete datasource and delete model check for dependents before removing anything. If the datasource is used by a custom model, or the model has charts or dashboards built on it, the command is blocked with an HTTP 409 and an impact report (affected chart/dashboard count, plus the full list of affected charts). Re-run the same command with --confirm to proceed anyway:

lhc semantic delete model my-model-id
# Error: This clean-up affects 3 chart(s) and 1 composed dashboard(s). Call again with confirmed=true to proceed.

lhc semantic delete model my-model-id --confirm

If there are no dependents, the command succeeds immediately without requiring --confirm.

Where these commands run​

A direct call — over the CLI or the REST API — runs the work in-process on the central semantic service, whose working directory is a 128 MiB in-memory disk. That is sized for short operations such as validate and schema. It is convenient for a quick manual run or during development, but a load of production-sized data will fail there as an out-of-memory kill or No space left on device.

For production loads, use the corresponding scheduled job instead (FULL_LOAD_DATASOURCE, INCREMENTAL_LOAD_DATASOURCE, CREATE_SEMANTIC_DATASOURCE, and the equivalents for models). Each runs in its own pod with disk-backed storage. See Jobs.

When you trigger one of these operations directly and a scheduled-job equivalent exists, the CLI prints an advisory to stderr and the job still starts:

WARN: This full_load is running in-process on the central semantic service, whose working
directory is a 128Mi tmpfs -- large loads fail there as an out-of-memory kill or 'No space
left on device'. Fine for a quick manual run; for production-sized data use the dedicated
Airflow job (FULL_LOAD_DATASOURCE), which executes in its own pod with disk-backed storage.

Standard output stays clean — 2>/dev/null leaves only the started (job: ...) line. With --output json, the advisory is not printed separately; it is carried in the warning field of the response payload. validate and schema have no scheduled-job equivalent and never carry this advisory.

Sub-Groups​