lhc semantic
Semantic extraction, hierarchies, and descriptions management.
Extraction Commands
| Command | Description |
|---|---|
lhc semantic create datasource <datasource-id> | Start semantic extraction for a datasource |
lhc semantic create model <model-id> | Start semantic extraction for a model |
lhc semantic update datasource <datasource-id> | Re-run extraction for a datasource |
lhc semantic update model <model-id> | Re-run extraction for a model |
lhc semantic delete datasource <datasource-id> | Delete extraction data for a datasource (--confirm if used by a custom model) |
lhc semantic delete model <model-id> | Delete extraction data for a model (--confirm if it has dependent charts/dashboards) |
lhc semantic clear datasource <datasource-id> | Permanently delete all ClickHouse data for a datasource (requires --confirm) |
lhc semantic load <datasource-id> | Trigger data load for a datasource |
lhc semantic schema <datasource-id> | Get schema for a datasource |
lhc semantic validate <datasource-id> | Validate datasource connection and configuration |
lhc semantic views <datasource-id> | Get semantic views for a datasource |
lhc semantic materialization <model-id> | Show whether a model's star views are materialized; --enable switches them to refreshable materialized views (faster charts, refresh cost at load time), --disable back to plain views |
ID types:
<datasource-id>— numeric datasource ID (e.g.1220). Uselhc datasources listto find it.<model-id>— the model's string identifier (e.g.my-sales-model), the same ID used inlhc models get <model-id>.
Async Commands
create, update, and load are async (HTTP 202). The CLI prints the job ID on success:
started (job: datasource_load_42_a1b2c3d4)
Tracking progress
create/update with --wait polls the job and prints one line per phase change — not a
repeated line per poll interval. Progress is a phase counter (phase_index/phase_count), not an
estimated percentage, plus a sub-counter for phases with a natural per-table loop (e.g.
descriptions, star-view analysis):
job semantic_extraction_2_f9b5765d: status=running progress=1/7 Scanning schema
job semantic_extraction_2_f9b5765d: status=running progress=3/7 Building descriptions (35 of 35 tables)
job semantic_extraction_2_f9b5765d: status=completed progress=7/7 Semantic extraction completed successfully
Use this to tell a slow run from a stuck one: the job status also carries an updated_at
timestamp of the last progress change, updated on every processed table within phases that have
one, not just on phase transitions.
To check status without blocking, use the non-polling equivalents:
lhc semantic jobs get <job-id>
lhc semantic jobs cancel <job-id>
semantic jobs inspects these Redis-backed extraction/update jobs specifically — it is distinct
from lhc jobs status <run-id>, which queries scheduled Airflow job runs (see Jobs).
Confirm Gate on Delete
delete datasource and delete model check for dependents before removing anything. If the
datasource is used by a custom model, or the model has charts or dashboards built on it, the
command is blocked with an HTTP 409 and an impact report (affected chart/dashboard count, plus
the full list of affected charts). Re-run the same command with --confirm to proceed anyway:
lhc semantic delete model my-model-id
# Error: This clean-up affects 3 chart(s) and 1 composed dashboard(s). Call again with confirmed=true to proceed.
lhc semantic delete model my-model-id --confirm
If there are no dependents, the command succeeds immediately without requiring --confirm.
Where these commands run
A direct call — over the CLI or the REST API — runs the work in-process on the central
semantic service, whose working directory is a 128 MiB in-memory disk. That is sized for
short operations such as validate and schema. It is convenient for a quick manual run or
during development, but a load of production-sized data will fail there as an out-of-memory
kill or No space left on device.
For production loads, use the corresponding scheduled job instead (FULL_LOAD_DATASOURCE,
INCREMENTAL_LOAD_DATASOURCE, CREATE_SEMANTIC_DATASOURCE, and the equivalents for models).
Each runs in its own pod with disk-backed storage. See Jobs.
When you trigger one of these operations directly and a scheduled-job equivalent exists, the CLI prints an advisory to stderr and the job still starts:
WARN: This full_load is running in-process on the central semantic service, whose working
directory is a 128Mi tmpfs -- large loads fail there as an out-of-memory kill or 'No space
left on device'. Fine for a quick manual run; for production-sized data use the dedicated
Airflow job (FULL_LOAD_DATASOURCE), which executes in its own pod with disk-backed storage.
Standard output stays clean — 2>/dev/null leaves only the started (job: ...) line. With
--output json, the advisory is not printed separately; it is carried in the warning field
of the response payload. validate and schema have no scheduled-job equivalent and never
carry this advisory.
Sub-Groups
- datasources — Create, manage, export, and import datasources
- hierarchies — Manage semantic hierarchies
- descriptions — Manage semantic descriptions