lhc datasources
Manage Lakehousecat semantic datasources. Datasources define the database connections that the AI uses to answer analytical questions.
Commands
datasources list
lhc datasources list [--type <type>]
| Flag | Description |
|---|---|
--type | Filter by datasource type, e.g. postgresql, mysql, clickhouse |
datasources get <id>
lhc datasources get <id>
datasources create
lhc datasources create \
--name <name> \
--type <datasource-type> \
--connection-string <connection-string>
| Flag | Required | Description |
|---|---|---|
--name | yes | Datasource name |
--type | yes | Datasource type (e.g. postgresql, mysql, clickhouse) |
--connection-string | yes | Full connection URI for the database |
For the full list of supported types, see Datasource Types.
datasources update <id>
Update a datasource's name, connection string, or state:
# Rename
lhc datasources update <id> --name "New Name"
# Correct a wrong connection string
lhc datasources update <id> --connection-string "postgresql://user:pass@host:5432/db"
# Disable a datasource
lhc datasources update <id> --is-enabled=false
# Lock a datasource (prevents editing and deletion)
lhc datasources update <id> --is-locked=true
# Combine multiple flags
lhc datasources update <id> --is-enabled=true --is-locked=false
| Flag | Description |
|---|---|
--name | New display name |
--connection-string | New connection URI for the datasource |
--is-enabled | Enable (true) or disable (false) the datasource — disabled datasources are unavailable for use |
--is-locked | Lock (true) or unlock (false) the datasource — locked datasources cannot be edited or deleted |
--is-visible | Control datasource visibility |
At least one flag must be provided. Flags accept =true or =false syntax (e.g. --is-locked=false).
datasources delete <id>
lhc datasources delete <id>
datasources share <id>
Share a datasource with a user or group. Exactly one of --user or --group is required:
lhc datasources share 1223 --user 3e7c3d4f-16ff-4914-8250-bbe693ab7224
lhc datasources share 1223 --group 1 --permission write
| Flag | Default | Description |
|---|---|---|
--user | — | Target user ID (UUID) |
--group | — | Target group ID (integer) |
--permission | read | Permission level: read, write, owner |
datasources unshare <id>
Remove a share from a datasource:
lhc datasources unshare 1223 --user 3e7c3d4f-16ff-4914-8250-bbe693ab7224
lhc datasources unshare 1223 --group 1
datasources search [query]
Search matches a substring of the datasource name or type. Connection settings are encrypted at rest and are never searched.
lhc datasources search "prod" [--type <type>] [--enabled]
lhc datasources search --tag RED # every red datasource, across all pages
| Flag | Description |
|---|---|
--type | Filter by datasource type, e.g. postgresql, mysql, clickhouse |
--enabled | Filter by enabled status — only applied when the flag is explicitly passed; omitting it returns enabled and disabled datasources alike |
--tag | Restrict to datasources tagged with this color (repeatable, OR-combined; valid colors: RED, YELLOW, GREEN, BLUE, PURPLE). Tags are per user and assigned in the UI — the CLI only filters by them |
The query is optional when --tag is given, but at least one of the two is required.
datasources statistics
lhc datasources statistics
Returns aggregate statistics across all datasources.
datasources export <id>
Export a datasource definition to a portable .lhc.json file:
# Print to stdout
lhc datasources export <id>
# Write to file
lhc datasources export <id> -f datasource.lhc.json
| Flag | Description |
|---|---|
-f, --output-file | Write output to this file (default: stdout) |
The export file can be imported into any Lakehousecat instance.
datasources import <file>
Import a datasource from a .lhc.json export file:
lhc datasources import datasource.lhc.json [--name <override-name>]
| Flag | Description |
|---|---|
--name | Override the datasource name from the export file |
--force-id <id> | Admin only. Pin the imported datasource to this exact numeric ID instead of letting the server assign the next free one. |
--skip-credential-check | Admin only. Import without requiring a valid connection URI. |
If a datasource with the same name already exists, the import is aborted with an HTTP 409 error. Use --name to import under a different name, or delete the existing datasource first.
Restoring alongside pre-migrated analytics data
--force-id and --skip-credential-check exist for one specific case: you have already restored a
datasource's analytical data into the warehouse (the datasource_<id> dataset) under a known ID, and
now need the datasource record to line up with it. A normal import assigns a new, sequential ID, which
no longer matches the warehouse dataset name or the view references built on top of it.
lhc datasources import datasource.lhc.json --force-id 17 --skip-credential-check
Use --skip-credential-check when the original connection is not reachable (or not needed, because
analytics reads only from the warehouse copy). Both flags require an admin API key.
If the export contains descriptions or hierarchies that cannot be re-created (for example schema drift between
versions), the datasource is still created. The command prints one warning: line per skipped item on stderr and exits
with a non-zero status. The warnings are returned only by the import call (import_warnings in the JSON response) and are not stored.
datasources upload <type> <id> <file>
Upload a local file for a file-based datasource (sqlite, duckdb):
lhc datasources upload sqlite 42 mydb.db
lhc datasources upload duckdb 42 mydb.duckdb
<type> must be sqlite or duckdb. This replaces the file backing that datasource.
Filters
Filters restrict what a datasource loads into ClickHouse — which schemas or tables are included,
which columns are selected per table, and row-level WHERE conditions. They take effect on the
next lhc semantic load — they do not retroactively change data already loaded. Use
lhc semantic schema <id> to see the tables and columns available to filter on before setting
anything.
datasources filters get <id>
lhc datasources filters get 765 --output json
Shows the filters currently set on the datasource.
datasources filters set <id>
lhc datasources filters set 765 --include-tables orders,customers
lhc datasources filters set 765 --exclude-tables '*_temp'
lhc datasources filters set 765 --filters-file ./filters.json
lhc datasources filters set 765 --filters-json '{"include_tables":["orders"]}'
lhc datasources filters set 765 --path-pattern 'data/2024/**.parquet' --path-recursive
lhc datasources filters set 765 --filters-file ./filters.json --replace
| Flag | Description |
|---|---|
--filters-json <json> | Full filter block as a JSON string |
--filters-file <path> | Full filter block read from a file (needed for column selections/filters — too deeply nested for shell quoting) |
--include-tables, --exclude-tables, --include-schemas, --exclude-schemas | Comfort flags for the common cases |
--path-pattern, --path-recursive | Object-storage path filtering (S3/GCS-style datasources) |
--replace | Discard the datasource's existing filters instead of merging into them |
By default, set merges the given filters into whatever the datasource already has — a field left
out keeps its previous value. --filters-json/--filters-file and the comfort flags can be
combined; where both set the same field, the comfort flag wins. The connection string and every
other connection setting are preserved automatically — you never need to re-supply
--connection-string just to change filters.
datasources filters clear <id>
lhc datasources filters clear 765 --confirm
Removes every filter from the datasource — the next lhc semantic load loads everything
unfiltered. Requires --confirm.
filters set/clear and incremental set/clear (below) both read the datasource, then write it
back with an optimistic-lock token (row_version). If another write landed on the same datasource
in between — e.g. someone ran filters set and incremental set around the same time — the second
write gets an HTTP 409 instead of silently discarding the first one. Re-run the command (it
re-reads the current state) if you hit this.
Incremental Load Settings
Incremental settings control merge-based loading: whether it is enabled, the default write
disposition, and per-table configuration (merge keys, incremental column, write disposition,
detection method). lhc semantic load --incremental only triggers the load — it requires these
settings to already exist, it does not create them.
datasources incremental get <id>
lhc datasources incremental get 765 --output json
Shows the incremental load settings currently set on the datasource.
datasources incremental set <id>
lhc datasources incremental set 765 --enabled
lhc datasources incremental set 765 --default-write-disposition merge
lhc datasources incremental set 765 --incremental-json '{"enabled":true,"table_configs":[{"table_name":"orders","merge_keys":["id"],"incremental_column":"updated_at"}]}'
lhc datasources incremental set 765 --incremental-file ./incremental.json
lhc datasources incremental set 765 --incremental-file ./incremental.json --replace
| Flag | Description |
|---|---|
--incremental-json <json> | Full incremental_settings block as a JSON string |
--incremental-file <path> | Full incremental_settings block read from a file (needed for table_configs — too deeply nested for shell quoting) |
--enabled | Comfort flag to enable/disable incremental loading |
--default-write-disposition | Default write disposition for tables without their own override (e.g. merge, append, replace) |
--replace | Discard the datasource's existing incremental settings instead of merging into them |
By default, set merges the given fields into whatever incremental_settings the datasource
already has — a field left out keeps its previous value. table_configs, when provided, replaces
the entire list — it is not merged per table. Where a comfort flag and --incremental-json/
--incremental-file set the same field, the comfort flag wins. The connection string and every
other connection setting are preserved automatically.
Applies only to connection-based datasource types (SQL databases, object storage, Delta/Hudi/
Iceberg) — file uploads and timeline datasources have no connection_settings to attach this to.
datasources incremental clear <id>
lhc datasources incremental clear 765 --confirm
Removes incremental settings from the datasource — lhc semantic load --incremental will no
longer have merge keys/write disposition to work with. Requires --confirm.
Operations
Toggle a datasource's automated jobs (admin/builder only). Each operation maps to a dedicated, per-datasource job definition — enabling it here is what creates that job definition.
datasources operations <id>
lhc datasources operations 765 --enable-clear
lhc datasources operations 765 --enable-full-load --disable-incremental-load
lhc datasources operations 765 --enable-create --enable-update --enable-delete
| Flag pair | Operation |
|---|---|
--enable-full-load / --disable-full-load | Full load |
--enable-incremental-load / --disable-incremental-load | Incremental load |
--enable-clear / --disable-clear | Clear datasource |
--enable-create / --disable-create | Create semantic layer |
--enable-update / --disable-update | Update semantic layer |
--enable-delete / --disable-delete | Delete semantic layer |
At least one flag pair must be given; unset operations are left unchanged. Once enabled, trigger
the resulting job definition like any other: lhc jobs trigger <job-definition-id> (its id comes
back in this command's own response).
The backend reports is_enabled=false either way, even when the underlying delete is silently
skipped (e.g. the job definition has an active run). If a job might still be triggerable after a
disable, verify with lhc jobs get <job-definition-id> rather than trusting this command's
response alone.
Relationships
Relationships connect columns across tables. They are auto-discovered from declared foreign keys or curated manually, and they feed the quality of generated charts and hierarchies. A manually curated relationship counts as ground truth and survives automated re-discovery.
datasources relationships schema <id>
Lists the tables and columns available for curation. Run this before creating a relationship — the loaded table identifiers are not necessarily the same as the names in the source system:
lhc datasources relationships schema 42 --output json
datasources relationships list <id>
Lists all relationships of a datasource, including disabled and orphaned ones:
lhc datasources relationships list 42 --output json
datasources relationships create <id>
Manually create a relationship between two tables:
lhc datasources relationships create 42 --table1 orders --column1 customer_id \
--table2 customers --column2 id --strength 0.9
| Flag | Required | Default | Description |
|---|---|---|---|
--table1 | yes | — | Table on one side of the relationship |
--column1 | yes | — | Column on the table1 side |
--table2 | yes | — | Table on the other side |
--column2 | yes | — | Column on the table2 side |
--strength | no | 1.0 | Weighting between 0.0 and 1.0 |
The four endpoints are the relationship's identity and cannot be changed later — to move a relationship to a different column, delete it and create a new one.
datasources relationships update <id> <relationship-id>
Change a relationship's strength or its enabled/locked flags. At least one flag is required:
lhc datasources relationships update 42 rel_abc123 --strength 0.5
lhc datasources relationships update 42 rel_abc123 --disable
lhc datasources relationships update 42 rel_abc123 --lock
| Flag | Description |
|---|---|
--strength | New weighting between 0.0 and 1.0 |
--enable / --disable | Disabling withholds the relationship from all consumers without deleting it |
--lock / --unlock | Lock the relationship against further edits |
--enable/--disable and --lock/--unlock are mutually exclusive pairs.
datasources relationships confirm <id> <relationship-id>
Confirm an automatically detected relationship. It becomes manual — protected from re-discovery by later semantic extractions — and its foreign-key column gets a curated foreign key. Which side is the key is taken from the data profile; if the data cannot tell, the relationship is confirmed without a foreign key. Views inherit the foreign key after the next semantic update and model update.
lhc datasources relationships confirm 42 52264696af909dcc
There is no un-confirm: to withdraw a confirmed relationship, delete it.
datasources relationships delete <id> <relationship-id>
lhc datasources relationships delete 42 rel_abc123
Cloning
datasources clone <id>
Clone a datasource's metadata only — name, type, connection settings, LLM settings. No descriptions, hierarchies, or data:
lhc datasources clone 42 --new-name "Prod PG copy"
| Flag | Required | Description |
|---|---|---|
--new-name | yes | Name for the cloned datasource |
datasources deep-clone <id>
Clone a datasource including its descriptions and hierarchies, and — unless --no-data is given —
its extracted schema and data as well:
lhc datasources deep-clone 42 --new-name "Prod PG full copy"
lhc datasources deep-clone 42 --new-name "Prod PG metadata only" --no-data
| Flag | Required | Description |
|---|---|---|
--new-name | yes | Name for the cloned datasource |
--no-data | no | Skip the data copy — metadata, descriptions, and hierarchies only |
For large datasources the data copy can take a while. The command returns once the clone has been
started, so poll datasources clone-status on the new datasource instead of treating the response
as completion.
datasources clone-status <id>
Status of a clone — use this to poll a running deep-clone:
lhc datasources clone-status 87
Sharing and counts
datasources sharing <id>
Show the sharing summary for one datasource — who it is shared with and at which permission level:
lhc datasources sharing 42
datasources shared-by-me / datasources shared-with-me
List datasources you have shared with others, or datasources others have shared with you:
lhc datasources shared-by-me
lhc datasources shared-with-me
datasources count
Total number of datasources:
lhc datasources count