Skip to main content
Version: Next

lhc datasources

Manage Lakehousecat semantic datasources. Datasources define the database connections that the AI uses to answer analytical questions.

Commands​

datasources list​

lhc datasources list [--type <type>]
FlagDescription
--typeFilter by datasource type, e.g. postgresql, mysql, clickhouse

datasources get <id>​

lhc datasources get <id>

datasources create​

lhc datasources create \
--name <name> \
--type <datasource-type> \
--connection-string <connection-string>
FlagRequiredDescription
--nameyesDatasource name
--typeyesDatasource type (e.g. postgresql, mysql, clickhouse)
--connection-stringyesFull connection URI for the database

For the full list of supported types, see Datasource Types.


datasources update <id>​

Update a datasource's name, connection string, or state:

# Rename
lhc datasources update <id> --name "New Name"

# Correct a wrong connection string
lhc datasources update <id> --connection-string "postgresql://user:pass@host:5432/db"

# Disable a datasource
lhc datasources update <id> --is-enabled=false

# Lock a datasource (prevents editing and deletion)
lhc datasources update <id> --is-locked=true

# Combine multiple flags
lhc datasources update <id> --is-enabled=true --is-locked=false
FlagDescription
--nameNew display name
--connection-stringNew connection URI for the datasource
--is-enabledEnable (true) or disable (false) the datasource — disabled datasources are unavailable for use
--is-lockedLock (true) or unlock (false) the datasource — locked datasources cannot be edited or deleted
--is-visibleControl datasource visibility

At least one flag must be provided. Flags accept =true or =false syntax (e.g. --is-locked=false).


datasources delete <id>​

lhc datasources delete <id>

datasources share <id>​

Share a datasource with a user or group. Exactly one of --user or --group is required:

lhc datasources share 1223 --user 3e7c3d4f-16ff-4914-8250-bbe693ab7224
lhc datasources share 1223 --group 1 --permission write
FlagDefaultDescription
--user—Target user ID (UUID)
--group—Target group ID (integer)
--permissionreadPermission level: read, write, owner

datasources unshare <id>​

Remove a share from a datasource:

lhc datasources unshare 1223 --user 3e7c3d4f-16ff-4914-8250-bbe693ab7224
lhc datasources unshare 1223 --group 1

datasources search [query]​

Search matches a substring of the datasource name or type. Connection settings are encrypted at rest and are never searched.

lhc datasources search "prod" [--type <type>] [--enabled]
lhc datasources search --tag RED # every red datasource, across all pages
FlagDescription
--typeFilter by datasource type, e.g. postgresql, mysql, clickhouse
--enabledFilter by enabled status — only applied when the flag is explicitly passed; omitting it returns enabled and disabled datasources alike
--tagRestrict to datasources tagged with this color (repeatable, OR-combined; valid colors: RED, YELLOW, GREEN, BLUE, PURPLE). Tags are per user and assigned in the UI — the CLI only filters by them

The query is optional when --tag is given, but at least one of the two is required.


datasources statistics​

lhc datasources statistics

Returns aggregate statistics across all datasources.


datasources export <id>​

Export a datasource definition to a portable .lhc.json file:

# Print to stdout
lhc datasources export <id>

# Write to file
lhc datasources export <id> -f datasource.lhc.json
FlagDescription
-f, --output-fileWrite output to this file (default: stdout)

The export file can be imported into any Lakehousecat instance.


datasources import <file>​

Import a datasource from a .lhc.json export file:

lhc datasources import datasource.lhc.json [--name <override-name>]
FlagDescription
--nameOverride the datasource name from the export file
--force-id <id>Admin only. Pin the imported datasource to this exact numeric ID instead of letting the server assign the next free one.
--skip-credential-checkAdmin only. Import without requiring a valid connection URI.

If a datasource with the same name already exists, the import is aborted with an HTTP 409 error. Use --name to import under a different name, or delete the existing datasource first.

Restoring alongside pre-migrated analytics data​

--force-id and --skip-credential-check exist for one specific case: you have already restored a datasource's analytical data into the warehouse (the datasource_<id> dataset) under a known ID, and now need the datasource record to line up with it. A normal import assigns a new, sequential ID, which no longer matches the warehouse dataset name or the view references built on top of it.

lhc datasources import datasource.lhc.json --force-id 17 --skip-credential-check

Use --skip-credential-check when the original connection is not reachable (or not needed, because analytics reads only from the warehouse copy). Both flags require an admin API key.

If the export contains descriptions or hierarchies that cannot be re-created (for example schema drift between versions), the datasource is still created. The command prints one warning: line per skipped item on stderr and exits with a non-zero status. The warnings are returned only by the import call (import_warnings in the JSON response) and are not stored.


datasources upload <type> <id> <file>​

Upload a local file for a file-based datasource (sqlite, duckdb):

lhc datasources upload sqlite 42 mydb.db
lhc datasources upload duckdb 42 mydb.duckdb

<type> must be sqlite or duckdb. This replaces the file backing that datasource.


Filters​

Filters restrict what a datasource loads into ClickHouse — which schemas or tables are included, which columns are selected per table, and row-level WHERE conditions. They take effect on the next lhc semantic load — they do not retroactively change data already loaded. Use lhc semantic schema <id> to see the tables and columns available to filter on before setting anything.

datasources filters get <id>​

lhc datasources filters get 765 --output json

Shows the filters currently set on the datasource.

datasources filters set <id>​

lhc datasources filters set 765 --include-tables orders,customers
lhc datasources filters set 765 --exclude-tables '*_temp'
lhc datasources filters set 765 --filters-file ./filters.json
lhc datasources filters set 765 --filters-json '{"include_tables":["orders"]}'
lhc datasources filters set 765 --path-pattern 'data/2024/**.parquet' --path-recursive
lhc datasources filters set 765 --filters-file ./filters.json --replace
FlagDescription
--filters-json <json>Full filter block as a JSON string
--filters-file <path>Full filter block read from a file (needed for column selections/filters — too deeply nested for shell quoting)
--include-tables, --exclude-tables, --include-schemas, --exclude-schemasComfort flags for the common cases
--path-pattern, --path-recursiveObject-storage path filtering (S3/GCS-style datasources)
--replaceDiscard the datasource's existing filters instead of merging into them

By default, set merges the given filters into whatever the datasource already has — a field left out keeps its previous value. --filters-json/--filters-file and the comfort flags can be combined; where both set the same field, the comfort flag wins. The connection string and every other connection setting are preserved automatically — you never need to re-supply --connection-string just to change filters.

datasources filters clear <id>​

lhc datasources filters clear 765 --confirm

Removes every filter from the datasource — the next lhc semantic load loads everything unfiltered. Requires --confirm.

Concurrent writes are rejected, not silently overwritten

filters set/clear and incremental set/clear (below) both read the datasource, then write it back with an optimistic-lock token (row_version). If another write landed on the same datasource in between — e.g. someone ran filters set and incremental set around the same time — the second write gets an HTTP 409 instead of silently discarding the first one. Re-run the command (it re-reads the current state) if you hit this.

Incremental Load Settings​

Incremental settings control merge-based loading: whether it is enabled, the default write disposition, and per-table configuration (merge keys, incremental column, write disposition, detection method). lhc semantic load --incremental only triggers the load — it requires these settings to already exist, it does not create them.

datasources incremental get <id>​

lhc datasources incremental get 765 --output json

Shows the incremental load settings currently set on the datasource.

datasources incremental set <id>​

lhc datasources incremental set 765 --enabled
lhc datasources incremental set 765 --default-write-disposition merge
lhc datasources incremental set 765 --incremental-json '{"enabled":true,"table_configs":[{"table_name":"orders","merge_keys":["id"],"incremental_column":"updated_at"}]}'
lhc datasources incremental set 765 --incremental-file ./incremental.json
lhc datasources incremental set 765 --incremental-file ./incremental.json --replace
FlagDescription
--incremental-json <json>Full incremental_settings block as a JSON string
--incremental-file <path>Full incremental_settings block read from a file (needed for table_configs — too deeply nested for shell quoting)
--enabledComfort flag to enable/disable incremental loading
--default-write-dispositionDefault write disposition for tables without their own override (e.g. merge, append, replace)
--replaceDiscard the datasource's existing incremental settings instead of merging into them

By default, set merges the given fields into whatever incremental_settings the datasource already has — a field left out keeps its previous value. table_configs, when provided, replaces the entire list — it is not merged per table. Where a comfort flag and --incremental-json/ --incremental-file set the same field, the comfort flag wins. The connection string and every other connection setting are preserved automatically.

Applies only to connection-based datasource types (SQL databases, object storage, Delta/Hudi/ Iceberg) — file uploads and timeline datasources have no connection_settings to attach this to.

datasources incremental clear <id>​

lhc datasources incremental clear 765 --confirm

Removes incremental settings from the datasource — lhc semantic load --incremental will no longer have merge keys/write disposition to work with. Requires --confirm.

Operations​

Toggle a datasource's automated jobs (admin/builder only). Each operation maps to a dedicated, per-datasource job definition — enabling it here is what creates that job definition.

datasources operations <id>​

lhc datasources operations 765 --enable-clear
lhc datasources operations 765 --enable-full-load --disable-incremental-load
lhc datasources operations 765 --enable-create --enable-update --enable-delete
Flag pairOperation
--enable-full-load / --disable-full-loadFull load
--enable-incremental-load / --disable-incremental-loadIncremental load
--enable-clear / --disable-clearClear datasource
--enable-create / --disable-createCreate semantic layer
--enable-update / --disable-updateUpdate semantic layer
--enable-delete / --disable-deleteDelete semantic layer

At least one flag pair must be given; unset operations are left unchanged. Once enabled, trigger the resulting job definition like any other: lhc jobs trigger <job-definition-id> (its id comes back in this command's own response).

Disable reports success even when the job definition survives

The backend reports is_enabled=false either way, even when the underlying delete is silently skipped (e.g. the job definition has an active run). If a job might still be triggerable after a disable, verify with lhc jobs get <job-definition-id> rather than trusting this command's response alone.

Relationships​

Relationships connect columns across tables. They are auto-discovered from declared foreign keys or curated manually, and they feed the quality of generated charts and hierarchies. A manually curated relationship counts as ground truth and survives automated re-discovery.

datasources relationships schema <id>​

Lists the tables and columns available for curation. Run this before creating a relationship — the loaded table identifiers are not necessarily the same as the names in the source system:

lhc datasources relationships schema 42 --output json

datasources relationships list <id>​

Lists all relationships of a datasource, including disabled and orphaned ones:

lhc datasources relationships list 42 --output json

datasources relationships create <id>​

Manually create a relationship between two tables:

lhc datasources relationships create 42 --table1 orders --column1 customer_id \
--table2 customers --column2 id --strength 0.9
FlagRequiredDefaultDescription
--table1yes—Table on one side of the relationship
--column1yes—Column on the table1 side
--table2yes—Table on the other side
--column2yes—Column on the table2 side
--strengthno1.0Weighting between 0.0 and 1.0

The four endpoints are the relationship's identity and cannot be changed later — to move a relationship to a different column, delete it and create a new one.

datasources relationships update <id> <relationship-id>​

Change a relationship's strength or its enabled/locked flags. At least one flag is required:

lhc datasources relationships update 42 rel_abc123 --strength 0.5
lhc datasources relationships update 42 rel_abc123 --disable
lhc datasources relationships update 42 rel_abc123 --lock
FlagDescription
--strengthNew weighting between 0.0 and 1.0
--enable / --disableDisabling withholds the relationship from all consumers without deleting it
--lock / --unlockLock the relationship against further edits

--enable/--disable and --lock/--unlock are mutually exclusive pairs.

datasources relationships confirm <id> <relationship-id>​

Confirm an automatically detected relationship. It becomes manual — protected from re-discovery by later semantic extractions — and its foreign-key column gets a curated foreign key. Which side is the key is taken from the data profile; if the data cannot tell, the relationship is confirmed without a foreign key. Views inherit the foreign key after the next semantic update and model update.

lhc datasources relationships confirm 42 52264696af909dcc

There is no un-confirm: to withdraw a confirmed relationship, delete it.

datasources relationships delete <id> <relationship-id>​

lhc datasources relationships delete 42 rel_abc123

Cloning​

datasources clone <id>​

Clone a datasource's metadata only — name, type, connection settings, LLM settings. No descriptions, hierarchies, or data:

lhc datasources clone 42 --new-name "Prod PG copy"
FlagRequiredDescription
--new-nameyesName for the cloned datasource

datasources deep-clone <id>​

Clone a datasource including its descriptions and hierarchies, and — unless --no-data is given — its extracted schema and data as well:

lhc datasources deep-clone 42 --new-name "Prod PG full copy"
lhc datasources deep-clone 42 --new-name "Prod PG metadata only" --no-data
FlagRequiredDescription
--new-nameyesName for the cloned datasource
--no-datanoSkip the data copy — metadata, descriptions, and hierarchies only

For large datasources the data copy can take a while. The command returns once the clone has been started, so poll datasources clone-status on the new datasource instead of treating the response as completion.

datasources clone-status <id>​

Status of a clone — use this to poll a running deep-clone:

lhc datasources clone-status 87

Sharing and counts​

datasources sharing <id>​

Show the sharing summary for one datasource — who it is shared with and at which permission level:

lhc datasources sharing 42

datasources shared-by-me / datasources shared-with-me​

List datasources you have shared with others, or datasources others have shared with you:

lhc datasources shared-by-me
lhc datasources shared-with-me

datasources count​

Total number of datasources:

lhc datasources count