Vendor Drivers
Some databases need a native driver that only the database vendor may distribute. Lakehousecat does
not ship these drivers in its images. Instead, you fetch the driver once, from the vendor, and accept
the vendor's licence terms yourself. Lakehousecat provides reference job definitions that do the
fetching and then use the driver for loading. The reference job definitions live in the
lhc-jobs repository of the lakehousecat GitHub organisation. They are examples to
copy and adapt, not a package registry.
This applies to:
| Database | Driver | Architecture | Provisioning job definition |
|---|---|---|---|
| Microsoft SQL Server | Microsoft ODBC Driver 18 for SQL Server | amd64, arm64 | msodbcsql_driver_provisioning |
| Azure Synapse Analytics | Microsoft ODBC Driver 18 for SQL Server | amd64, arm64 | azuresynapse_driver_provisioning |
| IBM DB2 | IBM Db2 CLI driver | amd64 only | db2_driver_provisioning |
Every other datasource type works directly after you create it and needs none of the steps below.
The databases above are fully supported. The driver is not preinstalled because it is the vendor's software under the vendor's terms. You obtain it once, and you are the one who accepts those terms. The licence texts are yours to read; Lakehousecat neither grants nor checks them.
How it works
Each database has three job definitions with the same prefix:
| Job definition | Purpose |
|---|---|
<prefix>_driver_provisioning | Fetches the driver from the vendor and stores it in your instance's File Storage. Run once per driver version. Reports a file ID. |
<prefix>_load | Loads one datasource through the regular semantic pipeline, using the stored driver. Filters and incremental loading work as for any other datasource. |
<prefix>_schema_validate | Validates the connection and discovers the schema for one datasource, using the stored driver. |
The prefixes are msodbcsql, azuresynapse and db2.
One-time setup
- Import the provisioning job definition (see Job Definitions).
Set
LHC_ACCEPT_VENDOR_EULAto"yes"only after you have read and accepted the vendor's terms. The definition ships with"no", and the job refuses to run until you change it. - Run it. The job downloads the driver, stores it in File Storage, and prints a
LHC_VENDOR_DRIVER_FILE_ID. - Import the load job definition and set three values in its environment variables:
DATASOURCE_ID(the datasource to load),LHC_VENDOR_DRIVER_FILE_ID(from step 2) andLHC_ACCEPT_VENDOR_EULA("yes"). - Run or schedule the load job. Do the same with
<prefix>_schema_validatewhen you want to validate the connection or discover the schema.
Provisioning is a one-time step per instance. Loads can then run as often as you like, including on a schedule.
Network access
The provisioning job needs outbound HTTPS access to the vendor's download location. The reference
definitions declare this through egress_profile: external-https.
| Database | Download location |
|---|---|
| SQL Server, Azure Synapse | packages.microsoft.com |
| IBM DB2 | pypi.org and files.pythonhosted.org (IBM's driver is published inside the ibm_db wheel; no IBM account is needed) |
The load and validate jobs need no extra egress rule, they only reach your database.
If your cluster has no internet access, build the driver archive yourself and upload it with
lhc files upload, then use the resulting file ID.
Why validate and load fail outside a job
Validating or loading one of these datasources directly through the platform API, for example with
lhc semantic validate <id> or from the datasource form, returns a message that the native driver
is not available in this container and points to the job definitions above. This is intended: a
driver that Lakehousecat does not ship runs only inside a job task pod.
Once a load has succeeded, the Filters tab reads the schema from your warehouse, so it works without the driver.
Things to know
- The driver lives in File Storage. If you delete all files there, the driver is gone and the
provisioning job has to run again; a load then fails with a
404on the driver file. Running the provisioning job again stores the driver under a new file ID. UpdateLHC_VENDOR_DRIVER_FILE_IDin every load and validate definition that uses this driver, including a Synapse definition that reuses the SQL Server driver. - Azure Synapse uses the same Microsoft driver as SQL Server. Its provisioning definition is separate so that a Synapse-only setup doesn't depend on SQL Server naming. If you already provisioned the driver for SQL Server, you can reuse its file ID for Synapse.
- Whoever imports or edits these definitions needs the Admin or Builder role.