Skip to main content
Version: Next

Vendor Drivers

Some databases need a native driver that only the database vendor may distribute. Lakehousecat does not ship these drivers in its images. Instead, you fetch the driver once, from the vendor, and accept the vendor's licence terms yourself. Lakehousecat provides reference job definitions that do the fetching and then use the driver for loading. The reference job definitions live in the lhc-jobs repository of the lakehousecat GitHub organisation. They are examples to copy and adapt, not a package registry.

This applies to:

DatabaseDriverArchitectureProvisioning job definition
Microsoft SQL ServerMicrosoft ODBC Driver 18 for SQL Serveramd64, arm64msodbcsql_driver_provisioning
Azure Synapse AnalyticsMicrosoft ODBC Driver 18 for SQL Serveramd64, arm64azuresynapse_driver_provisioning
IBM DB2IBM Db2 CLI driveramd64 onlydb2_driver_provisioning

Every other datasource type works directly after you create it and needs none of the steps below.

note

The databases above are fully supported. The driver is not preinstalled because it is the vendor's software under the vendor's terms. You obtain it once, and you are the one who accepts those terms. The licence texts are yours to read; Lakehousecat neither grants nor checks them.

How it works​

Each database has three job definitions with the same prefix:

Job definitionPurpose
<prefix>_driver_provisioningFetches the driver from the vendor and stores it in your instance's File Storage. Run once per driver version. Reports a file ID.
<prefix>_loadLoads one datasource through the regular semantic pipeline, using the stored driver. Filters and incremental loading work as for any other datasource.
<prefix>_schema_validateValidates the connection and discovers the schema for one datasource, using the stored driver.

The prefixes are msodbcsql, azuresynapse and db2.

One-time setup​

  1. Import the provisioning job definition (see Job Definitions). Set LHC_ACCEPT_VENDOR_EULA to "yes" only after you have read and accepted the vendor's terms. The definition ships with "no", and the job refuses to run until you change it.
  2. Run it. The job downloads the driver, stores it in File Storage, and prints a LHC_VENDOR_DRIVER_FILE_ID.
  3. Import the load job definition and set three values in its environment variables: DATASOURCE_ID (the datasource to load), LHC_VENDOR_DRIVER_FILE_ID (from step 2) and LHC_ACCEPT_VENDOR_EULA ("yes").
  4. Run or schedule the load job. Do the same with <prefix>_schema_validate when you want to validate the connection or discover the schema.

Provisioning is a one-time step per instance. Loads can then run as often as you like, including on a schedule.

Network access​

The provisioning job needs outbound HTTPS access to the vendor's download location. The reference definitions declare this through egress_profile: external-https.

DatabaseDownload location
SQL Server, Azure Synapsepackages.microsoft.com
IBM DB2pypi.org and files.pythonhosted.org (IBM's driver is published inside the ibm_db wheel; no IBM account is needed)

The load and validate jobs need no extra egress rule, they only reach your database.

If your cluster has no internet access, build the driver archive yourself and upload it with lhc files upload, then use the resulting file ID.

Why validate and load fail outside a job​

Validating or loading one of these datasources directly through the platform API, for example with lhc semantic validate <id> or from the datasource form, returns a message that the native driver is not available in this container and points to the job definitions above. This is intended: a driver that Lakehousecat does not ship runs only inside a job task pod.

Once a load has succeeded, the Filters tab reads the schema from your warehouse, so it works without the driver.

Things to know​

  • The driver lives in File Storage. If you delete all files there, the driver is gone and the provisioning job has to run again; a load then fails with a 404 on the driver file. Running the provisioning job again stores the driver under a new file ID. Update LHC_VENDOR_DRIVER_FILE_ID in every load and validate definition that uses this driver, including a Synapse definition that reuses the SQL Server driver.
  • Azure Synapse uses the same Microsoft driver as SQL Server. Its provisioning definition is separate so that a Synapse-only setup doesn't depend on SQL Server naming. If you already provisioned the driver for SQL Server, you can reuse its file ID for Synapse.
  • Whoever imports or edits these definitions needs the Admin or Builder role.