Skip to main content
Version: Next

IBM DB2

Connect an IBM DB2 database to Lakehousecat using a connection URI.

Connection Fields​

FieldRequiredDescription
Connection URIYesFull connection string. Example: db2+ibm_db://user:password@host:50000/database
Use SSHNoEnable if the database is behind an SSH bastion host. See SSH Tunneling.

Prerequisites​

  • Admin or Builder role in Lakehousecat
  • A Db2 user with SELECT access to the target schemas and tables
  • Default port: 50000
  • The IBM Db2 driver, provisioned once through a job definition. It is not preinstalled — see Provision the IBM driver below.
  • Your Lakehousecat instance must be deployed with architecture: "amd64". IBM does not ship a Linux/ARM64 driver for Db2, so this datasource type is unavailable on an arm64 instance. See the architecture field in your deployment configuration. Every other datasource type is unaffected by this — it is specific to Db2.

Notes​

  • Use a read-only database user to follow the principle of least privilege.
  • Incremental Load is supported; use a timestamp column (e.g., UPDATED_AT) as the merge key.
  • Db2 schema names are typically uppercase — ensure filter settings use the correct casing.

Provision the IBM driver (one-time)​

IBM Db2 is fully supported, but IBM's Db2 CLI driver is IBM's software under IBM's terms, so Lakehousecat does not ship it. You fetch it once through a job definition, and you accept IBM's licence terms while doing so. IBM's licence files are stored unchanged in the downloaded archive under clidriver/license/.

  1. Import the db2_driver_provisioning job definition, set LHC_ACCEPT_VENDOR_EULA to "yes" and run it. The job downloads IBM's ibm_db wheel from PyPI, verifies its SHA-256 checksum and stores the driver in File Storage. It prints a LHC_VENDOR_DRIVER_FILE_ID.
  2. Import db2_load and db2_schema_validate. In each, set DATASOURCE_ID, LHC_VENDOR_DRIVER_FILE_ID (from step 1) and LHC_ACCEPT_VENDOR_EULA to "yes".
  3. Load through db2_load, also on a schedule. Validate the connection and discover the schema through db2_schema_validate.

Requirements:

  • Outbound HTTPS from the provisioning job to pypi.org and files.pythonhosted.org. Without internet access, build the archive yourself (clidriver/ and ibm_db.libs/ side by side) and upload it with lhc files upload.
  • An amd64 instance. On arm64 the provisioning job stops with an explanation.

Validate and load calls made directly against the platform API don't work for Db2 and return a message that points to these job definitions. After a first successful load, the Filters tab reads the schema from your warehouse and needs no driver.

For the general mechanism, see Vendor Drivers.

Next Steps​

After provisioning the driver, load the datasource through the db2_load job definition, then open the Operations tab to trigger semantic extraction.