Skip to main content
Version: Next

Apache Iceberg

Connect an Apache Iceberg table to Lakehousecat either directly by its metadata location in S3-compatible object storage, or through a REST catalog (e.g. Nessie, Polaris, Lakekeeper).

Connection Fields​

FieldRequiredDescription
URLYes (metadata path)Full URL of the Iceberg table in object storage. Example: http://seaweedfs-s3:8333/iceberg/tablename
Access Key IDYesThe access key for object storage authentication
Secret Access KeyYesThe secret key for object storage authentication
S3 EndpointNoCustom S3-compatible endpoint address (MinIO, Ceph, Wasabi, Hetzner, etc.). Leave empty for AWS. Applies to both the metadata path and the catalog path.
Catalog URINoREST catalog endpoint, e.g. http://nessie-host:19120/iceberg/main. Include the branch/ref in the path where the catalog supports it (Nessie: /iceberg/<branch>). Leave empty to use the metadata path instead.
WarehouseNoWarehouse identifier as configured on the catalog server. Only used when Catalog URI is set.
TokenNoBearer token for catalog authentication, where required. Masked after saving.

Two Ways to Connect​

  • Metadata path (default): point directly at the table's metadata directory in object storage. Works with any Hadoop-style warehouse, no catalog infrastructure needed.
  • Catalog path: set Catalog URI (and optionally Warehouse/Token) to resolve tables through a REST catalog instead. Namespaces are folded into the table name (e.g. lhc.customers → lhc__customers); the catalog is the source of truth for table identity and the current snapshot, not the file layout.

Prerequisites​

  • Admin or Builder role in Lakehousecat
  • An Apache Iceberg table accessible via an S3-compatible endpoint, or a reachable REST catalog
  • Object storage credentials with read access to the table path
  • For the catalog path: a catalog token if the catalog requires authentication

Notes​

  • The URL (metadata path) must point directly to the Iceberg table root directory (the folder containing the metadata).
  • Network connectivity from the Lakehousecat backend to the object storage endpoint — and, when used, the catalog endpoint — is required.
  • The metadata path and catalog path are independent; setting a Catalog URI does not remove the need for object storage credentials, since table data still lives in S3-compatible storage.
  • AWS Glue is not supported as a catalog type yet.

Next Steps​

After creating the datasource, go to the Operations tab to trigger semantic extraction.