Skip to main content
Version: 0.0.42

Apache Hive

Connect an Apache Hive metastore to Lakehousecat using a connection URI. Hive is a data warehouse system built on top of Hadoop that provides a SQL-like query language (HiveQL) for querying and managing large datasets stored in distributed storage.

Connection Fields​

FieldRequiredDescription
Connection URIYesFull connection string including host, port, and database. Example: hive+pyhive://host:10000/database
Use SSHNoEnable if the HiveServer2 endpoint is behind an SSH bastion host. When enabled, SSH settings fields appear: Host, Port, Username, Password, Private Key, and Private Key Passphrase. See SSH Tunneling.

URI Format​

hive+pyhive://host:port/database
hive+pyhive://user:password@host:port/database

Example (local development):

hive+pyhive://localhost:10000/lhc

Prerequisites​

  • Admin or Builder role in Lakehousecat
  • A running HiveServer2 instance (port 10000 by default)
  • A Hive user with SELECT access to the databases and tables you want to expose
  • Network connectivity from the Lakehousecat backend to HiveServer2 (direct or via SSH)

Notes​

  • The connection targets HiveServer2 via the Thrift Binary Transport protocol on port 10000 (default).
  • Authentication mode depends on your Hive configuration — the URI supports NONE, LDAP, and KERBEROS-based authentication through additional URI parameters if required.
  • Use a read-only Hive user to follow the principle of least privilege.
  • For large Hive metastores, configure schema and table filters in the Filters tab after creation to limit the scope of semantic extraction.

Next Steps​

After creating the datasource, go to the Operations tab to trigger semantic extraction.