Apache Hive
Connect an Apache Hive metastore to Lakehousecat using a connection URI. Hive is a data warehouse system built on top of Hadoop that provides a SQL-like query language (HiveQL) for querying and managing large datasets stored in distributed storage.
Connection Fields
| Field | Required | Description |
|---|---|---|
| Connection URI | Yes | Full connection string including host, port, and database. Example: hive+pyhive://host:10000/database |
| Use SSH | No | Enable if the HiveServer2 endpoint is behind an SSH bastion host. When enabled, SSH settings fields appear: Host, Port, Username, Password, Private Key, and Private Key Passphrase. See SSH Tunneling. |
URI Format
hive+pyhive://host:port/database
hive+pyhive://user:password@host:port/database
Example (local development):
hive+pyhive://localhost:10000/lhc
Prerequisites
- Admin or Builder role in Lakehousecat
- A running HiveServer2 instance (port 10000 by default)
- A Hive user with SELECT access to the databases and tables you want to expose
- Network connectivity from the Lakehousecat backend to HiveServer2 (direct or via SSH)
Notes
- The connection targets HiveServer2 via the Thrift Binary Transport protocol on port 10000 (default).
- Authentication mode depends on your Hive configuration — the URI supports
NONE,LDAP, andKERBEROS-based authentication through additional URI parameters if required. - Use a read-only Hive user to follow the principle of least privilege.
- For large Hive metastores, configure schema and table filters in the Filters tab after creation to limit the scope of semantic extraction.
Next Steps
After creating the datasource, go to the Operations tab to trigger semantic extraction.