Apache Drill
Connect an Apache Drill cluster to Lakehousecat using a connection URI. Drill is a schema-free SQL query engine for Hadoop, NoSQL systems, and cloud object storage. It can query JSON, CSV, Parquet, and other file formats directly without requiring a predefined schema.
Connection Fields
| Field | Required | Description |
|---|---|---|
| Connection URI | Yes | Full connection string including host, port, and storage plugin. Example: drill+sadrill://host:8047/dfs |
| Use SSH | No | Enable if the Drill Drillbit endpoint is behind an SSH bastion host. When enabled, SSH settings fields appear: Host, Port, Username, Password, Private Key, and Private Key Passphrase. See SSH Tunneling. |
URI Format
drill+sadrill://host:port/storage_plugin
drill+sadrill://host:port/storage_plugin/workspace
Example (local development):
drill+sadrill://localhost:8047/dfs
Prerequisites
- Admin or Builder role in Lakehousecat
- A running Apache Drill cluster with the REST API enabled (port 8047 by default)
- Network connectivity from the Lakehousecat backend to the Drill Drillbit (direct or via SSH)
Storage Plugins
The path segment in the URI refers to a storage plugin configured in your Drill instance. Common storage plugins:
| Plugin | Description |
|---|---|
dfs | Local or distributed file system (HDFS, S3, etc.) |
cp | Classpath file system (Drill's built-in sample data) |
hive | Hive Metastore integration |
mongo | MongoDB integration |
Notes
- Drill does not require a schema to be defined before querying — it infers schema at query time from the underlying data.
- For large storage plugins with many files, configure table filters in the Filters tab after creation to limit the scope of semantic extraction.
- Authentication can be configured in Drill's
drill-override.conf(Plain SASL, Kerberos, etc.).
Next Steps
After creating the datasource, go to the Operations tab to trigger semantic extraction.