Skip to main content
Version: 0.0.41

Apache Drill

Connect an Apache Drill cluster to Lakehousecat using a connection URI. Drill is a schema-free SQL query engine for Hadoop, NoSQL systems, and cloud object storage. It can query JSON, CSV, Parquet, and other file formats directly without requiring a predefined schema.

Connection Fields​

FieldRequiredDescription
Connection URIYesFull connection string including host, port, and storage plugin. Example: drill+sadrill://host:8047/dfs
Use SSHNoEnable if the Drill Drillbit endpoint is behind an SSH bastion host. When enabled, SSH settings fields appear: Host, Port, Username, Password, Private Key, and Private Key Passphrase. See SSH Tunneling.

URI Format​

drill+sadrill://host:port/storage_plugin
drill+sadrill://host:port/storage_plugin/workspace

Example (local development):

drill+sadrill://localhost:8047/dfs

Prerequisites​

  • Admin or Builder role in Lakehousecat
  • A running Apache Drill cluster with the REST API enabled (port 8047 by default)
  • Network connectivity from the Lakehousecat backend to the Drill Drillbit (direct or via SSH)

Storage Plugins​

The path segment in the URI refers to a storage plugin configured in your Drill instance. Common storage plugins:

PluginDescription
dfsLocal or distributed file system (HDFS, S3, etc.)
cpClasspath file system (Drill's built-in sample data)
hiveHive Metastore integration
mongoMongoDB integration

Notes​

  • Drill does not require a schema to be defined before querying — it infers schema at query time from the underlying data.
  • For large storage plugins with many files, configure table filters in the Filters tab after creation to limit the scope of semantic extraction.
  • Authentication can be configured in Drill's drill-override.conf (Plain SASL, Kerberos, etc.).

Next Steps​

After creating the datasource, go to the Operations tab to trigger semantic extraction.