Apache Spark SQL
Connect an Apache Spark SQL endpoint to Lakehousecat using a connection URI. Apache Spark SQL exposes data via the Spark Thrift Server — a HiveServer2-compatible interface for SQL queries over Spark DataFrames and tables.
Connection Fields
| Field | Required | Description |
|---|---|---|
| Connection URI | Yes | Full connection string including host, port, and database name. Example: sparksql://host:10000/mydb |
| Use SSH | No | Enable if the Thrift Server is behind an SSH bastion host. When enabled, SSH settings fields appear. See SSH Tunneling. |
Prerequisites
- Admin or Builder role in Lakehousecat
- A running Spark Thrift Server (start with
./sbin/start-thriftserver.shin your Spark installation) - Network connectivity from the Lakehousecat backend to the Thrift Server (direct or via SSH)
- Spark Thrift Server default port: 10000
Notes
- The connection URI scheme
sparksql://is automatically translated tohive://internally, using the HiveServer2-compatible interface. - Tables must be registered in the Spark catalog (e.g., via
CREATE TABLE,spark.catalog.createTable, or Hive Metastore). - Supports standard SQL features: SELECT, WHERE, GROUP BY — no DML (no INSERT/UPDATE/DELETE on Spark tables from Lakehousecat).
- Incremental Load is supported; use a timestamp column (e.g.,
updated_at) as the incremental column. - Spark SQL tables use Parquet or ORC format by default — column types are fully supported.
Next Steps
After creating the datasource, go to the Operations tab to trigger semantic extraction.