Skip to main content
Version: Next

Apache Spark SQL

Connect an Apache Spark SQL endpoint to Lakehousecat using a connection URI. Apache Spark SQL exposes data via the Spark Thrift Server — a HiveServer2-compatible interface for SQL queries over Spark DataFrames and tables.

Connection Fields​

FieldRequiredDescription
Connection URIYesFull connection string including host, port, and database name. Example: sparksql://host:10000/mydb
Use SSHNoEnable if the Thrift Server is behind an SSH bastion host. When enabled, SSH settings fields appear. See SSH Tunneling.

Prerequisites​

  • Admin or Builder role in Lakehousecat
  • A running Spark Thrift Server (start with ./sbin/start-thriftserver.sh in your Spark installation)
  • Network connectivity from the Lakehousecat backend to the Thrift Server (direct or via SSH)
  • Spark Thrift Server default port: 10000

Notes​

  • The connection URI scheme sparksql:// is automatically translated to hive:// internally, using the HiveServer2-compatible interface.
  • Tables must be registered in the Spark catalog (e.g., via CREATE TABLE, spark.catalog.createTable, or Hive Metastore).
  • Supports standard SQL features: SELECT, WHERE, GROUP BY — no DML (no INSERT/UPDATE/DELETE on Spark tables from Lakehousecat).
  • Incremental Load is supported; use a timestamp column (e.g., updated_at) as the incremental column.
  • Spark SQL tables use Parquet or ORC format by default — column types are fully supported.

Next Steps​

After creating the datasource, go to the Operations tab to trigger semantic extraction.