Skip to main content
Version: 0.0.42

Lakehousecat Architecture

Lakehousecat is built on a cloud-native, Kubernetes-based architecture. Four interaction surfaces — UI, REST API, CLI, and Apache Superset — connect to a set of specialized backend services, supported by a workflow engine and a multi-store persistence layer.

Architecture Overview​

Interaction
UIREST APICLIApache Superset
Backend
LHC ServiceLLM ServiceSemantic ServiceAnalytics ServiceAudio ServiceAgent Service
Workflow
Apache Airflow
Storage
PostgreSQLClickHouseValkeySeaweedFS

How a Request Flows​

The steps below trace a typical user question — from input in the browser to the final answer with a visualization.

1
User sends a question
UI
Typed or spoken in a session. The UI packages the question with session context and forwards it to the backend.
2
Authentication & session context
LHC Service
The LHC Service validates the JWT and returns the user profile, provider model configuration, and RBAC permissions for the session.
3
LLM receives the request
LLM Service
The LLM Service determines whether the question requires a data query. If yes, it requests schema context from the Semantic Service before generating SQL.
4
Semantic context loaded
Semantic ServicePostgreSQL
The Semantic Service loads table schemas, column descriptions, and business terminology from PostgreSQL and returns them to the LLM as grounding context.
5
SQL generated and executed
LLM ServiceAnalytics ServiceClickHouse
The LLM generates a SQL query and hands it to the Analytics Service, which executes it against ClickHouse and returns the result set.
6
Chart built in Superset
Analytics ServiceApache Superset
The Analytics Service uses the result set to build a chart definition in Superset. The chart reference and data summary are returned to the LLM.
7
Answer and chart delivered
LLM ServiceUI
The LLM composes the final answer text with the chart reference and streams it to the UI. The user sees the explanation and the interactive chart side by side.

Key Points​

  • LHC Service is the authority for identity and metadata — it does not process queries directly.
  • LLM Service orchestrates the response: it decides whether a question requires data (→ Semantic + Analytics) or can be answered directly from context.
  • Semantic Service provides the layer of meaning: without accurate descriptions and schema context, the LLM cannot generate correct SQL.
  • Analytics Service executes the query and builds the visualization. Chart data is stored in ClickHouse; chart definitions are persisted in Superset.
  • Agent Service handles autonomous multi-step tasks — it can trigger datasource operations, chart generation, and model management on behalf of users.
  • Persistence happens at multiple points: the session, the generated chart, and the query history are all stored so they can be revisited.

Interaction Surfaces​

Lakehousecat exposes four interaction surfaces. All four share the same backend and the same authentication model.

UI​

The Lakehousecat web application is built with Svelte. The UI adapts based on user roles, showing different workspaces and capabilities depending on access level. All communication with backend services happens over HTTP/HTTPS with JWT-based authentication.

REST API​

All capabilities available in the UI are also accessible through the REST API. The API is the foundation for the CLI and for custom integrations. API access is authenticated with API keys generated in the UI under Account Settings → Security → API Keys. See the API Reference for available endpoints organized by use case.

CLI​

The lhc command-line interface provides a terminal-first path to all management operations: datasources, custom models, users, sessions, analytics, and jobs. It wraps the REST API and is available for macOS, Linux, and Windows. See the CLI Reference for installation and command reference.

Apache Superset​

Charts generated through Lakehousecat sessions are backed by Superset chart definitions. Users can open any chart directly in Superset to apply advanced customizations — additional filters, drill-downs, cross-filters, and layout changes that go beyond the session interface. Superset is bundled with every Lakehousecat instance and accessible from the Analytics workspace.

Backend​

LHC Service​

The central backend service — the authority for identity and metadata:

  • User authentication and session management
  • Role-based access control (RBAC) enforcement
  • Configuration and metadata management
  • Coordination between services

LLM Service​

Integrates large language models for conversational analytics:

  • Natural language to SQL translation
  • Provider model integration (OpenAI, Anthropic, Google, Azure, AWS Bedrock)
  • Custom model execution
  • Result interpretation and explanation

Semantic Service​

Provides the layer of meaning — enabling the LLM to understand your data:

  • Schema context and datasource metadata
  • Descriptions for tables and columns
  • Business terminology mapping
  • Datasource and model semantic layer management

Analytics Service​

Executes queries and builds visualizations:

  • Executes SQL queries against ClickHouse
  • Generates chart data (bar, line, pie, area, scatter, table, big number, pivot table)
  • Handles data transformations and aggregations
  • Builds and persists chart definitions in Superset

Audio Service​

Enables voice input for chat sessions:

  • Speech-to-text transcription (speech input only)
  • Allows users to interact with Custom Models and create charts using natural language via microphone
  • Integrates directly into the session prompt workflow

Agent Service​

Runs autonomous, multi-step operations on behalf of users:

  • Executes sequences of Lakehousecat actions based on a goal (e.g., load a datasource, build a model, generate a chart)
  • Orchestrates multiple backend services within a single agent task
  • Accessible through the AI Agent interface in the UI

See the AI Agent guide for user-facing capabilities.

Workflow​

Apache Airflow​

Workflow orchestration and task monitoring:

  • Schedules and executes data ingestion jobs (incremental and full-load)
  • Manages job dependencies and execution order
  • Monitors job status

Job definitions are created and managed through the Lakehousecat Operations workspace.

Storage​

StoreTypePurpose
PostgreSQLRelationalMetadata for data sources, custom models, user accounts, permissions, and job definitions
ClickHouseColumnarAnalytical query execution — large dataset aggregations and time-series analysis
ValkeyIn-memorySession storage and query result caching
SeaweedFSObject storageSystem logs, backups, and data archiving

User Roles and Access Control​

Lakehousecat implements role-based access control (RBAC) with three roles:

RolePermissions
AdministratorFull system access — user management, provider model configuration, scaling, system settings
BuilderCreate and configure data sources, custom models, prompts, and job definitions
UserInteract with custom models in sessions, generate charts and dashboards

Access control is enforced by the LHC Service and propagated to all backend services via JWT tokens.

note

Bulk export of raw query results from the interface is not available in the current release. Resource definitions — data sources, custom models, charts, and dashboards — can be exported and re-imported as portable files using the lhc command-line tool.

Data Processing​

Lakehousecat processes data in batch mode. Two patterns are supported:

  • Incremental Load — loads only new or changed records since the last run; efficient for large datasets
  • Full Load — complete dataset refresh; ensures consistency for smaller or fully replaced datasets

Deployment​

Lakehousecat is deployed and managed via a Kubernetes Operator. The operator is the central control component of the platform:

  • Deploys and manages all services as Kubernetes resources
  • Manages instance configuration through a Custom Resource (CR)
  • Handles upgrades and configuration changes
  • The operator Helm chart is freely available at https://charts.lakehousecat.com

Service Communication​

  • Interaction → Backend: REST over HTTP/HTTPS
  • Backend → Storage: Direct database connections within the Kubernetes namespace

Production Deployments​

Stateless services (UI, LHC, LLM, Semantic, Analytics, Audio, Agent, Superset, Airflow) support horizontal scaling via multiple replicas. Stateful services (PostgreSQL, ClickHouse) use persistent volumes and replication. Kubernetes handles pod restarts automatically on failure.

Security​

  • Authentication: JWT-based sessions and API keys
  • Authorization: RBAC enforced at the LHC Service level
  • Data at rest: Sensitive data is encrypted in the database. Persistent volume encryption is a Kubernetes cluster responsibility — managed by the cluster administrator, not by Lakehousecat.
  • Data in transit: HTTPS with TLS certificates can be configured at the ingress level
Network Policies

The Lakehousecat Operator automatically deploys Kubernetes NetworkPolicy resources within the application namespace, restricting service-to-service traffic to the specific connections each service needs (e.g. only the LHC backend may reach PostgreSQL, only services calling AI providers may reach external HTTPS). Whether these policies are actually enforced depends on your cluster's CNI (network plugin) — the Operator detects this and surfaces a warning if the cluster cannot enforce them. Network policies for the cluster as a whole (outside the application namespace) remain the responsibility of the Kubernetes cluster administrator.

Scaling​

Each service scales independently:

  • Horizontal (add replicas): UI, LHC, LLM, Semantic, Analytics, Audio, Agent, Superset, Airflow
  • Vertical (CPU/memory): PostgreSQL, ClickHouse, SeaweedFS

See the Scaling documentation for configuration details.

Observability (Bring Your Own Monitoring)​

The Operator does not deploy a monitoring stack. Lakehousecat exposes observability data through open interfaces, intended for consumption by the customer's existing tools:

  • Application metrics: each backend service exposes a /metrics endpoint in Prometheus exposition format — scrape it with your Prometheus, Datadog Agent, Grafana Mimir, AWS Managed Prometheus, etc.
  • Application logs: all service logs are stored as structured records in SeaweedFS (backend_logs/) and can be forwarded with any S3-compatible log collector.
  • Cluster metrics: pod/node-level data (CPU, memory, restarts) come from standard Kubernetes tooling (kube-state-metrics, cAdvisor) that the customer already operates at cluster scope.

Service Overview​

ServiceLayerRequired
UIInteractionYes
Apache SupersetInteractionYes
LHC ServiceBackendYes
LLM ServiceBackendYes
Semantic ServiceBackendYes
Analytics ServiceBackendYes
Audio ServiceBackendYes
Agent ServiceBackendYes
Apache AirflowWorkflowYes
PostgreSQLStorageYes
ClickHouseStorageYes
ValkeyStorageYes
SeaweedFSStorageYes