Lakehousecat Architecture
Lakehousecat is built on a cloud-native, Kubernetes-based architecture. Four interaction surfaces — UI, REST API, CLI, and Apache Superset — connect to a set of specialized backend services, supported by a workflow engine and a multi-store persistence layer.
Architecture Overview
How a Request Flows
The steps below trace a typical user question — from input in the browser to the final answer with a visualization.
Key Points
- LHC Service is the authority for identity and metadata — it does not process queries directly.
- LLM Service orchestrates the response: it decides whether a question requires data (→ Semantic + Analytics) or can be answered directly from context.
- Semantic Service provides the layer of meaning: without accurate descriptions and schema context, the LLM cannot generate correct SQL.
- Analytics Service executes the query and builds the visualization. Chart data is stored in ClickHouse; chart definitions are persisted in Superset.
- Agent Service handles autonomous multi-step tasks — it can trigger datasource operations, chart generation, and model management on behalf of users.
- Persistence happens at multiple points: the session, the generated chart, and the query history are all stored so they can be revisited.
Interaction Surfaces
Lakehousecat exposes four interaction surfaces. All four share the same backend and the same authentication model.
UI
The Lakehousecat web application is built with Svelte. The UI adapts based on user roles, showing different workspaces and capabilities depending on access level. All communication with backend services happens over HTTP/HTTPS with JWT-based authentication.
REST API
All capabilities available in the UI are also accessible through the REST API. The API is the foundation for the CLI and for custom integrations. API access is authenticated with API keys generated in the UI under Account Settings → Security → API Keys. See the API Reference for available endpoints organized by use case.
CLI
The lhc command-line interface provides a terminal-first path to all management operations: datasources, custom models, users, sessions, analytics, and jobs. It wraps the REST API and is available for macOS, Linux, and Windows. See the CLI Reference for installation and command reference.
Apache Superset
Charts generated through Lakehousecat sessions are backed by Superset chart definitions. Users can open any chart directly in Superset to apply advanced customizations — additional filters, drill-downs, cross-filters, and layout changes that go beyond the session interface. Superset is bundled with every Lakehousecat instance and accessible from the Analytics workspace.
Backend
LHC Service
The central backend service — the authority for identity and metadata:
- User authentication and session management
- Role-based access control (RBAC) enforcement
- Configuration and metadata management
- Coordination between services
LLM Service
Integrates large language models for conversational analytics:
- Natural language to SQL translation
- Provider model integration (OpenAI, Anthropic, Google, Azure, AWS Bedrock)
- Custom model execution
- Result interpretation and explanation
Semantic Service
Provides the layer of meaning — enabling the LLM to understand your data:
- Schema context and datasource metadata
- Descriptions for tables and columns
- Business terminology mapping
- Datasource and model semantic layer management
Analytics Service
Executes queries and builds visualizations:
- Executes SQL queries against ClickHouse
- Generates chart data (bar, line, pie, area, scatter, table, big number, pivot table)
- Handles data transformations and aggregations
- Builds and persists chart definitions in Superset
Audio Service
Enables voice input for chat sessions:
- Speech-to-text transcription (speech input only)
- Allows users to interact with Custom Models and create charts using natural language via microphone
- Integrates directly into the session prompt workflow
Agent Service
Runs autonomous, multi-step operations on behalf of users:
- Executes sequences of Lakehousecat actions based on a goal (e.g., load a datasource, build a model, generate a chart)
- Orchestrates multiple backend services within a single agent task
- Accessible through the AI Agent interface in the UI
See the AI Agent guide for user-facing capabilities.
Workflow
Apache Airflow
Workflow orchestration and task monitoring:
- Schedules and executes data ingestion jobs (incremental and full-load)
- Manages job dependencies and execution order
- Monitors job status
Job definitions are created and managed through the Lakehousecat Operations workspace.
Storage
| Store | Type | Purpose |
|---|---|---|
| PostgreSQL | Relational | Metadata for data sources, custom models, user accounts, permissions, and job definitions |
| ClickHouse | Columnar | Analytical query execution — large dataset aggregations and time-series analysis |
| Valkey | In-memory | Session storage and query result caching |
| SeaweedFS | Object storage | System logs, backups, and data archiving |
User Roles and Access Control
Lakehousecat implements role-based access control (RBAC) with three roles:
| Role | Permissions |
|---|---|
| Administrator | Full system access — user management, provider model configuration, scaling, system settings |
| Builder | Create and configure data sources, custom models, prompts, and job definitions |
| User | Interact with custom models in sessions, generate charts and dashboards |
Access control is enforced by the LHC Service and propagated to all backend services via JWT tokens.
Bulk export of raw query results from the interface is not available in the current release. Resource definitions — data sources, custom models, charts, and dashboards — can be exported and re-imported as portable files using the lhc command-line tool.
Data Processing
Lakehousecat processes data in batch mode. Two patterns are supported:
- Incremental Load — loads only new or changed records since the last run; efficient for large datasets
- Full Load — complete dataset refresh; ensures consistency for smaller or fully replaced datasets
Deployment
Lakehousecat is deployed and managed via a Kubernetes Operator. The operator is the central control component of the platform:
- Deploys and manages all services as Kubernetes resources
- Manages instance configuration through a Custom Resource (CR)
- Handles upgrades and configuration changes
- The operator Helm chart is freely available at
https://charts.lakehousecat.com
Service Communication
- Interaction → Backend: REST over HTTP/HTTPS
- Backend → Storage: Direct database connections within the Kubernetes namespace
Production Deployments
Stateless services (UI, LHC, LLM, Semantic, Analytics, Audio, Agent, Superset, Airflow) support horizontal scaling via multiple replicas. Stateful services (PostgreSQL, ClickHouse) use persistent volumes and replication. Kubernetes handles pod restarts automatically on failure.
Security
- Authentication: JWT-based sessions and API keys
- Authorization: RBAC enforced at the LHC Service level
- Data at rest: Sensitive data is encrypted in the database. Persistent volume encryption is a Kubernetes cluster responsibility — managed by the cluster administrator, not by Lakehousecat.
- Data in transit: HTTPS with TLS certificates can be configured at the ingress level
The Lakehousecat Operator automatically deploys Kubernetes NetworkPolicy resources within the application namespace, restricting service-to-service traffic to the specific connections each service needs (e.g. only the LHC backend may reach PostgreSQL, only services calling AI providers may reach external HTTPS). Whether these policies are actually enforced depends on your cluster's CNI (network plugin) — the Operator detects this and surfaces a warning if the cluster cannot enforce them. Network policies for the cluster as a whole (outside the application namespace) remain the responsibility of the Kubernetes cluster administrator.
Scaling
Each service scales independently:
- Horizontal (add replicas): UI, LHC, LLM, Semantic, Analytics, Audio, Agent, Superset, Airflow
- Vertical (CPU/memory): PostgreSQL, ClickHouse, SeaweedFS
See the Scaling documentation for configuration details.
Observability (Bring Your Own Monitoring)
The Operator does not deploy a monitoring stack. Lakehousecat exposes observability data through open interfaces, intended for consumption by the customer's existing tools:
- Application metrics: each backend service exposes a
/metricsendpoint in Prometheus exposition format — scrape it with your Prometheus, Datadog Agent, Grafana Mimir, AWS Managed Prometheus, etc. - Application logs: all service logs are stored as structured records in SeaweedFS (
backend_logs/) and can be forwarded with any S3-compatible log collector. - Cluster metrics: pod/node-level data (CPU, memory, restarts) come from standard Kubernetes tooling (kube-state-metrics, cAdvisor) that the customer already operates at cluster scope.
Service Overview
| Service | Layer | Required |
|---|---|---|
| UI | Interaction | Yes |
| Apache Superset | Interaction | Yes |
| LHC Service | Backend | Yes |
| LLM Service | Backend | Yes |
| Semantic Service | Backend | Yes |
| Analytics Service | Backend | Yes |
| Audio Service | Backend | Yes |
| Agent Service | Backend | Yes |
| Apache Airflow | Workflow | Yes |
| PostgreSQL | Storage | Yes |
| ClickHouse | Storage | Yes |
| Valkey | Storage | Yes |
| SeaweedFS | Storage | Yes |