Lakehousecat Architecture
Lakehousecat is built on a cloud-native, Kubernetes-based architecture. The system is organized into four layers: Frontend, Backend, Workflow & Visualization, and Storage.
Architecture Overview
How a Request Flows
The diagram below shows the path of a typical user question — from input in the browser to the final answer with a visualization.
Key Points
- LHC Service is the authority for identity and metadata — it does not process queries directly.
- LLM Service orchestrates the response: it decides whether a question requires data (→ Semantic + Analytics) or can be answered directly from context.
- Semantic Service provides the layer of meaning: without accurate descriptions and schema context, the LLM cannot generate correct SQL.
- Analytics Service executes the query and builds the visualization. Chart data is stored in ClickHouse; chart definitions are persisted in Superset.
- Persistence happens at multiple points: the session, the generated chart, and the query history are all stored so they can be revisited.
Frontend
UI
The Lakehousecat web application is built with Svelte. The UI adapts based on user roles, showing different workspaces and capabilities depending on access level. All communication with backend services happens over HTTP/HTTPS with JWT-based authentication.
Backend
LHC Service
The central backend service — the authority for identity and metadata:
- User authentication and session management
- Role-based access control (RBAC) enforcement
- Configuration and metadata management
- Coordination between services
LLM Service
Integrates large language models for conversational analytics:
- Natural language to SQL translation
- Provider model integration (OpenAI, Anthropic, Google, Azure, AWS Bedrock)
- Custom model execution
- Result interpretation and explanation
Semantic Service
Provides the layer of meaning — enabling the LLM to understand your data:
- Schema context and datasource metadata
- Descriptions for tables and columns
- Business terminology mapping
- Datasource and model semantic layer management
Analytics Service
Executes queries and builds visualizations:
- Executes SQL queries against ClickHouse
- Generates chart data (bar, line, pie, area, scatter, table, big number, pivot table)
- Handles data transformations and aggregations
- Builds and persists chart definitions in Superset
Audio Service
Enables voice input for chat sessions:
- Speech-to-text transcription (speech input only)
- Allows users to interact with Custom Models and create charts using natural language via microphone
- Integrates directly into the session prompt workflow
Workflow & Visualization
Apache Airflow
Workflow orchestration and task monitoring:
- Schedules and executes data ingestion jobs (incremental and full-load)
- Manages job dependencies and execution order
- Monitors job status
Job definitions are created and managed through the Lakehousecat Operations workspace.
Apache Superset
Visualization engine for charts and dashboards:
- Interactive chart rendering and dashboards
- Advanced filtering and drill-down
- Chart export and embedding in external applications
Storage
| Store | Type | Purpose |
|---|---|---|
| PostgreSQL | Relational | Metadata for data sources, custom models, user accounts, permissions, and job definitions |
| ClickHouse | Columnar | Analytical query execution — large dataset aggregations and time-series analysis |
| Valkey | In-memory | Session storage and query result caching |
| SeaweedFS | Object storage | System logs, backups, and data archiving |
User Roles and Access Control
Lakehousecat implements role-based access control (RBAC) with three roles:
| Role | Permissions |
|---|---|
| Administrator | Full system access — user management, provider model configuration, scaling, system settings |
| Builder | Create and configure data sources, custom models, prompts, and job definitions |
| User | Interact with custom models in sessions, generate charts and dashboards |
Access control is enforced by the LHC Service and propagated to all backend services via JWT tokens.
Export and download of data and results is not available in the current release.
Data Processing
Lakehousecat processes data in batch mode. Two patterns are supported:
- Incremental Load — loads only new or changed records since the last run; efficient for large datasets
- Full Load — complete dataset refresh; ensures consistency for smaller or fully replaced datasets
Deployment
Lakehousecat is deployed and managed via a Kubernetes Operator. The operator is the central control component of the platform:
- Deploys and manages all services as Kubernetes resources
- Manages instance configuration through a Custom Resource (CR)
- Handles upgrades and configuration changes
- The operator Helm chart is freely available at
https://charts.lakehousecat.com
Service Communication
- Frontend → Backend: REST over HTTP/HTTPS
- Backend → Storage: Direct database connections within the Kubernetes namespace
Production Deployments
Stateless services (UI, LHC, LLM, Semantic, Analytics, Audio, Superset, Airflow) support horizontal scaling via multiple replicas. Stateful services (PostgreSQL, ClickHouse) use persistent volumes and replication. Kubernetes handles pod restarts automatically on failure.
Security
- Authentication: JWT-based sessions and API keys
- Authorization: RBAC enforced at the LHC Service level
- Data at rest: Sensitive data is encrypted in the database. Persistent volume encryption is a Kubernetes cluster responsibility — managed by the cluster administrator, not by Lakehousecat.
- Data in transit: HTTPS with TLS certificates can be configured at the ingress level
Network isolation policies are managed by the Kubernetes cluster administrator, not by Lakehousecat.
Scaling
Each service scales independently:
- Horizontal (add replicas): UI, API, LHC, LLM, Semantic, Analytics, Audio, Superset, Airflow
- Vertical (CPU/memory): PostgreSQL, ClickHouse, SeaweedFS
- Autoscaling: HPA can be configured for stateless services
See the Scaling documentation for configuration details.
Monitoring (Optional)
A monitoring stack can be enabled through the Lakehousecat CR:
- Prometheus + Grafana: metrics collection and pre-configured dashboards
- Logs: all service logs are stored in SeaweedFS and are accessible for review
Monitoring is optional and can be disabled for smaller deployments to reduce resource consumption.
Service Overview
| Service | Layer | Required |
|---|---|---|
| UI | Frontend | Yes |
| LHC Service | Backend | Yes |
| LLM Service | Backend | Yes |
| Semantic Service | Backend | Yes |
| Analytics Service | Backend | Yes |
| Audio Service | Backend | Yes |
| Apache Airflow | Workflow | Yes |
| Apache Superset | Visualization | Yes |
| PostgreSQL | Storage | Yes |
| ClickHouse | Storage | Yes |
| Valkey | Storage | Yes |
| SeaweedFS | Storage | Yes |
| Prometheus + Grafana | Monitoring | Optional |