Skip to main content
Version: 0.0.41

Lakehousecat Architecture

Lakehousecat is built on a cloud-native, Kubernetes-based architecture. The system is organized into four layers: Frontend, Backend, Workflow & Visualization, and Storage.

Architecture Overview​

How a Request Flows​

The diagram below shows the path of a typical user question — from input in the browser to the final answer with a visualization.

Key Points​

  • LHC Service is the authority for identity and metadata — it does not process queries directly.
  • LLM Service orchestrates the response: it decides whether a question requires data (→ Semantic + Analytics) or can be answered directly from context.
  • Semantic Service provides the layer of meaning: without accurate descriptions and schema context, the LLM cannot generate correct SQL.
  • Analytics Service executes the query and builds the visualization. Chart data is stored in ClickHouse; chart definitions are persisted in Superset.
  • Persistence happens at multiple points: the session, the generated chart, and the query history are all stored so they can be revisited.

Frontend​

UI​

The Lakehousecat web application is built with Svelte. The UI adapts based on user roles, showing different workspaces and capabilities depending on access level. All communication with backend services happens over HTTP/HTTPS with JWT-based authentication.

Backend​

LHC Service​

The central backend service — the authority for identity and metadata:

  • User authentication and session management
  • Role-based access control (RBAC) enforcement
  • Configuration and metadata management
  • Coordination between services

LLM Service​

Integrates large language models for conversational analytics:

  • Natural language to SQL translation
  • Provider model integration (OpenAI, Anthropic, Google, Azure, AWS Bedrock)
  • Custom model execution
  • Result interpretation and explanation

Semantic Service​

Provides the layer of meaning — enabling the LLM to understand your data:

  • Schema context and datasource metadata
  • Descriptions for tables and columns
  • Business terminology mapping
  • Datasource and model semantic layer management

Analytics Service​

Executes queries and builds visualizations:

  • Executes SQL queries against ClickHouse
  • Generates chart data (bar, line, pie, area, scatter, table, big number, pivot table)
  • Handles data transformations and aggregations
  • Builds and persists chart definitions in Superset

Audio Service​

Enables voice input for chat sessions:

  • Speech-to-text transcription (speech input only)
  • Allows users to interact with Custom Models and create charts using natural language via microphone
  • Integrates directly into the session prompt workflow

Workflow & Visualization​

Apache Airflow​

Workflow orchestration and task monitoring:

  • Schedules and executes data ingestion jobs (incremental and full-load)
  • Manages job dependencies and execution order
  • Monitors job status

Job definitions are created and managed through the Lakehousecat Operations workspace.

Apache Superset​

Visualization engine for charts and dashboards:

  • Interactive chart rendering and dashboards
  • Advanced filtering and drill-down
  • Chart export and embedding in external applications

Storage​

StoreTypePurpose
PostgreSQLRelationalMetadata for data sources, custom models, user accounts, permissions, and job definitions
ClickHouseColumnarAnalytical query execution — large dataset aggregations and time-series analysis
ValkeyIn-memorySession storage and query result caching
SeaweedFSObject storageSystem logs, backups, and data archiving

User Roles and Access Control​

Lakehousecat implements role-based access control (RBAC) with three roles:

RolePermissions
AdministratorFull system access — user management, provider model configuration, scaling, system settings
BuilderCreate and configure data sources, custom models, prompts, and job definitions
UserInteract with custom models in sessions, generate charts and dashboards

Access control is enforced by the LHC Service and propagated to all backend services via JWT tokens.

note

Export and download of data and results is not available in the current release.

Data Processing​

Lakehousecat processes data in batch mode. Two patterns are supported:

  • Incremental Load — loads only new or changed records since the last run; efficient for large datasets
  • Full Load — complete dataset refresh; ensures consistency for smaller or fully replaced datasets

Deployment​

Lakehousecat is deployed and managed via a Kubernetes Operator. The operator is the central control component of the platform:

  • Deploys and manages all services as Kubernetes resources
  • Manages instance configuration through a Custom Resource (CR)
  • Handles upgrades and configuration changes
  • The operator Helm chart is freely available at https://charts.lakehousecat.com

Service Communication​

  • Frontend → Backend: REST over HTTP/HTTPS
  • Backend → Storage: Direct database connections within the Kubernetes namespace

Production Deployments​

Stateless services (UI, LHC, LLM, Semantic, Analytics, Audio, Superset, Airflow) support horizontal scaling via multiple replicas. Stateful services (PostgreSQL, ClickHouse) use persistent volumes and replication. Kubernetes handles pod restarts automatically on failure.

Security​

  • Authentication: JWT-based sessions and API keys
  • Authorization: RBAC enforced at the LHC Service level
  • Data at rest: Sensitive data is encrypted in the database. Persistent volume encryption is a Kubernetes cluster responsibility — managed by the cluster administrator, not by Lakehousecat.
  • Data in transit: HTTPS with TLS certificates can be configured at the ingress level
Network Policies

Network isolation policies are managed by the Kubernetes cluster administrator, not by Lakehousecat.

Scaling​

Each service scales independently:

  • Horizontal (add replicas): UI, API, LHC, LLM, Semantic, Analytics, Audio, Superset, Airflow
  • Vertical (CPU/memory): PostgreSQL, ClickHouse, SeaweedFS
  • Autoscaling: HPA can be configured for stateless services

See the Scaling documentation for configuration details.

Monitoring (Optional)​

A monitoring stack can be enabled through the Lakehousecat CR:

  • Prometheus + Grafana: metrics collection and pre-configured dashboards
  • Logs: all service logs are stored in SeaweedFS and are accessible for review

Monitoring is optional and can be disabled for smaller deployments to reduce resource consumption.

Service Overview​

ServiceLayerRequired
UIFrontendYes
LHC ServiceBackendYes
LLM ServiceBackendYes
Semantic ServiceBackendYes
Analytics ServiceBackendYes
Audio ServiceBackendYes
Apache AirflowWorkflowYes
Apache SupersetVisualizationYes
PostgreSQLStorageYes
ClickHouseStorageYes
ValkeyStorageYes
SeaweedFSStorageYes
Prometheus + GrafanaMonitoringOptional