Deployment Overview
Lakehousecat is deployed on Kubernetes using a Kubernetes Operator that manages the complete lifecycle of your data platform instances. This operator-based approach provides automated deployment, configuration management, scaling, and upgrades across any Kubernetes environment.
How Lakehousecat Deployment Works
The Lakehousecat deployment consists of two main components:
1. Lakehousecat Operator (Cluster-Wide)
The Operator is a Kubernetes controller that:
- Manages Lakehousecat instances across multiple namespaces
- Validates licenses via the Lakehousecat license service (
license.lakehousecat.com) - Synchronizes container images from AWS ECR (public and private) to your cluster
- Orchestrates 40+ microservices (Airflow, PostgreSQL, Valkey, ClickHouse, Superset, SeaweedFS, etc.)
- Handles upgrades, scaling, and configuration changes
Install once per cluster - The operator can manage multiple Lakehousecat instances.
2. Lakehousecat Instance (Namespace-Specific)
Each instance is defined as a Custom Resource (CR) that specifies:
- License credentials (handshake key)
- Version and architecture (amd64/arm64)
- Storage configuration (S3, SeaweedFS, etc.)
- Domain and TLS settings
- Resource allocations and scaling parameters
Deploy multiple instances - Each instance runs in its own namespace with isolated resources.
Operator Permissions
The Lakehousecat Operator follows the principle of least privilege. It does not require cluster-admin rights and does not use ClusterRole or ClusterRoleBinding resources. All RBAC is namespace-scoped.
The operator requires the following Kubernetes permissions:
| Resource Group | Resources | Access |
|---|---|---|
lhc.lakehousecat.com | lakehousecats, lakehousecats/status, lakehousecats/finalizers | Full lifecycle management of the operator's own CRD |
Core ("") | namespaces | get, list, watch, create — required to provision instance namespaces |
Core ("") | secrets, configmaps, services, persistentvolumeclaims, serviceaccounts, pods, pods/log, events | Manage instance resources within namespaces |
apps | deployments, statefulsets, daemonsets, replicasets | Deploy and manage application workloads |
batch | jobs, cronjobs | Orchestrate data ingestion workflows |
networking.k8s.io | ingresses, networkpolicies | Configure ingress and network policies |
rbac.authorization.k8s.io | roles, rolebindings | Namespace-scoped RBAC for instance service accounts |
autoscaling | horizontalpodautoscalers | Enable HPA-based autoscaling |
policy | poddisruptionbudgets | Ensure availability during cluster maintenance |
coordination.k8s.io | leases | Operator leader election |
discovery.k8s.io | endpointslices | Service discovery |
The operator uses Role and RoleBinding (namespace-scoped), not ClusterRole or ClusterRoleBinding. The only cluster-level resources accessed are namespaces and leases — required for namespace provisioning and leader election.
Prerequisites
Before deploying Lakehousecat, ensure you have:
Required Software
- Kubernetes: Version 1.29 or higher
- Helm: Version 3.8 or higher
- kubectl: Configured and connected to your cluster
- Handshake Key: Obtained from portal.lakehousecat.com after registration (free)
Cluster Resources (Minimum)
For Single-User Evaluation:
- 10 CPU cores (minimum)
- 32 GB RAM (minimum)
- 50 GB free disk space (SSD recommended)
- Supported OS: macOS, Linux
For Production Multi-User Deployment:
- 16+ CPU cores
- 64+ GB RAM
- 500+ GB persistent storage (scales with data volume)
- Persistent volume provisioner (for database storage)
Network Requirements
- Outbound HTTPS access to:
license.lakehousecat.com(license validation)- AWS ECR (public and private) — container image pull; handled automatically by the operator
- DNS resolution for internal Kubernetes service discovery
- Ingress controller (required for external access via domain)
Deployment Options
Choose the deployment option that best fits your needs:
Local Evaluation
Best for: Single-user evaluation, development, testing
Two options are available:
- TUI Installer — Recommended. A terminal application that automates the full setup: Minikube cluster, Helm repository, operator, and instance. No Kubernetes experience required.
- Manual Setup — For advanced users who need full control over the configuration.
Resource Requirements: 32 GB RAM (minimum), 10 CPU cores, 50 GB free disk, macOS or Linux
Kubernetes Cluster
Best for: Production deployments, multi-user environments, scalability
Deploy Lakehousecat on managed or self-hosted Kubernetes clusters:
- AWS EKS (Elastic Kubernetes Service)
- Google GKE (Google Kubernetes Engine)
- Azure AKS (Azure Kubernetes Service)
- On-Premise - Self-managed Kubernetes clusters
Resource Requirements: 16+ CPU cores, 64+ GB RAM, 500+ GB storage
Deployment Process Overview
Regardless of which option you choose, the deployment follows the same basic steps:
Step 1: Prepare Kubernetes Cluster
Set up your Kubernetes environment (Minikube, EKS, GKE, AKS, or on-premise).
Step 2: Install Lakehousecat Operator
Install the operator using Helm:
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update
helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait
Step 3: Configure Handshake Secret
Create a secret with your handshake key from the portal:
kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator
Step 4: Deploy Lakehousecat Instance
Create a Lakehousecat Custom Resource (CR):
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: my-instance
spec:
license:
handshakeSecretName: "lhc-handshake-secret"
customer:
namespace: "lhc-instance"
adminemail: "admin@company.com"
version: "0.0.40"
architecture: "amd64" # amd64 | arm64
storage:
bucketName: "my-data-bucket"
region: "us-east-1"
accessKeyID: "AKIA..."
secretAccessKey: "..."
Apply the configuration:
kubectl apply -f lakehousecat-instance.yaml
Step 5: Verify Deployment
Monitor the deployment:
# Check status of your specific instance (replace my-instance and lhc-instance with your values)
kubectl get lakehousecat my-instance -n lhc-instance
# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator
# Check pods in instance namespace
kubectl get pods -n lhc-instance
Verifying a Successful Deployment
Check all instances across all namespaces:
kubectl get lakehousecats -A
Phase: Ready means the Operator has reconciled all components and all pods are online and healthy. If the phase shows Pending, the Operator is still reconciling — check the CR description for details:
kubectl describe lakehousecat <instance-name> -n <namespace>
Check individual pod status:
kubectl -n <instance-namespace> get pods
All pods should show Running status with all containers ready. The full set of expected pods in a healthy deployment:
| Pod | Component |
|---|---|
airflow-api-server | Airflow API |
airflow-dag-processor | DAG processing |
airflow-pgbouncer | Airflow DB connection pool |
airflow-scheduler | Airflow scheduler |
airflow-statsd | Airflow metrics |
airflow-triggerer | Airflow triggerer |
airflow-worker | Airflow task worker |
analytics | Analytics service |
audio | Audio transcription service |
clickhouse-shard0-0 | ClickHouse analytics DB |
clickhouse-zookeeper-0 | ClickHouse coordination |
lhc | LHC Core backend |
lhc-registry | Internal registry |
llm | LLM integration service |
seaweedfs-s3 | Object storage |
postgresql-primary-0 | Primary PostgreSQL |
postgresql-read-0 | PostgreSQL read replica |
valkey-primary | Valkey primary |
valkey-replica | Valkey replica |
semantic | Semantic processing service |
superset | Superset web application |
superset-worker | Superset Celery worker |
ui | Lakehousecat web UI |
First start timing: On initial deployment, expect 30–60 minutes before all pods reach Running state, depending on your cluster's available resources and image pull speed from AWS ECR.
Step 6: Access Lakehousecat
Once deployed, access Lakehousecat via:
Port-forward (for testing):
kubectl port-forward svc/lhc -n lhc-instance 42021:42021
# Access at http://localhost:42021
Ingress (for production): Configure domain and TLS in the Lakehousecat CR (see cluster guides for details).
Subscription Tiers
Lakehousecat offers multiple subscription tiers to match your needs:
| Tier | Users | Data Sources | Use Case |
|---|---|---|---|
| FREE | 1 | 10 | Evaluation, development |
| STANDARD | Up to 25 | 50 | Small teams |
| PREMIUM | Up to 100 | Unlimited | Medium organizations |
| ENTERPRISE | 300+ | Unlimited | Large enterprises |
A handshake key is required for all tiers, including FREE. Register at portal.lakehousecat.com — no paid subscription is required to get started.
Multi-Cloud and On-Premise Support
The Lakehousecat Operator is designed to run on any Kubernetes cluster, regardless of where it's hosted:
- ✅ AWS EKS - Tested and supported
- ⚠️ Google GKE - Not yet tested by Lakehousecat
- ⚠️ Azure AKS - Not yet tested by Lakehousecat
- ✅ On-Premise - Self-managed Kubernetes clusters
- ✅ Minikube - Local development and evaluation
The operator abstracts cloud-specific differences, providing a consistent deployment experience across all environments.
What Gets Deployed
When you create a Lakehousecat instance, the operator automatically deploys:
Application Services
- UI - SvelteKit web application
- LHC Core - Backend API and business logic
- Analytics - Query execution and chart generation
- LLM - Large language model integration
- Semantic - Natural language processing
- Audio - Audio transcription services
Data Storage
- PostgreSQL - Primary relational database
- ClickHouse - Columnar analytics database
- Valkey - In-memory cache and messaging
- SeaweedFS - S3-compatible object storage
Workflow Orchestration
- Apache Airflow - ETL pipeline orchestration
- Scheduler, Workers, Webserver, Triggerer, DAG Processor
Business Intelligence
- Apache Superset - Data visualization and dashboards
- Web application, Celery workers
Optional Monitoring
- Prometheus - Metrics collection
- Grafana - Metrics visualization
All services are configured, networked, and managed automatically by the operator.
Upgrade and Maintenance
Upgrades and scaling are managed through the Operator after initial deployment.
- For upgrade instructions (backup recommendations, Helm commands, verification): see Upgrade
- For scaling configuration: see Scaling
Do not upgrade Airflow, ClickHouse, or Superset via Helm or any other means outside of the Lakehousecat Operator. Independent upgrades break compatibility and are not supported. Component version updates are delivered as part of Lakehousecat releases.
Next Steps
Choose your deployment path:
- TUI Installer - Fastest path: automated local evaluation setup
- Manual Setup - Advanced local evaluation with full configuration control
- Kubernetes Cluster - Deploy to AWS EKS, Google GKE, Azure AKS, or on-premise
Support
- Documentation: https://docs.lakehousecat.com
- Portal: https://portal.lakehousecat.com
- Support: portal.lakehousecat.com
- Community: Discord (coming soon)
- Operator README: Detailed operator documentation in the Helm chart