Kubernetes Cluster Deployment
Deploy Lakehousecat to production-grade Kubernetes clusters in the cloud or on-premise. This guide covers deployment to managed Kubernetes services and self-hosted clusters.
Overview
Kubernetes cluster deployment is recommended for:
- Production environments with high availability requirements
- Multi-user deployments (teams and organizations)
- Scalable workloads with dynamic resource needs
- Enterprise deployments with compliance and security requirements
Supported Kubernetes Platforms
Lakehousecat Operator runs on any standard Kubernetes 1.29+ cluster:
Managed Kubernetes Services
| Platform | Status | Guide |
|---|---|---|
| AWS EKS | ✅ Tested & Supported | → EKS Guide |
| Google GKE | ⚠️ Not yet tested | → GKE Guide |
| Azure AKS | ⚠️ Not yet tested | → AKS Guide |
Self-Managed Kubernetes
| Platform | Status | Guide |
|---|---|---|
| On-Premise | ✅ Compatible | → On-Premise Guide |
| Bare Metal | ✅ Compatible | Use On-Premise Guide |
| Private Cloud | ✅ Compatible | Use On-Premise Guide |
Prerequisites
Before deploying to a Kubernetes cluster, ensure you have:
Cluster Requirements
- Kubernetes: Version 1.29 or higher
- kubectl: Installed and configured for your cluster
- Helm: Version 3.8 or higher
- Persistent Storage: Dynamic volume provisioner (StorageClass)
- LoadBalancer or Ingress Controller: For external access
Resource Requirements
Minimum Production Configuration:
- 16 CPU cores
- 64 GB RAM
- 500 GB persistent storage
Recommended Production Configuration:
- 32+ CPU cores
- 128+ GB RAM
- 1+ TB persistent storage
- Auto-scaling enabled
Network Requirements
- Outbound HTTPS access to:
license.lakehousecat.com(license validation)- AWS ECR (public and private) — container image pull; handled automatically by the operator
- Ingress capability for external user access
- TLS certificates for HTTPS (Let's Encrypt recommended)
License
- Handshake key from portal.lakehousecat.com (required for all tiers, including FREE)
General Deployment Process
The deployment process is consistent across all Kubernetes platforms:
Step 1: Provision Kubernetes Cluster
Create a Kubernetes cluster on your chosen platform (EKS, GKE, AKS, or on-premise).
Key Configuration:
- Kubernetes version: 1.29 or higher
- Node pool: Appropriate instance types for your workload
- Storage: Dynamic volume provisioner enabled
- Networking: LoadBalancer or Ingress controller
See platform-specific guides for detailed instructions.
Step 2: Configure kubectl
Ensure kubectl is configured to connect to your cluster:
# Verify cluster access
kubectl cluster-info
kubectl get nodes
# Check available storage classes
kubectl get storageclass
Step 3: Install Lakehousecat Operator
Install the operator using Helm:
# Add Lakehousecat Helm repository (OCI registry)
helm pull oci://public.ecr.aws/lakehousecat/charts/lakehousecat-operator --version 0.0.42
# Or add traditional Helm repo
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update
# Install operator
helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait
Production Values (optional):
replicaCount: 1
operator:
logging:
level: info
format: json
leaderElection:
enabled: false
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2000m
memory: 2Gi
metrics:
serviceMonitor:
enabled: true
namespace: monitoring
interval: 30s
networkPolicy:
enabled: true
Install with custom values:
helm install lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
--create-namespace \
-f operator-values.yaml \
--wait
Verify operator deployment:
kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 4: Create Handshake Secret
Create a Kubernetes secret with your handshake key:
kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY_FROM_PORTAL \
-n lhc-operator
Verify secret creation:
kubectl get secret lhc-handshake-secret -n lhc-operator
Step 5: Create Required Secrets
All sensitive values — model API keys, the Mapbox token, external S3 credentials, OAuth client secrets — are read by the operator from pre-created Kubernetes secrets, never from the CR. Create the secrets you need before deploying the instance.
See Required Secrets for the full list, the exact keys per provider, and ready-made helper scripts. At minimum you created the handshake secret in Step 4; add the optional secrets for any feature you intend to use (models, Mapbox, external S3, SSO).
Missing optional secrets do not block the deployment — the affected feature is simply skipped and reported via a status condition. You can add secrets later and re-trigger a reconcile.
Step 5b: Configure Storage
Lakehousecat requires persistent storage for databases and object storage.
Option 1: Cloud Object Storage
Use cloud-native S3-compatible storage. Credentials come from the
lhc-storage-s3-secret (see Required Secrets),
and the CR references it:
spec:
objectStorage:
type: s3
region: "us-east-1"
bucketPrefix: "my-company-lakehousecat"
credentialsSecretName: "lhc-storage-s3-secret" # keys: access-key, secret-key
Option 2: In-Cluster SeaweedFS (Default)
Deploy SeaweedFS within the cluster (managed by the operator, no credentials needed):
spec:
objectStorage:
type: seaweedfs
volume:
persistence:
size: "500Gi"
storageClass: "gp3" # Use your cluster's storage class
Step 6: Deploy Lakehousecat Instance
Create a Lakehousecat Custom Resource with production configuration:
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"
customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"
version: "0.0.42"
architecture: "amd64"
# Storage configuration — external S3 (credentials via pre-created secret).
# Omit this block entirely to use the default in-cluster SeaweedFS.
objectStorage:
type: s3
region: "us-east-1"
bucketPrefix: "my-company-lakehousecat"
credentialsSecretName: "lhc-storage-s3-secret" # keys: access-key, secret-key
# Domain configuration with TLS
domain:
baseDomain: "lakehousecat.company.com"
tls:
enabled: true
letsEncrypt: true
email: "admin@company.com"
# Production resource allocations
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "2"
memory: "8Gi"
llm:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"
postgresql:
persistence:
size: "100Gi"
storageClass: "gp3"
resources:
requests:
cpu: "2"
memory: "8Gi"
clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "gp3"
resources:
requests:
cpu: "4"
memory: "32Gi"
valkey:
persistence:
size: "20Gi"
# Airflow configuration
airflow:
worker:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
resources:
requests:
cpu: "2"
memory: "8Gi"
# Monitoring — exposes /metrics endpoints on backend services
# for scraping by the customer's existing monitoring stack.
# The Operator does NOT deploy Prometheus or Grafana.
monitoring:
metricsEnabled: true
Apply the configuration:
kubectl apply -f lakehousecat-production.yaml
Step 7: Monitor Deployment
Watch the deployment progress:
# Check instance status
kubectl get lakehousecat -w
# Check all pods
kubectl get pods -n lakehousecat-prod -w
# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator
# Check specific service
kubectl describe lakehousecat production
Step 8: Configure DNS and Ingress
Once deployed, configure DNS to point to your cluster's LoadBalancer or Ingress:
Get LoadBalancer IP/Hostname:
kubectl get svc -n lakehousecat-prod lakehousecat-ui
Configure DNS: Create A or CNAME records:
lakehousecat.company.com → <LOADBALANCER_IP_OR_HOSTNAME>
Verify TLS:
curl -I https://lakehousecat.company.com
Step 9: Access Lakehousecat
Open your browser to your configured domain:
https://lakehousecat.company.com
Create your admin account and begin onboarding users.
High Availability Configuration
For mission-critical deployments, configure high availability:
Multi-Replica Services
spec:
analytics:
replicaCount: 3
lhc:
replicaCount: 3
llm:
replicaCount: 2
ui:
replicaCount: 3
Database Replication
spec:
postgresql:
replicaCount: 3
replication:
enabled: true
clickhouse:
shards: 3
replicaCount: 2 # 2 replicas per shard
Pod Disruption Budgets
The operator automatically creates Pod Disruption Budgets (PDBs) for critical services to ensure availability during cluster maintenance.
Security Best Practices
Network Policies
Enable network policies to isolate namespaces:
# Operator values
networkPolicy:
enabled: true
RBAC
The operator creates minimal RBAC permissions. Review and adjust as needed:
kubectl get rolebindings -n lakehousecat-prod
kubectl get clusterrolebindings | grep lakehousecat
Secrets Management
Use external secret management for sensitive data:
- AWS: AWS Secrets Manager or Parameter Store
- Google Cloud: Secret Manager
- Azure: Key Vault
Integrate using Kubernetes External Secrets Operator.
TLS/HTTPS
Always enable TLS for production:
spec:
domain:
tls:
enabled: true
letsEncrypt: true # Or provide your own certificates
Monitoring and Observability
The Operator does not deploy a monitoring stack. It exposes observability data through open interfaces — connect your own tooling.
Prometheus Metrics — /metrics Endpoint
Each backend service (lhc, llm, analytics, semantic, audio) exposes a /metrics endpoint in the Prometheus exposition format. The endpoint is unauthenticated by design (cluster-internal, follows Prometheus convention) and toggled via monitoring.metricsEnabled (default true).
For customers running prometheus-operator, a ServiceMonitor such as the following scrapes Lakehousecat services:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: lakehousecat-services
namespace: lakehousecat-prod
spec:
selector:
matchLabels:
app.kubernetes.io/part-of: lakehousecat
endpoints:
- port: http
path: /metrics
interval: 30s
Datadog, New Relic and similar agents support equivalent auto-discovery via Kubernetes service annotations — consult your agent's documentation.
Logging
Centralize logs using your preferred logging solution. All backend services produce structured application logs that are written to the SeaweedFS backend_logs/ bucket and to stdout/stderr. Compatible collectors include:
- AWS: CloudWatch Logs
- Google Cloud: Cloud Logging
- Azure: Azure Monitor
- ELK Stack: Elasticsearch, Logstash, Kibana
- Loki: Grafana Loki
- Datadog / New Relic / Splunk: via their respective Kubernetes agents
Backup and Disaster Recovery
Database Backups
Configure automated backups for PostgreSQL and ClickHouse:
spec:
postgresql:
backup:
enabled: true
schedule: "0 2 * * *" # Daily at 2 AM
retention: 30 # Keep 30 days
clickhouse:
backup:
enabled: true
schedule: "0 3 * * *"
retention: 30
Object Storage Backups
Enable versioning on your S3/Cloud Storage bucket for data protection.
Disaster Recovery Plan
- Regular backups of all databases
- GitOps for Lakehousecat CR configurations
- Multi-region deployment for critical workloads
- RTO/RPO definitions and testing
Scaling
Scale your deployment as needs grow. See the Scaling documentation for detailed guidance.
Quick scaling example:
# Scale analytics service
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"replicaCount":5}}}'
# Enable autoscaling
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"autoscaling":{"enabled":true,"minReplicas":3,"maxReplicas":10}}}}'
Upgrading
Operator Upgrade
helm upgrade lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
--version 0.0.42
Instance Upgrade
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"version":"0.0.42"}}'
The operator handles rolling upgrades with zero downtime.
Platform-Specific Guides
For detailed, platform-specific deployment instructions:
- AWS EKS - Amazon Elastic Kubernetes Service
- Google GKE - Google Kubernetes Engine
- Azure AKS - Azure Kubernetes Service
- On-Premise - Self-managed Kubernetes clusters
Troubleshooting
Pods Not Starting
# Check pod events
kubectl describe pod -n lakehousecat-prod <pod-name>
# Check operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator
# Check instance status
kubectl describe lakehousecat production
License Validation Failures
# Verify handshake secret
kubectl get secret lhc-handshake-secret -n lhc-operator -o yaml
# Check network connectivity to provider backend
kubectl run -it --rm debug --image=alpine --restart=Never -- \
wget -O- https://license.lakehousecat.com/health
Storage Issues
# Check PVCs
kubectl get pvc -n lakehousecat-prod
# Check storage class
kubectl get storageclass
# Describe PVC for events
kubectl describe pvc -n lakehousecat-prod <pvc-name>
Support
- Documentation: https://docs.lakehousecat.com
- Portal: https://portal.lakehousecat.com
- Support: portal.lakehousecat.com
- Community: Discord (coming soon)