Kubernetes Cluster Deployment
Deploy Lakehousecat to production-grade Kubernetes clusters in the cloud or on-premise. This guide covers deployment to managed Kubernetes services and self-hosted clusters.
Overview
Kubernetes cluster deployment is recommended for:
- Production environments with high availability requirements
- Multi-user deployments (teams and organizations)
- Scalable workloads with dynamic resource needs
- Enterprise deployments with compliance and security requirements
Supported Kubernetes Platforms
Lakehousecat Operator runs on any standard Kubernetes 1.29+ cluster:
Managed Kubernetes Services
| Platform | Status | Guide |
|---|---|---|
| AWS EKS | ✅ Tested & Supported | → EKS Guide |
| Google GKE | ⚠️ Not yet tested | → GKE Guide |
| Azure AKS | ⚠️ Not yet tested | → AKS Guide |
Self-Managed Kubernetes
| Platform | Status | Guide |
|---|---|---|
| On-Premise | ✅ Compatible | → On-Premise Guide |
| Bare Metal | ✅ Compatible | Use On-Premise Guide |
| Private Cloud | ✅ Compatible | Use On-Premise Guide |
Prerequisites
Before deploying to a Kubernetes cluster, ensure you have:
Cluster Requirements
- Kubernetes: Version 1.29 or higher
- kubectl: Installed and configured for your cluster
- Helm: Version 3.8 or higher
- Persistent Storage: Dynamic volume provisioner (StorageClass)
- LoadBalancer or Ingress Controller: For external access
Resource Requirements
Minimum Production Configuration:
- 16 CPU cores
- 64 GB RAM
- 500 GB persistent storage
Recommended Production Configuration:
- 32+ CPU cores
- 128+ GB RAM
- 1+ TB persistent storage
- Auto-scaling enabled
Network Requirements
- Outbound HTTPS access to:
license.lakehousecat.com(license validation)- AWS ECR (public and private) — container image pull; handled automatically by the operator
- Ingress capability for external user access
- TLS certificates for HTTPS (Let's Encrypt recommended)
License
- Handshake key from portal.lakehousecat.com (required for all tiers, including FREE)
General Deployment Process
The deployment process is consistent across all Kubernetes platforms:
Step 1: Provision Kubernetes Cluster
Create a Kubernetes cluster on your chosen platform (EKS, GKE, AKS, or on-premise).
Key Configuration:
- Kubernetes version: 1.29 or higher
- Node pool: Appropriate instance types for your workload
- Storage: Dynamic volume provisioner enabled
- Networking: LoadBalancer or Ingress controller
See platform-specific guides for detailed instructions.
Step 2: Configure kubectl
Ensure kubectl is configured to connect to your cluster:
# Verify cluster access
kubectl cluster-info
kubectl get nodes
# Check available storage classes
kubectl get storageclass
Step 3: Install Lakehousecat Operator
Install the operator using Helm:
# Add Lakehousecat Helm repository (OCI registry)
helm pull oci://public.ecr.aws/lakehousecat/charts/lakehousecat-operator --version 0.0.40
# Or add traditional Helm repo
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update
# Install operator
helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait
Production Values (optional):
replicaCount: 1
operator:
logging:
level: info
format: json
leaderElection:
enabled: false
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2000m
memory: 2Gi
metrics:
serviceMonitor:
enabled: true
namespace: monitoring
interval: 30s
networkPolicy:
enabled: true
Install with custom values:
helm install lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
--create-namespace \
-f operator-values.yaml \
--wait
Verify operator deployment:
kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 4: Create Handshake Secret
Create a Kubernetes secret with your handshake key:
kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY_FROM_PORTAL \
-n lhc-operator
Verify secret creation:
kubectl get secret lhc-handshake-secret -n lhc-operator
Step 5: Configure Storage
Lakehousecat requires persistent storage for databases and object storage.
Option 1: Cloud Object Storage (Recommended)
Use cloud-native object storage:
- AWS: S3
- Google Cloud: Cloud Storage
- Azure: Blob Storage
Create IAM credentials with appropriate permissions and configure in the Lakehousecat CR.
Option 2: In-Cluster SeaweedFS
Deploy SeaweedFS within the cluster (managed by operator):
spec:
seaweedfs:
enabled: true
persistence:
size: "500Gi"
storageClass: "gp3" # Use your cluster's storage class
Step 6: Deploy Lakehousecat Instance
Create a Lakehousecat Custom Resource with production configuration:
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"
customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"
version: "0.0.40"
architecture: "amd64"
# Storage configuration (S3 example)
storage:
bucketName: "my-company-lakehousecat-data"
region: "us-east-1"
accessKeyID: "AKIA..."
secretAccessKey: "..."
# Domain configuration with TLS
domain:
baseDomain: "lakehousecat.company.com"
tls:
enabled: true
letsEncrypt: true
email: "admin@company.com"
# Production resource allocations
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "2"
memory: "8Gi"
llm:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"
postgresql:
persistence:
size: "100Gi"
storageClass: "gp3"
resources:
requests:
cpu: "2"
memory: "8Gi"
clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "gp3"
resources:
requests:
cpu: "4"
memory: "32Gi"
valkey:
persistence:
size: "20Gi"
# Airflow configuration
airflow:
worker:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
resources:
requests:
cpu: "2"
memory: "8Gi"
# Monitoring (optional)
monitoring:
enabled: true
prometheus:
persistence:
size: "50Gi"
grafana:
persistence:
size: "10Gi"
Apply the configuration:
kubectl apply -f lakehousecat-production.yaml
Step 7: Monitor Deployment
Watch the deployment progress:
# Check instance status
kubectl get lakehousecat -w
# Check all pods
kubectl get pods -n lakehousecat-prod -w
# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator
# Check specific service
kubectl describe lakehousecat production
Step 8: Configure DNS and Ingress
Once deployed, configure DNS to point to your cluster's LoadBalancer or Ingress:
Get LoadBalancer IP/Hostname:
kubectl get svc -n lakehousecat-prod lakehousecat-ui
Configure DNS: Create A or CNAME records:
lakehousecat.company.com → <LOADBALANCER_IP_OR_HOSTNAME>
Verify TLS:
curl -I https://lakehousecat.company.com
Step 9: Access Lakehousecat
Open your browser to your configured domain:
https://lakehousecat.company.com
Create your admin account and begin onboarding users.
High Availability Configuration
For mission-critical deployments, configure high availability:
Multi-Replica Services
spec:
analytics:
replicaCount: 3
lhc:
replicaCount: 3
llm:
replicaCount: 2
ui:
replicaCount: 3
Database Replication
spec:
postgresql:
replicaCount: 3
replication:
enabled: true
clickhouse:
shards: 3
replicaCount: 2 # 2 replicas per shard
Pod Disruption Budgets
The operator automatically creates Pod Disruption Budgets (PDBs) for critical services to ensure availability during cluster maintenance.
Security Best Practices
Network Policies
Enable network policies to isolate namespaces:
# Operator values
networkPolicy:
enabled: true
RBAC
The operator creates minimal RBAC permissions. Review and adjust as needed:
kubectl get rolebindings -n lakehousecat-prod
kubectl get clusterrolebindings | grep lakehousecat
Secrets Management
Use external secret management for sensitive data:
- AWS: AWS Secrets Manager or Parameter Store
- Google Cloud: Secret Manager
- Azure: Key Vault
Integrate using Kubernetes External Secrets Operator.
TLS/HTTPS
Always enable TLS for production:
spec:
domain:
tls:
enabled: true
letsEncrypt: true # Or provide your own certificates
Monitoring and Observability
Prometheus Metrics
The operator exposes Prometheus metrics:
# Enable ServiceMonitor for Prometheus Operator
metrics:
serviceMonitor:
enabled: true
namespace: monitoring
Grafana Dashboards
If monitoring is enabled in the Lakehousecat CR, Grafana is automatically deployed with pre-configured dashboards.
Access Grafana:
kubectl port-forward -n lakehousecat-prod svc/lakehousecat-grafana 3000:80
Logging
Centralize logs using your preferred logging solution:
- AWS: CloudWatch Logs
- Google Cloud: Cloud Logging
- Azure: Azure Monitor
- ELK Stack: Elasticsearch, Logstash, Kibana
- Loki: Grafana Loki
Backup and Disaster Recovery
Database Backups
Configure automated backups for PostgreSQL and ClickHouse:
spec:
postgresql:
backup:
enabled: true
schedule: "0 2 * * *" # Daily at 2 AM
retention: 30 # Keep 30 days
clickhouse:
backup:
enabled: true
schedule: "0 3 * * *"
retention: 30
Object Storage Backups
Enable versioning on your S3/Cloud Storage bucket for data protection.
Disaster Recovery Plan
- Regular backups of all databases
- GitOps for Lakehousecat CR configurations
- Multi-region deployment for critical workloads
- RTO/RPO definitions and testing
Scaling
Scale your deployment as needs grow. See the Scaling documentation for detailed guidance.
Quick scaling example:
# Scale analytics service
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"replicaCount":5}}}'
# Enable autoscaling
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"autoscaling":{"enabled":true,"minReplicas":3,"maxReplicas":10}}}}'
Upgrading
Operator Upgrade
helm upgrade lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
--version 0.0.40
Instance Upgrade
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"version":"0.0.40"}}'
The operator handles rolling upgrades with zero downtime.
Platform-Specific Guides
For detailed, platform-specific deployment instructions:
- AWS EKS - Amazon Elastic Kubernetes Service
- Google GKE - Google Kubernetes Engine
- Azure AKS - Azure Kubernetes Service
- On-Premise - Self-managed Kubernetes clusters
Troubleshooting
Pods Not Starting
# Check pod events
kubectl describe pod -n lakehousecat-prod <pod-name>
# Check operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator
# Check instance status
kubectl describe lakehousecat production
License Validation Failures
# Verify handshake secret
kubectl get secret lhc-handshake-secret -n lhc-operator -o yaml
# Check network connectivity to provider backend
kubectl run -it --rm debug --image=alpine --restart=Never -- \
wget -O- https://license.lakehousecat.com/health
Storage Issues
# Check PVCs
kubectl get pvc -n lakehousecat-prod
# Check storage class
kubectl get storageclass
# Describe PVC for events
kubectl describe pvc -n lakehousecat-prod <pvc-name>
Support
- Documentation: https://docs.lakehousecat.com
- Portal: https://portal.lakehousecat.com
- Support: portal.lakehousecat.com
- Community: Discord (coming soon)