Skip to main content
Version: Next

Kubernetes Cluster Deployment

Deploy Lakehousecat to production-grade Kubernetes clusters in the cloud or on-premise. This guide covers deployment to managed Kubernetes services and self-hosted clusters.

Overview​

Kubernetes cluster deployment is recommended for:

  • Production environments with high availability requirements
  • Multi-user deployments (teams and organizations)
  • Scalable workloads with dynamic resource needs
  • Enterprise deployments with compliance and security requirements

Supported Kubernetes Platforms​

Lakehousecat Operator runs on any standard Kubernetes 1.29+ cluster:

Managed Kubernetes Services​

PlatformStatusGuide
AWS EKS✅ Tested & Supported→ EKS Guide
Google GKE✅ Tested & Supported→ GKE Guide
Azure AKS✅ Tested & Supported→ AKS Guide

Self-Managed Kubernetes​

PlatformStatusGuide
Hetzner Cloud (kube-hetzner, k3s)✅ Tested & Supported→ Hetzner Guide
On-Premise✅ Compatible→ On-Premise Guide
Bare Metal✅ CompatibleUse On-Premise Guide
Private Cloud✅ CompatibleUse On-Premise Guide

Prerequisites​

Before deploying to a Kubernetes cluster, ensure you have:

Cluster Requirements​

  • Kubernetes: Version 1.29 or higher
  • kubectl: Installed and configured for your cluster
  • Helm: Version 3.8 or higher
  • Persistent Storage: Dynamic volume provisioner (StorageClass)
  • LoadBalancer or Ingress Controller: For external access

Resource Requirements​

Minimum Production Configuration:

  • 16 CPU cores
  • 64 GB RAM
  • 500 GB persistent storage

Recommended Production Configuration:

  • 32+ CPU cores
  • 128+ GB RAM
  • 1+ TB persistent storage
  • Auto-scaling enabled

Network Requirements​

  • Outbound HTTPS access to:
    • license.lakehousecat.com (license validation)
    • ghcr.io (Lakehousecat Custom Images + Operator — public, no credentials required) and the official upstream registries of the third-party services the operator deploys — all container image pulls are handled automatically by the operator
  • Ingress capability for external user access
  • TLS certificates for HTTPS (Let's Encrypt recommended)

License​

General Deployment Process​

The deployment process is consistent across all Kubernetes platforms:

Step 1: Provision Kubernetes Cluster​

Create a Kubernetes cluster on your chosen platform (EKS, GKE, AKS, Hetzner Cloud, or on-premise).

Key Configuration:

  • Kubernetes version: 1.29 or higher
  • Node pool: Appropriate instance types for your workload
  • Storage: Dynamic volume provisioner enabled
  • Networking: LoadBalancer or Ingress controller

See platform-specific guides for detailed instructions.

Reference Terraform repositories

Ready-made Terraform configurations for the cluster layer are available in the lakehousecat GitHub organisation, one repository per provider: lhc-kube-aws, lhc-kube-azure, lhc-kube-gcp and lhc-kube-hetzner. They are a standard starter setup for evaluation and simple production clusters, not a managed service. The application rollout (Step 2 onwards) is the same whether you use them or the platform guides below.

Step 2: Configure kubectl​

Ensure kubectl is configured to connect to your cluster:

# Verify cluster access
kubectl cluster-info
kubectl get nodes

# Check available storage classes
kubectl get storageclass

Step 3: Install Lakehousecat Operator​

Install the operator using Helm:

# Add Lakehousecat Helm repository (OCI registry)
helm pull oci://public.ecr.aws/lakehousecat/charts/lakehousecat-operator --version 0.0.42

# Or add traditional Helm repo
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

# Install operator
# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lakehousecat-prod

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait

Production Values (optional):

operator-values.yaml
# Required: the instance namespaces the operator may manage. It gets rights only
# there (one RoleBinding each) and does not create them. Another instance later:
# create its namespace, add it here, run helm upgrade.
targetNamespaces:
- lakehousecat-prod

replicaCount: 1

operator:
logging:
level: info
format: json
leaderElection:
enabled: false

resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: 2000m
memory: 2Gi

metrics:
serviceMonitor:
enabled: true
namespace: monitoring
interval: 30s

networkPolicy:
enabled: true

Install with custom values:

helm install lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
--create-namespace \
-f operator-values.yaml \
--wait

Verify operator deployment:

kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 4: Create Handshake Secret​

Create a Kubernetes secret with your handshake key:

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY_FROM_PORTAL \
-n lhc-operator

Verify secret creation:

kubectl get secret lhc-handshake-secret -n lhc-operator

Step 5: Create Required Secrets​

All sensitive values — model API keys, the Mapbox token, external S3 credentials, OAuth client secrets — are read by the operator from pre-created Kubernetes secrets, never from the CR. Create the secrets you need before deploying the instance.

See Required Secrets for the full list, the exact keys per provider, and ready-made helper scripts. At minimum you created the handshake secret in Step 4; add the optional secrets for any feature you intend to use (models, Mapbox, external S3, SSO).

Graceful by design

Missing optional secrets do not block the deployment — the affected feature is simply skipped and reported via a status condition. You can add secrets later and re-trigger a reconcile.

Step 5b: Configure Storage​

Lakehousecat requires persistent storage for databases and object storage.

Option 1: Cloud Object Storage​

Use cloud-native S3-compatible storage. Credentials come from the lhc-storage-s3-secret (see Required Secrets), and the CR references it:

spec:
objectStorage:
type: s3
region: "us-east-1"
bucketPrefix: "my-company-lakehousecat"
credentialsSecretName: "lhc-storage-s3-secret" # keys: access-key, secret-key

Set endpoint to target an S3-compatible store other than native AWS S3 (e.g. Google Cloud Storage's S3-Interop API). This also enables hybrid / multi-cloud setups: the Kubernetes cluster can run on any provider (or on-premise) while Lakehousecat stores objects in S3 elsewhere — all that's required is network access to the endpoint and valid credentials. See Required Secrets — Custom S3-compatible endpoints for details. Azure Blob Storage is not supported as an objectStorage backend — it does not speak the S3 protocol.

Option 2: In-Cluster SeaweedFS (Default)​

Deploy SeaweedFS within the cluster (managed by the operator, no credentials needed):

spec:
objectStorage:
type: seaweedfs
volume:
persistence:
size: "500Gi"
storageClass: "gp3" # Use your cluster's storage class

Step 6: Deploy Lakehousecat Instance​

Create a Lakehousecat Custom Resource with production configuration:

lakehousecat-production.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "amd64"

# Storage configuration — external S3 (credentials via pre-created secret).
# Omit this block entirely to use the default in-cluster SeaweedFS.
objectStorage:
type: s3
region: "us-east-1"
bucketPrefix: "my-company-lakehousecat"
credentialsSecretName: "lhc-storage-s3-secret" # keys: access-key, secret-key

# Domain configuration with TLS
domain:
baseDomain: "lakehousecat.company.com"
tls:
enabled: true
letsEncrypt: true
email: "admin@company.com"

# Production resource allocations
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "2"
memory: "8Gi"

llm:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"

ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"

postgresql:
persistenceSize: "100Gi"
storageClass: "gp3"
resources:
requests:
cpu: "2"
memory: "8Gi"

clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "gp3"
resources:
requests:
cpu: "4"
memory: "32Gi"

valkey:
persistence:
size: "20Gi"

# Airflow configuration
airflow:
worker:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
resources:
requests:
cpu: "2"
memory: "8Gi"

# Monitoring — exposes /metrics endpoints on backend services
# for scraping by the customer's existing monitoring stack.
# The Operator does NOT deploy Prometheus or Grafana.
monitoring:
metricsEnabled: true

Apply the configuration:

kubectl apply -f lakehousecat-production.yaml

Step 7: Monitor Deployment​

Watch the deployment progress:

# Check instance status
kubectl get lakehousecat -w

# Check all pods
kubectl get pods -n lakehousecat-prod -w

# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

# Check specific service
kubectl describe lakehousecat production

Step 8: Configure DNS and Ingress​

Once deployed, configure DNS to point to your cluster's LoadBalancer or Ingress:

Get LoadBalancer IP/Hostname:

kubectl get svc -n lakehousecat-prod lakehousecat-ui

Configure DNS: Create A or CNAME records:

lakehousecat.company.com → <LOADBALANCER_IP_OR_HOSTNAME>

Verify TLS:

curl -I https://lakehousecat.company.com

Step 9: Access Lakehousecat​

Open your browser to your configured domain:

https://lakehousecat.company.com

Create your admin account and begin onboarding users.

High Availability Configuration​

For mission-critical deployments, configure high availability:

Multi-Replica Services​

spec:
analytics:
replicaCount: 3
lhc:
replicaCount: 3
llm:
replicaCount: 2
ui:
replicaCount: 3

Database Replication​

spec:
postgresql:
replicaCount: 3
replication:
enabled: true

clickhouse:
shards: 3
replicaCount: 2 # 2 replicas per shard

Pod Disruption Budgets​

The operator automatically creates Pod Disruption Budgets (PDBs) for critical services to ensure availability during cluster maintenance.

Security Best Practices​

Network Policies​

# Operator values
networkPolicy:
enabled: true

networkPolicy.enabled: true makes the Operator create the NetworkPolicy objects (roughly 15 of them) that isolate Lakehousecat's namespaces. That is a prerequisite for isolation, not a guarantee of it: NetworkPolicy is a plain Kubernetes API object — the API server accepts it and otherwise does nothing with it. Whether it's actually enforced depends entirely on the cluster's CNI. A cluster whose CNI doesn't implement NetworkPolicy (plain Flannel without an add-on, a bare VPC-native CNI, kindnet, ...) accepts all 15 objects and enforces none of them, with nothing in the cluster indicating that.

  • GKE: enforced by default via Dataplane V2 (--enable-dataplane-v2 at cluster creation) — see Deploy to Google GKE for the verification steps.
  • EKS / AKS / on-premise: confirm with your cluster administrator that the CNI enforces NetworkPolicy (e.g. the VPC-CNI network-policy agent on EKS, Calico/Cilium/Azure NPM on AKS, Calico/Cilium on a self-managed cluster) — this is not on by default on every managed Kubernetes offering.

Verify enforcement after deploying, on any platform — don't just trust the cluster's default:

kubectl get lakehousecat -A

The NP-Enforced column reports the Operator's own live detection. If it reads False, every NetworkPolicy in the cluster is inert regardless of networkPolicy.enabled.

RBAC​

The operator chart binds the operator's rights per instance namespace (targetNamespaces); the only cluster-wide binding is a read-mostly role for cluster-scoped objects (see Operator Permissions). Review them:

kubectl get rolebindings -n lakehousecat-prod
kubectl get clusterrolebindings | grep lhc-operator
kubectl auth can-i --list --as=system:serviceaccount:lhc-operator:lhc-operator -n kube-system

Secrets Management​

Use external secret management for sensitive data:

  • AWS: AWS Secrets Manager or Parameter Store
  • Google Cloud: Secret Manager
  • Azure: Key Vault

Integrate using Kubernetes External Secrets Operator.

TLS/HTTPS​

Always enable TLS for production:

spec:
domain:
tls:
enabled: true
letsEncrypt: true # Or provide your own certificates

Monitoring and Observability​

The Operator does not deploy a monitoring stack. It exposes observability data through open interfaces — connect your own tooling.

Prometheus Metrics — /metrics Endpoint​

Each backend service (lhc, llm, analytics, semantic, audio) exposes a /metrics endpoint in the Prometheus exposition format. The endpoint is unauthenticated by design (cluster-internal, follows Prometheus convention) and toggled via monitoring.metricsEnabled (default true).

For customers running prometheus-operator, a ServiceMonitor such as the following scrapes Lakehousecat services:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: lakehousecat-services
namespace: lakehousecat-prod
spec:
selector:
matchLabels:
app.kubernetes.io/part-of: lakehousecat
endpoints:
- port: http
path: /metrics
interval: 30s

Datadog, New Relic and similar agents support equivalent auto-discovery via Kubernetes service annotations — consult your agent's documentation.

Logging​

Centralize logs using your preferred logging solution. All backend services produce structured application logs that are written to the SeaweedFS backend_logs/ bucket and to stdout/stderr. Compatible collectors include:

  • AWS: CloudWatch Logs
  • Google Cloud: Cloud Logging
  • Azure: Azure Monitor
  • ELK Stack: Elasticsearch, Logstash, Kibana
  • Loki: Grafana Loki
  • Datadog / New Relic / Splunk: via their respective Kubernetes agents

Backup and Disaster Recovery​

Database Backups​

Configure automated backups for PostgreSQL and ClickHouse:

spec:
postgresql:
backup:
enabled: true
schedule: "0 2 * * *" # Daily at 2 AM
retention: 30 # Keep 30 days

clickhouse:
backup:
enabled: true
schedule: "0 3 * * *"
retention: 30

Object Storage Backups​

Enable versioning on your S3/Cloud Storage bucket for data protection.

Disaster Recovery Plan​

  1. Regular backups of all databases
  2. GitOps for Lakehousecat CR configurations
  3. Multi-region deployment for critical workloads
  4. RTO/RPO definitions and testing

Scaling​

Scale your deployment as needs grow. See the Scaling documentation for detailed guidance.

Quick scaling example:

# Scale analytics service
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"replicaCount":5}}}'

# Enable autoscaling
kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"analytics":{"autoscaling":{"enabled":true,"minReplicas":3,"maxReplicas":10}}}}'

Upgrading​

Operator Upgrade​

helm upgrade lhc-operator lakehousecat/lakehousecat-operator \
-n lhc-operator \
-f operator-values.yaml \
--version 0.0.42

Pass the same values (including targetNamespaces) on every upgrade; without them the chart refuses to render.

Instance Upgrade​

kubectl patch lakehousecat production --type=merge \
-p '{"spec":{"version":"0.0.42"}}'

The operator handles rolling upgrades with zero downtime.

Platform-Specific Guides​

For detailed, platform-specific deployment instructions:

Troubleshooting​

Pods Not Starting​

# Check pod events
kubectl describe pod -n lakehousecat-prod <pod-name>

# Check operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

# Check instance status
kubectl describe lakehousecat production

License Validation Failures​

# Verify handshake secret
kubectl get secret lhc-handshake-secret -n lhc-operator -o yaml

# Check network connectivity to provider backend
kubectl run -it --rm debug --image=alpine --restart=Never -- \
wget -O- https://license.lakehousecat.com/health

Storage Issues​

# Check PVCs
kubectl get pvc -n lakehousecat-prod

# Check storage class
kubectl get storageclass

# Describe PVC for events
kubectl describe pvc -n lakehousecat-prod <pvc-name>

Support​