Deploy to Google GKE
Deploy Lakehousecat to Google Kubernetes Engine (GKE), Google Cloud's managed Kubernetes service.
Overview
Deployment Status: ⚠️ Not yet tested by Lakehousecat — community-contributed guide
This environment has not yet been tested by Lakehousecat. The guide is provided as a reference based on standard Kubernetes deployment patterns. If you encounter environment-specific issues, please report them via portal.lakehousecat.com.
Lakehousecat does not currently support Google Cloud Storage as an object storage backend. SeaweedFS is deployed internally by the Operator and handles all object storage needs. Google Cloud SQL is not supported — Lakehousecat manages its own PostgreSQL and ClickHouse instances.
GCP-specific configurations — IAM policies, networking, and monitoring integrations — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.
Prerequisites
Google Cloud Setup
- Google Cloud Platform account with billing enabled
- Appropriate IAM permissions for GKE cluster creation
Required Tools
- gcloud CLI: Google Cloud command-line tool
- kubectl: Kubernetes command-line tool
- Helm: Version 3.8+
Install Tools (macOS)
brew install google-cloud-sdk
gcloud components install kubectl
brew install helm
Authenticate with Google Cloud
gcloud auth login
gcloud config set project YOUR_PROJECT_ID
gcloud config set compute/region us-central1
Step 1: Provision GKE Cluster
Option 1: GKE Standard (Recommended)
gcloud container clusters create lakehousecat-prod \
--region us-central1 \
--cluster-version 1.29 \
--machine-type n2-standard-16 \
--num-nodes 3 \
--min-nodes 3 \
--max-nodes 10 \
--enable-autoscaling \
--enable-autorepair \
--disk-type pd-ssd \
--disk-size 200 \
--enable-ip-alias \
--addons HorizontalPodAutoscaling,HttpLoadBalancing,GcePersistentDiskCsiDriver \
--labels environment=production,app=lakehousecat
Common instance types:
n2-standard-16: 16 vCPU, 64 GB RAM (production)n2-standard-8: 8 vCPU, 32 GB RAM (smaller workloads)
Option 2: GKE Autopilot
gcloud container clusters create-auto lakehousecat-autopilot \
--region us-central1 \
--cluster-version 1.29
Autopilot: Google manages nodes, scaling, and security — per-pod billing. Standard: You manage nodes and instance types.
Verify Cluster
gcloud container clusters get-credentials lakehousecat-prod --region us-central1
kubectl cluster-info
kubectl get nodes
Step 2: Configure GKE-Specific Components
Verify Storage Class
GKE provides multiple storage classes out of the box:
kubectl get storageclass
Recommended for production: premium-rwo (SSD-backed persistent disks)
Ingress Controller
An ingress controller is required to expose Lakehousecat via a domain. GKE's built-in HTTP Load Balancing addon (enabled above) can serve as the ingress controller. Alternatively, install NGINX Ingress:
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update
helm install ingress-nginx ingress-nginx/ingress-nginx \
--namespace ingress-nginx \
--create-namespace \
--set controller.service.type=LoadBalancer
GKE ingress routing for Lakehousecat has not been end-to-end tested. The exact ingress configuration may require adjustments depending on the ingress controller used.
Step 3: Install Lakehousecat Operator
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update
helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait
Verify:
kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 4: Create Handshake Secret
kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator
Step 5: Deploy Lakehousecat Instance
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"
customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"
version: "0.0.40"
architecture: "amd64"
# SSD-backed persistent storage
postgresql:
persistence:
size: "100Gi"
storageClass: "premium-rwo"
clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "premium-rwo"
valkey:
persistence:
size: "20Gi"
storageClass: "premium-rwo"
seaweedfs:
enabled: true
persistence:
size: "100Gi"
storageClass: "premium-rwo"
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"
Apply:
kubectl apply -f lakehousecat-gke.yaml
Step 6: Monitor Deployment
kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 7: Configure DNS and Access
Get Load Balancer IP
kubectl get svc -n lakehousecat-prod lakehousecat-ui
Configure Cloud DNS
gcloud dns record-sets create lakehousecat.company.com. \
--zone=company-zone \
--type=A \
--ttl=300 \
--rrdatas=<EXTERNAL-IP>
Troubleshooting
Pods Pending (Insufficient Resources)
gcloud container clusters resize lakehousecat-prod \
--region us-central1 \
--num-nodes 5
Persistent Disk Mounting Issues
kubectl get pods -n kube-system | grep csi
kubectl describe storageclass premium-rwo
Load Balancer Not Provisioned
kubectl describe svc lakehousecat-ui -n lakehousecat-prod
kubectl get pods -n kube-system | grep l7
Next Steps
- Scaling - Configure autoscaling for production
- Monitoring - Review monitoring options
- Security - Review security configuration
Support
- GKE Documentation: https://cloud.google.com/kubernetes-engine/docs
- Lakehousecat Portal: https://portal.lakehousecat.com