Skip to main content
Version: 0.0.41

Deploy to Google GKE

Deploy Lakehousecat to Google Kubernetes Engine (GKE), Google Cloud's managed Kubernetes service.

Overview​

Deployment Status: ⚠️ Not yet tested by Lakehousecat — community-contributed guide

Not Yet Tested

This environment has not yet been tested by Lakehousecat. The guide is provided as a reference based on standard Kubernetes deployment patterns. If you encounter environment-specific issues, please report them via portal.lakehousecat.com.

Cloud Storage not supported

Lakehousecat does not currently support Google Cloud Storage as an object storage backend. SeaweedFS is deployed internally by the Operator and handles all object storage needs. Google Cloud SQL is not supported — Lakehousecat manages its own PostgreSQL and ClickHouse instances.

Cluster administration

GCP-specific configurations — IAM policies, networking, and monitoring integrations — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.

Prerequisites​

Google Cloud Setup​

  • Google Cloud Platform account with billing enabled
  • Appropriate IAM permissions for GKE cluster creation

Required Tools​

  • gcloud CLI: Google Cloud command-line tool
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+

Install Tools (macOS)​

brew install google-cloud-sdk
gcloud components install kubectl
brew install helm

Authenticate with Google Cloud​

gcloud auth login
gcloud config set project YOUR_PROJECT_ID
gcloud config set compute/region us-central1

Step 1: Provision GKE Cluster​

gcloud container clusters create lakehousecat-prod \
--region us-central1 \
--cluster-version 1.29 \
--machine-type n2-standard-16 \
--num-nodes 3 \
--min-nodes 3 \
--max-nodes 10 \
--enable-autoscaling \
--enable-autorepair \
--disk-type pd-ssd \
--disk-size 200 \
--enable-ip-alias \
--addons HorizontalPodAutoscaling,HttpLoadBalancing,GcePersistentDiskCsiDriver \
--labels environment=production,app=lakehousecat

Common instance types:

  • n2-standard-16: 16 vCPU, 64 GB RAM (production)
  • n2-standard-8: 8 vCPU, 32 GB RAM (smaller workloads)

Option 2: GKE Autopilot​

gcloud container clusters create-auto lakehousecat-autopilot \
--region us-central1 \
--cluster-version 1.29
Autopilot vs Standard

Autopilot: Google manages nodes, scaling, and security — per-pod billing. Standard: You manage nodes and instance types.

Verify Cluster​

gcloud container clusters get-credentials lakehousecat-prod --region us-central1
kubectl cluster-info
kubectl get nodes

Step 2: Configure GKE-Specific Components​

Verify Storage Class​

GKE provides multiple storage classes out of the box:

kubectl get storageclass

Recommended for production: premium-rwo (SSD-backed persistent disks)

Ingress Controller​

An ingress controller is required to expose Lakehousecat via a domain. GKE's built-in HTTP Load Balancing addon (enabled above) can serve as the ingress controller. Alternatively, install NGINX Ingress:

helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update

helm install ingress-nginx ingress-nginx/ingress-nginx \
--namespace ingress-nginx \
--create-namespace \
--set controller.service.type=LoadBalancer
Ingress configuration not yet validated

GKE ingress routing for Lakehousecat has not been end-to-end tested. The exact ingress configuration may require adjustments depending on the ingress controller used.

Step 3: Install Lakehousecat Operator​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait

Verify:

kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 4: Create Handshake Secret​

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 5: Deploy Lakehousecat Instance​

lakehousecat-gke.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.40"
architecture: "amd64"

# SSD-backed persistent storage
postgresql:
persistence:
size: "100Gi"
storageClass: "premium-rwo"

clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "premium-rwo"

valkey:
persistence:
size: "20Gi"
storageClass: "premium-rwo"

seaweedfs:
enabled: true
persistence:
size: "100Gi"
storageClass: "premium-rwo"

analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"

ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"

Apply:

kubectl apply -f lakehousecat-gke.yaml

Step 6: Monitor Deployment​

kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 7: Configure DNS and Access​

Get Load Balancer IP​

kubectl get svc -n lakehousecat-prod lakehousecat-ui

Configure Cloud DNS​

gcloud dns record-sets create lakehousecat.company.com. \
--zone=company-zone \
--type=A \
--ttl=300 \
--rrdatas=<EXTERNAL-IP>

Troubleshooting​

Pods Pending (Insufficient Resources)​

gcloud container clusters resize lakehousecat-prod \
--region us-central1 \
--num-nodes 5

Persistent Disk Mounting Issues​

kubectl get pods -n kube-system | grep csi
kubectl describe storageclass premium-rwo

Load Balancer Not Provisioned​

kubectl describe svc lakehousecat-ui -n lakehousecat-prod
kubectl get pods -n kube-system | grep l7

Next Steps​

  • Scaling - Configure autoscaling for production
  • Monitoring - Review monitoring options
  • Security - Review security configuration

Support​