Deploy to Google GKE
Deploy Lakehousecat to Google Kubernetes Engine (GKE), Google Cloud's managed Kubernetes service.
Overview
A reference Terraform configuration for this platform is available as
lhc-kube-gcp in the lakehousecat GitHub organisation. It provisions the cluster layer
only; the application rollout below is the same either way.
Deployment Status: ✅ Verified by Lakehousecat — full stack deployed and validated on GKE
Lakehousecat does not have a native Google Cloud Storage integration. By default the Operator
deploys SeaweedFS in-cluster (no configuration needed). Two external-storage alternatives are
available as a hybrid setup: AWS S3 (fully supported), or Google Cloud Storage itself via
its S3-Interop API (endpoint: https://storage.googleapis.com with GCS HMAC keys — optional,
not yet verified end-to-end by Lakehousecat). See Configure Storage.
Google Cloud SQL is not supported — Lakehousecat manages its own PostgreSQL and ClickHouse instances.
GCP-specific configurations — IAM policies, networking, and monitoring integrations — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.
Prerequisites
Google Cloud Setup
- Google Cloud Platform account with billing enabled
- Appropriate IAM permissions for GKE cluster creation
Required Tools
- gcloud CLI: Google Cloud command-line tool
- kubectl: Kubernetes command-line tool
- Helm: Version 3.8+
- gke-gcloud-auth-plugin: required by
kubectlfor every request against a GKE cluster (see Cluster Access below)
Install Tools (macOS)
brew install google-cloud-sdk
gcloud components install kubectl
gcloud components install gke-gcloud-auth-plugin
brew install helm
On Homebrew, gcloud components install gke-gcloud-auth-plugin can print "Update done!" but still
exit with code 1, and the binary lands under /opt/homebrew/share/google-cloud-sdk/bin without
a symlink into /opt/homebrew/bin. If kubectl reports the plugin as missing afterwards, add that
directory to your PATH.
Authenticate with Google Cloud
gcloud auth login
gcloud config set project YOUR_PROJECT_ID
gcloud config set compute/region us-central1
Cluster access: the auth plugin has two modes
kubectl calls gke-gcloud-auth-plugin on every request against a GKE cluster. The plugin has two
credential sources, controlled by a single flag:
| Kubeconfig source | args | Credentials from |
|---|---|---|
gcloud container clusters get-credentials | (none) | the active gcloud account |
| A kubeconfig generated without that flag (e.g. via Terraform output) | --use_application_default_credentials | GOOGLE_APPLICATION_CREDENTIALS |
Without the flag, the plugin ignores Application Default Credentials entirely, even if
GOOGLE_APPLICATION_CREDENTIALS is set and valid, and fails with:
failed to retrieve access token: … (gcloud.config.config-helper)
You do not currently have an active account selected.
The message says "no active account," which points at the wrong problem — the credentials file is
fine, the plugin just isn't reading it. If you're generating kubeconfig outside of
gcloud container clusters get-credentials (for example from automation), pass
--use_application_default_credentials explicitly.
Check your SSD quota before creating the cluster
A fresh GCP project defaults to 500 GB SSD_TOTAL_GB per region — a quota shared by
pd-balanced and pd-ssd disks. Node boot disks are counted against this quota by default, and on
a 3+ node cluster they can consume most of it before the application's own volumes are provisioned.
Check the current quota first:
gcloud compute regions describe <region> --project <project> \
--format="flattened(quotas)" | grep -A2 SSD_TOTAL_GB
If it's tight, either request an increase under IAM & Admin → Quotas (usually approved within
minutes with active billing, but not guaranteed), or — the more reliable fix — set node boot disks
to pd-standard, which is billed against the much larger DISKS_TOTAL_GB quota instead. See
Verify Storage Class below for why pd-standard boot disks don't affect
application performance.
Step 1: Provision GKE Cluster
Reserve a static ingress IP (recommended)
gcloud compute addresses create lakehousecat-ingress --region us-central1
Reserving the address up front means it survives a cluster teardown and rebuild — in testing, the same IP was reused automatically on rebuild and DNS records didn't need to change. Without a reservation, a rebuild allocates a new ephemeral IP and DNS propagation delay follows.
Option 1: GKE Standard (Recommended)
gcloud container clusters create lakehousecat-prod \
--zone us-central1-a \
--cluster-version 1.29 \
--machine-type n2-standard-16 \
--num-nodes 3 \
--min-nodes 3 \
--max-nodes 10 \
--enable-autoscaling \
--enable-autorepair \
--disk-type pd-standard \
--disk-size 200 \
--default-max-pods-per-node 110 \
--enable-ip-alias \
--enable-dataplane-v2 \
--addons HorizontalPodAutoscaling,HttpLoadBalancing,GcePersistentDiskCsiDriver \
--labels environment=production,app=lakehousecat
Common instance types:
n2-standard-16: 16 vCPU, 64 GB RAM (production)n2-standard-8: 8 vCPU, 32 GB RAM (smaller workloads)
Use a zonal cluster (--zone), not a regional one (--region). Persistent Disks are
zone-bound; a regional cluster spreads nodes across three zones, and a rescheduled pod can end up
in a zone that doesn't hold its volume. A zonal cluster also fits the GKE free-tier allowance
(one zonal cluster, ~$73/month cluster management fee).
Boot disks use pd-standard here deliberately — see Check your SSD quota
above. Don't shrink the boot disk size to save quota instead: GKE reserves roughly half of a
node's boot disk as non-allocatable ephemeral storage (measured 43.8 GiB allocatable out of
94.3 GiB total). At the full Lakehousecat stack, per-node ephemeral usage reached 16–22.8 GiB — a
50 GB boot disk would leave too little headroom and risk ephemeral-storage evictions.
GKE assigns each node a contiguous CIDR block for pods, not individual addresses. At the default
of 110 pods/node that's a /24 per node — a /22 pod range only covers four nodes. Running out
shows up as pods stuck in ContainerCreating with "failed to assign IP," which looks like a
scheduling problem rather than an IP exhaustion issue. Size your pod IP range for the node count
you expect to scale to, not just the starting count.
--enable-dataplane-v2 above isn't optional decoration: GKE Dataplane V2 is what actually
enforces the NetworkPolicy objects the Operator creates (~15 of them). A NetworkPolicy is
just a Kubernetes API object — without an enforcing CNI it is accepted and silently does
nothing. Dataplane V2 is GKE's modern default for new clusters, but if you're customizing
cluster creation, don't drop this flag or switch to the legacy --enable-network-policy
Calico add-on path without understanding the tradeoff.
Verify enforcement is actually active, not just requested — check the cluster, not just the flag you passed:
kubectl get pods -n kube-system -l k8s-app=cilium
# anetd (Dataplane V2's Cilium-based dataplane) should show 1/1 or 2/2 Running on every node
kubectl get lakehousecat -A
# NP-Enforced column reads True once the Operator's own detector confirms enforcement
If NP-Enforced shows False, every NetworkPolicy in the cluster is inert regardless of what
networkPolicy.enabled says in the Operator's Helm values — see
Network Policies below.
Option 2: GKE Autopilot
gcloud container clusters create-auto lakehousecat-autopilot \
--region us-central1 \
--cluster-version 1.29
Autopilot: Google manages nodes, scaling, and security — per-pod billing. Standard: You manage nodes and instance types.
Verify Cluster
gcloud container clusters get-credentials lakehousecat-prod --zone us-central1-a
kubectl cluster-info
kubectl get nodes
Step 2: Configure GKE-Specific Components
Verify Storage Class
GKE provides three storage classes out of the box:
kubectl get storageclass
Recommended and verified: standard-rwo — the CSI driver's SSD-backed (pd-balanced) class.
Don't use plain standard: it's the older in-tree provisioner on pd-standard (HDD) — it works,
but is noticeably slower under load, and the difference isn't obvious until you're running real
workloads. premium-rwo (pd-ssd) is also available but not required for Lakehousecat's workload.
Ingress Controller
An ingress controller is required to expose Lakehousecat via a domain. GKE's built-in HTTP Load Balancing addon (enabled above) can serve as the ingress controller. Alternatively, install NGINX Ingress:
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update
helm install ingress-nginx ingress-nginx/ingress-nginx \
--namespace ingress-nginx \
--create-namespace \
--set controller.service.type=LoadBalancer
Step 3: Install Lakehousecat Operator
helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update
# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lakehousecat-prod
helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait
Verify:
kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 4: Create Handshake Secret
kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator
Step 5: Deploy Lakehousecat Instance
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"
customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"
version: "0.0.42"
architecture: "amd64"
# SSD-backed persistent storage
postgresql:
persistenceSize: "100Gi"
storageClass: "standard-rwo"
clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "standard-rwo"
valkey:
persistence:
size: "20Gi"
storageClass: "standard-rwo"
# In-cluster SeaweedFS (default). Omit this block entirely to use the same default,
# or replace it with `objectStorage: {type: s3, ...}` for external S3 — see
# "Configure Storage" in the Kubernetes Cluster guide.
objectStorage:
type: seaweedfs
volume:
persistence:
size: "100Gi"
storageClass: "standard-rwo"
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"
Apply:
kubectl apply -f lakehousecat-gke.yaml
Step 6: Monitor Deployment
kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w
kubectl logs -f deployment/lhc-operator -n lhc-operator
Step 7: Configure DNS and Access
Get Load Balancer IP
kubectl get svc -n lakehousecat-prod lakehousecat-ui
Configure Cloud DNS
gcloud dns record-sets create lakehousecat.company.com. \
--zone=company-zone \
--type=A \
--ttl=300 \
--rrdatas=<EXTERNAL-IP>
Troubleshooting
Pods Pending (Insufficient Resources)
gcloud container clusters resize lakehousecat-prod \
--zone us-central1-a \
--num-nodes 5
Pods Pending — deployment stalls mid-stack with "context deadline exceeded"
If clickhouse-shard0-0 or postgresql-1 stay Pending while everything else is running, and
kubectl describe pod only shows a generic volume-binding timeout:
FailedScheduling: running PreBind plugin "VolumeBinding": binding volumes:
context deadline exceeded
that's a follow-on symptom, not the cause. Check the PVC's own events instead:
kubectl describe pvc <pvc-name> -n <namespace>
ProvisioningFailed … QUOTA_EXCEEDED: Quota 'SSD_TOTAL_GB' exceeded. Limit: 500.0 in region <region>
This is the SSD_TOTAL_GB project quota described under Check your SSD quota
— it fails partway through the deployment, not at cluster creation, which is what makes it easy to
mistake for a scheduling or Operator problem. Fix by switching node boot disks to pd-standard (see
above) or requesting a quota increase.
Persistent Disk Mounting Issues
kubectl get pods -n kube-system | grep csi
kubectl describe storageclass standard-rwo
Load Balancer Not Provisioned
kubectl describe svc lakehousecat-ui -n lakehousecat-prod
kubectl get pods -n kube-system | grep l7
autoscaling-metrics-adapter in CrashLoopBackOff in kube-system
This is a GKE-internal system component, not part of Lakehousecat. On some GKE minor versions (observed on 1.36) it's deployed without its matching CRD and restarts continuously:
failed to list *v1beta1.AutoscalingMetric: the server could not find the
requested resource (get autoscalingmetrics.autoscaling.gke.io)
It only serves custom/external metrics, which Lakehousecat doesn't use — Horizontal Pod Autoscaling on CPU/memory continues to work normally via the standard metrics-server. Safe to ignore.
Instance stuck in LicenseError after rebuilding the cluster
Not GKE-specific, but relevant for disaster recovery on any platform: the license handshake is bound to the cluster's identity. If you tear down and rebuild the cluster, the new cluster gets a new identity, and re-registering the same instance fails:
Handshake key already used by another cluster. Each instance can only be
registered on one cluster.
There is currently no self-service path to re-bind an existing instance to a rebuilt cluster — contact Lakehousecat support to have the registration released.
Next Steps
- Scaling - Configure autoscaling for production
- Monitoring - Review monitoring options
- Security - Review security configuration
Support
- GKE Documentation: https://cloud.google.com/kubernetes-engine/docs
- Lakehousecat Portal: https://portal.lakehousecat.com