Skip to main content
Version: Next

Deploy to Google GKE

Deploy Lakehousecat to Google Kubernetes Engine (GKE), Google Cloud's managed Kubernetes service.

Overview​

Terraform reference repository

A reference Terraform configuration for this platform is available as lhc-kube-gcp in the lakehousecat GitHub organisation. It provisions the cluster layer only; the application rollout below is the same either way.

Deployment Status: ✅ Verified by Lakehousecat — full stack deployed and validated on GKE

Google Cloud Storage — supported only via its S3-Interop API

Lakehousecat does not have a native Google Cloud Storage integration. By default the Operator deploys SeaweedFS in-cluster (no configuration needed). Two external-storage alternatives are available as a hybrid setup: AWS S3 (fully supported), or Google Cloud Storage itself via its S3-Interop API (endpoint: https://storage.googleapis.com with GCS HMAC keys — optional, not yet verified end-to-end by Lakehousecat). See Configure Storage. Google Cloud SQL is not supported — Lakehousecat manages its own PostgreSQL and ClickHouse instances.

Cluster administration

GCP-specific configurations — IAM policies, networking, and monitoring integrations — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.

Prerequisites​

Google Cloud Setup​

  • Google Cloud Platform account with billing enabled
  • Appropriate IAM permissions for GKE cluster creation

Required Tools​

  • gcloud CLI: Google Cloud command-line tool
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+
  • gke-gcloud-auth-plugin: required by kubectl for every request against a GKE cluster (see Cluster Access below)

Install Tools (macOS)​

brew install google-cloud-sdk
gcloud components install kubectl
gcloud components install gke-gcloud-auth-plugin
brew install helm
gke-gcloud-auth-plugin install may report exit code 1

On Homebrew, gcloud components install gke-gcloud-auth-plugin can print "Update done!" but still exit with code 1, and the binary lands under /opt/homebrew/share/google-cloud-sdk/bin without a symlink into /opt/homebrew/bin. If kubectl reports the plugin as missing afterwards, add that directory to your PATH.

Authenticate with Google Cloud​

gcloud auth login
gcloud config set project YOUR_PROJECT_ID
gcloud config set compute/region us-central1

Cluster access: the auth plugin has two modes​

kubectl calls gke-gcloud-auth-plugin on every request against a GKE cluster. The plugin has two credential sources, controlled by a single flag:

Kubeconfig sourceargsCredentials from
gcloud container clusters get-credentials(none)the active gcloud account
A kubeconfig generated without that flag (e.g. via Terraform output)--use_application_default_credentialsGOOGLE_APPLICATION_CREDENTIALS

Without the flag, the plugin ignores Application Default Credentials entirely, even if GOOGLE_APPLICATION_CREDENTIALS is set and valid, and fails with:

failed to retrieve access token: … (gcloud.config.config-helper)
You do not currently have an active account selected.

The message says "no active account," which points at the wrong problem — the credentials file is fine, the plugin just isn't reading it. If you're generating kubeconfig outside of gcloud container clusters get-credentials (for example from automation), pass --use_application_default_credentials explicitly.

Check your SSD quota before creating the cluster​

A fresh GCP project defaults to 500 GB SSD_TOTAL_GB per region — a quota shared by pd-balanced and pd-ssd disks. Node boot disks are counted against this quota by default, and on a 3+ node cluster they can consume most of it before the application's own volumes are provisioned. Check the current quota first:

gcloud compute regions describe <region> --project <project> \
--format="flattened(quotas)" | grep -A2 SSD_TOTAL_GB

If it's tight, either request an increase under IAM & Admin → Quotas (usually approved within minutes with active billing, but not guaranteed), or — the more reliable fix — set node boot disks to pd-standard, which is billed against the much larger DISKS_TOTAL_GB quota instead. See Verify Storage Class below for why pd-standard boot disks don't affect application performance.

Step 1: Provision GKE Cluster​

gcloud compute addresses create lakehousecat-ingress --region us-central1

Reserving the address up front means it survives a cluster teardown and rebuild — in testing, the same IP was reused automatically on rebuild and DNS records didn't need to change. Without a reservation, a rebuild allocates a new ephemeral IP and DNS propagation delay follows.

gcloud container clusters create lakehousecat-prod \
--zone us-central1-a \
--cluster-version 1.29 \
--machine-type n2-standard-16 \
--num-nodes 3 \
--min-nodes 3 \
--max-nodes 10 \
--enable-autoscaling \
--enable-autorepair \
--disk-type pd-standard \
--disk-size 200 \
--default-max-pods-per-node 110 \
--enable-ip-alias \
--enable-dataplane-v2 \
--addons HorizontalPodAutoscaling,HttpLoadBalancing,GcePersistentDiskCsiDriver \
--labels environment=production,app=lakehousecat

Common instance types:

  • n2-standard-16: 16 vCPU, 64 GB RAM (production)
  • n2-standard-8: 8 vCPU, 32 GB RAM (smaller workloads)
Zonal, not regional

Use a zonal cluster (--zone), not a regional one (--region). Persistent Disks are zone-bound; a regional cluster spreads nodes across three zones, and a rescheduled pod can end up in a zone that doesn't hold its volume. A zonal cluster also fits the GKE free-tier allowance (one zonal cluster, ~$73/month cluster management fee).

Boot disk type vs. size

Boot disks use pd-standard here deliberately — see Check your SSD quota above. Don't shrink the boot disk size to save quota instead: GKE reserves roughly half of a node's boot disk as non-allocatable ephemeral storage (measured 43.8 GiB allocatable out of 94.3 GiB total). At the full Lakehousecat stack, per-node ephemeral usage reached 16–22.8 GiB — a 50 GB boot disk would leave too little headroom and risk ephemeral-storage evictions.

Pod CIDR sizing

GKE assigns each node a contiguous CIDR block for pods, not individual addresses. At the default of 110 pods/node that's a /24 per node — a /22 pod range only covers four nodes. Running out shows up as pods stuck in ContainerCreating with "failed to assign IP," which looks like a scheduling problem rather than an IP exhaustion issue. Size your pod IP range for the node count you expect to scale to, not just the starting count.

NetworkPolicy enforcement is on by default — don't turn it off

--enable-dataplane-v2 above isn't optional decoration: GKE Dataplane V2 is what actually enforces the NetworkPolicy objects the Operator creates (~15 of them). A NetworkPolicy is just a Kubernetes API object — without an enforcing CNI it is accepted and silently does nothing. Dataplane V2 is GKE's modern default for new clusters, but if you're customizing cluster creation, don't drop this flag or switch to the legacy --enable-network-policy Calico add-on path without understanding the tradeoff.

Verify enforcement is actually active, not just requested — check the cluster, not just the flag you passed:

kubectl get pods -n kube-system -l k8s-app=cilium
# anetd (Dataplane V2's Cilium-based dataplane) should show 1/1 or 2/2 Running on every node

kubectl get lakehousecat -A
# NP-Enforced column reads True once the Operator's own detector confirms enforcement

If NP-Enforced shows False, every NetworkPolicy in the cluster is inert regardless of what networkPolicy.enabled says in the Operator's Helm values — see Network Policies below.

Option 2: GKE Autopilot​

gcloud container clusters create-auto lakehousecat-autopilot \
--region us-central1 \
--cluster-version 1.29
Autopilot vs Standard

Autopilot: Google manages nodes, scaling, and security — per-pod billing. Standard: You manage nodes and instance types.

Verify Cluster​

gcloud container clusters get-credentials lakehousecat-prod --zone us-central1-a
kubectl cluster-info
kubectl get nodes

Step 2: Configure GKE-Specific Components​

Verify Storage Class​

GKE provides three storage classes out of the box:

kubectl get storageclass

Recommended and verified: standard-rwo — the CSI driver's SSD-backed (pd-balanced) class. Don't use plain standard: it's the older in-tree provisioner on pd-standard (HDD) — it works, but is noticeably slower under load, and the difference isn't obvious until you're running real workloads. premium-rwo (pd-ssd) is also available but not required for Lakehousecat's workload.

Ingress Controller​

An ingress controller is required to expose Lakehousecat via a domain. GKE's built-in HTTP Load Balancing addon (enabled above) can serve as the ingress controller. Alternatively, install NGINX Ingress:

helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update

helm install ingress-nginx ingress-nginx/ingress-nginx \
--namespace ingress-nginx \
--create-namespace \
--set controller.service.type=LoadBalancer

Step 3: Install Lakehousecat Operator​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lakehousecat-prod

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait

Verify:

kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 4: Create Handshake Secret​

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 5: Deploy Lakehousecat Instance​

lakehousecat-gke.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "amd64"

# SSD-backed persistent storage
postgresql:
persistenceSize: "100Gi"
storageClass: "standard-rwo"

clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "standard-rwo"

valkey:
persistence:
size: "20Gi"
storageClass: "standard-rwo"

# In-cluster SeaweedFS (default). Omit this block entirely to use the same default,
# or replace it with `objectStorage: {type: s3, ...}` for external S3 — see
# "Configure Storage" in the Kubernetes Cluster guide.
objectStorage:
type: seaweedfs
volume:
persistence:
size: "100Gi"
storageClass: "standard-rwo"

analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"

ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"

Apply:

kubectl apply -f lakehousecat-gke.yaml

Step 6: Monitor Deployment​

kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 7: Configure DNS and Access​

Get Load Balancer IP​

kubectl get svc -n lakehousecat-prod lakehousecat-ui

Configure Cloud DNS​

gcloud dns record-sets create lakehousecat.company.com. \
--zone=company-zone \
--type=A \
--ttl=300 \
--rrdatas=<EXTERNAL-IP>

Troubleshooting​

Pods Pending (Insufficient Resources)​

gcloud container clusters resize lakehousecat-prod \
--zone us-central1-a \
--num-nodes 5

Pods Pending — deployment stalls mid-stack with "context deadline exceeded"​

If clickhouse-shard0-0 or postgresql-1 stay Pending while everything else is running, and kubectl describe pod only shows a generic volume-binding timeout:

FailedScheduling: running PreBind plugin "VolumeBinding": binding volumes:
context deadline exceeded

that's a follow-on symptom, not the cause. Check the PVC's own events instead:

kubectl describe pvc <pvc-name> -n <namespace>
ProvisioningFailed … QUOTA_EXCEEDED: Quota 'SSD_TOTAL_GB' exceeded. Limit: 500.0 in region <region>

This is the SSD_TOTAL_GB project quota described under Check your SSD quota — it fails partway through the deployment, not at cluster creation, which is what makes it easy to mistake for a scheduling or Operator problem. Fix by switching node boot disks to pd-standard (see above) or requesting a quota increase.

Persistent Disk Mounting Issues​

kubectl get pods -n kube-system | grep csi
kubectl describe storageclass standard-rwo

Load Balancer Not Provisioned​

kubectl describe svc lakehousecat-ui -n lakehousecat-prod
kubectl get pods -n kube-system | grep l7

autoscaling-metrics-adapter in CrashLoopBackOff in kube-system​

This is a GKE-internal system component, not part of Lakehousecat. On some GKE minor versions (observed on 1.36) it's deployed without its matching CRD and restarts continuously:

failed to list *v1beta1.AutoscalingMetric: the server could not find the
requested resource (get autoscalingmetrics.autoscaling.gke.io)

It only serves custom/external metrics, which Lakehousecat doesn't use — Horizontal Pod Autoscaling on CPU/memory continues to work normally via the standard metrics-server. Safe to ignore.

Instance stuck in LicenseError after rebuilding the cluster​

Not GKE-specific, but relevant for disaster recovery on any platform: the license handshake is bound to the cluster's identity. If you tear down and rebuild the cluster, the new cluster gets a new identity, and re-registering the same instance fails:

Handshake key already used by another cluster. Each instance can only be
registered on one cluster.

There is currently no self-service path to re-bind an existing instance to a rebuilt cluster — contact Lakehousecat support to have the registration released.

Next Steps​

  • Scaling - Configure autoscaling for production
  • Monitoring - Review monitoring options
  • Security - Review security configuration

Support​