Skip to main content
Version: Next

Deploy to On-Premise Kubernetes

Deploy Lakehousecat to a self-managed Kubernetes cluster in your own data center, private cloud, or bare-metal infrastructure.

Deployment Status: ✅ Compatible (Kubernetes-standard)

Scope of this guide

This guide covers Lakehousecat Operator deployment only. Setting up and managing an on-premise Kubernetes cluster — including nodes, networking, storage, and ingress — is the responsibility of the cluster administrator. Consult your Kubernetes distribution documentation for cluster setup details.

Prerequisites​

Kubernetes Cluster​

You need a running Kubernetes cluster (1.29+) with:

  • kubectl access configured and connected to the cluster
  • Helm version 3.8 or higher installed
  • A default StorageClass with dynamic volume provisioning (e.g., local-path, NFS, Ceph, Longhorn)
  • An Ingress Controller deployed (e.g., NGINX, Traefik) — required for external access
  • Outbound HTTPS access to license.lakehousecat.com, ghcr.io (public, no credentials required), and the official upstream registries of the third-party services the operator deploys (for image pulls)

Verify your cluster is ready:

# Check cluster connection
kubectl cluster-info

# Verify nodes are Ready
kubectl get nodes

# Verify a default storage class exists
kubectl get storageclass

Handshake Key​

Register at portal.lakehousecat.com (free) to obtain your handshake key.

Step 1: Install Lakehousecat Operator​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lhc-instance

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lhc-instance}" \
--wait

Verify:

kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
kubectl get crd | grep lakehousecat

Step 2: Create Handshake Secret​

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 3: Deploy Lakehousecat Instance​

Create a Custom Resource for your on-premise deployment:

lakehousecat-onprem.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lhc-instance"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "amd64" # Use "arm64" for ARM nodes

# In-cluster SeaweedFS for object storage (recommended for on-premise; this is also the
# default if the objectStorage block is omitted entirely). External S3 is also supported
# here — see "Configure Storage" in the Kubernetes Cluster guide — but requires outbound
# network access from the cluster to the S3 endpoint, which on-premise networks don't
# always have.
objectStorage:
type: seaweedfs
volume:
persistence:
size: "500Gi"
storageClass: "your-storage-class" # Replace with your StorageClass

# Domain and TLS
domain:
baseDomain: "lakehousecat.company.internal"
tls:
enabled: true
secretName: "lakehousecat-tls-cert" # Pre-create this secret (see below)

# Database storage
postgresql:
persistenceSize: "100Gi"
storageClass: "your-storage-class"

clickhouse:
persistence:
size: "500Gi"
storageClass: "your-storage-class"
arm64 excludes one datasource type

architecture: "arm64" excludes IBM Db2 as a datasource type — IBM does not ship a Linux/ARM64 driver for it. Every other datasource type is unaffected. See IBM Db2 — Prerequisites for details.

Apply:

kubectl apply -f lakehousecat-onprem.yaml

TLS Certificate​

Provide your own certificate (from your internal CA or self-signed):

kubectl create secret tls lakehousecat-tls-cert \
--cert=your-certificate.crt \
--key=your-private.key \
-n lhc-instance

Step 4: Monitor Deployment​

# Watch instance status
kubectl get lakehousecat -w

# Watch pods come up
kubectl get pods -n lhc-instance -w

# Operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

Initial deployment takes about 10 minutes while images are pulled and all services start. The exact duration depends on the network bandwidth between your cluster and the image registries.

Step 5: Configure DNS and Access​

Get the external IP assigned by your ingress controller or load balancer:

kubectl get svc -n ingress-nginx ingress-nginx-controller

Add a DNS record pointing your domain to that IP. For testing, add to /etc/hosts:

echo "192.168.1.x  lakehousecat.company.internal" | sudo tee -a /etc/hosts

Access Lakehousecat at https://lakehousecat.company.internal.

Retrieve Admin Credentials​

kubectl get secret lhc-admin-secret -n lhc-instance \
-o jsonpath='{.data.LHC_ADMIN_EMAIL}' | base64 -d && echo

kubectl get secret lhc-admin-secret -n lhc-instance \
-o jsonpath='{.data.LHC_ADMIN_PASSWORD}' | base64 -d && echo

Troubleshooting​

Pods Stuck in Pending​

# Check pod events for resource or storage issues
kubectl describe pod <pod-name> -n lhc-instance

# Check available node resources
kubectl describe nodes

Persistent Volume Issues​

# Check PVC status
kubectl get pvc -n lhc-instance

# Check storage class is available
kubectl get storageclass

License Validation Failures​

Ensure outbound HTTPS access to license.lakehousecat.com is allowed from your cluster nodes.

kubectl get secret lhc-handshake-secret -n lhc-operator -o yaml

Next Steps​

  • Scaling — Configure resource allocation and autoscaling
  • Monitoring — Connect your monitoring stack to the /metrics endpoints and log bucket
  • Backup — Configure data backups

Support​