Skip to main content
Version: Next

Deploy to Hetzner Cloud

Deploy Lakehousecat to a Kubernetes cluster on Hetzner Cloud.

Overview​

Terraform reference repository

A reference Terraform configuration for this platform is available as lhc-kube-hetzner in the lakehousecat GitHub organisation. It provisions the cluster layer only; the application rollout below is the same either way.

Deployment Status: ✅ Verified by Lakehousecat — Hetzner was the first platform Lakehousecat was deployed and tested on.

Hetzner has no managed Kubernetes

Unlike EKS, GKE and AKS, Hetzner Cloud does not offer a managed Kubernetes service. You run the cluster yourself. The reference setup uses kube-hetzner, a Terraform module that builds a k3s cluster on Hetzner Cloud servers with openSUSE MicroOS as the node operating system. MicroOS is an immutable, container-optimised OS and a fixed requirement of kube-hetzner; it cannot be replaced by Ubuntu.

Cluster administration

Hetzner-specific configuration — projects, firewalls, networks and API tokens — is the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.

Prerequisites​

Hetzner Setup​

  • A Hetzner Cloud account with a project for the cluster and an API token with read and write permission for that project
  • An SSH key pair for the cluster nodes

Required Tools​

  • Terraform: 1.10.1 or higher
  • Packer: 1.11.0 or higher (needed once per Hetzner project, see Step 1)
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+
brew install terraform packer kubectl helm

Step 1: Create the MicroOS Snapshots (once per project)​

kube-hetzner boots its nodes from two MicroOS snapshots (x86 and ARM) that must exist in your Hetzner project before the first terraform apply. It finds them through the label microos-snapshot = yes. You build them once per Hetzner project; they persist and are reused by every later cluster deployment or rebuild.

export HCLOUD_TOKEN=<token for the target project>

packer init .
packer build hcloud-microos-snapshots.pkr.hcl

The build takes about ten minutes. Packer creates two temporary servers, installs MicroOS on them, takes a snapshot of each and deletes the servers again. Afterwards the project contains two snapshots:

  • OpenSUSE MicroOS x86 by Kube-Hetzner
  • OpenSUSE MicroOS ARM by Kube-Hetzner

If you ever delete the snapshots, run the build again before the next terraform apply. The Packer template ships with the reference repository.

Step 2: Provision the Cluster​

export HCLOUD_TOKEN=<token for the target project>
export TF_VAR_hcloud_token=$HCLOUD_TOKEN
export TF_VAR_ssh_public_key="$(cat ~/.ssh/<your-key>.pub)"

terraform init
terraform plan -out tf_plan
terraform apply tf_plan

Export the kubeconfig and verify access:

terraform output -raw kubeconfig > ~/.kube/lhc-hetzner.yaml
export KUBECONFIG=~/.kube/lhc-hetzner.yaml

kubectl get nodes

All nodes should report Ready.

The kube-hetzner module is configured with the following settings, which the rest of this guide relies on:

SettingValueWhy
ingress_controller"nginx"NGINX Ingress is installed by the module and exposed through a Hetzner Load Balancer
enable_cert_managertruecert-manager is installed for Let's Encrypt certificates
allow_scheduling_on_control_planetrueLets small clusters use the control plane nodes for workloads
automatically_upgrade_k3s / automatically_upgrade_osfalseUpgrades happen when you decide, not unattended
Load Balancerlb11 in the same location as the nodesEntry point for ingress traffic

Node sizes​

Size the node pool to the resource requirements of the Kubernetes Cluster guide and to your load and storage plan. Hetzner offers ARM (cax) and x86 (cx, cpx) server types; Lakehousecat runs on both. The architecture field of the Lakehousecat resource (Step 5) must match the architecture of your nodes.

One location per cluster

Hetzner volumes are bound to a single location. Keep all nodes of a cluster in the same location (for example nbg1) so that a rescheduled pod can always reach its volume.

NetworkPolicy enforcement​

Check that NetworkPolicies are enforced

The Operator creates about 15 NetworkPolicy objects. A NetworkPolicy is only an object in the Kubernetes API: if the cluster's network plugin (CNI) does not enforce it, it is accepted and silently does nothing. The kube-hetzner module's default network plugin is Flannel, which does not enforce NetworkPolicies on its own.

Configure the module with a network plugin that does — Cilium or Calico — by setting cni_plugin in the module, and verify enforcement after the Lakehousecat instance is running:

kubectl get lakehousecat -A
# The NP-Enforced column reads True once the Operator's own detector confirms enforcement

If NP-Enforced shows False, every NetworkPolicy in the cluster is inert, whatever networkPolicy.enabled says in the Operator's Helm values. See Network Policies.

Step 3: Verify Storage and Ingress​

kube-hetzner installs the Hetzner CSI driver, which provides the default storage class hcloud-volumes:

kubectl get storageclass

Use that class for the persistent volumes of the instance (Step 5). Hetzner volumes have a minimum size of 10 GiB and live in one location, see above.

The NGINX ingress controller and cert-manager were installed by the module. Check that the ingress controller has an external address:

kubectl get svc -A | grep -i ingress

The EXTERNAL-IP is the address of the Hetzner Load Balancer. Point a wildcard DNS record for your base domain (for example *.lakehousecat.company.com) at it.

Let's Encrypt ClusterIssuer​

Create a ClusterIssuer that answers the HTTP-01 challenge through the NGINX ingress class. The Operator uses the issuer letsencrypt-prod by default:

clusterissuer.yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: admin@company.com
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
ingressClassName: nginx
kubectl apply -f clusterissuer.yaml
kubectl describe clusterissuer letsencrypt-prod
Use the staging endpoint while testing

https://acme-staging-v02.api.letsencrypt.org/directory has no rate limits but issues certificates browsers do not trust. Use it for test setups and switch to the production endpoint afterwards.

Step 4: Install the Lakehousecat Operator​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lakehousecat-prod

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait

Create the handshake secret with the key from the portal:

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 5: Deploy the Lakehousecat Instance​

lakehousecat-hetzner.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret" # pragma: allowlist secret

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "arm64" # or "amd64" — must match your nodes

# Pin the ingress mode: the Operator's automatic detection is not needed here
ingress:
mode: "nginx"

domain:
baseDomain: "lakehousecat.company.com"
tls:
enabled: true
letsEncrypt: true

postgresql:
persistenceSize: "100Gi"
storageClass: "hcloud-volumes"

clickhouse:
persistence:
size: "500Gi"
storageClass: "hcloud-volumes"

valkey:
persistence:
size: "20Gi"
storageClass: "hcloud-volumes"

# In-cluster SeaweedFS (default). See "Configure Storage" in the Kubernetes Cluster guide
# for external S3.
objectStorage:
type: seaweedfs
volume:
persistence:
size: "100Gi"
storageClass: "hcloud-volumes"
kubectl apply -f lakehousecat-hetzner.yaml
kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w

Adjust sizes and replica counts to your load plan; see Scaling.

Troubleshooting​

Terraform fails because the MicroOS snapshots are missing​

terraform apply cannot find a boot image if the project has no snapshots labelled microos-snapshot = yes. Run the Packer build from Step 1 in that project.

Load Balancer has no external address​

kubectl get svc -A | grep -i ingress   # then: kubectl describe svc <name> -n <namespace>

Check that the Load Balancer type and location in the module configuration exist in your Hetzner project and that your project's Load Balancer quota is not exhausted.

Pods stay Pending after a node was replaced​

Hetzner volumes are bound to one location. If a pod is scheduled onto a node in another location it cannot attach its volume. Keep the whole cluster in one location.

Instance stuck in LicenseError after rebuilding the cluster​

The license handshake is bound to the cluster's identity. After tearing down and rebuilding the cluster, re-registering the same instance fails with Handshake key already used by another cluster. Contact Lakehousecat support to have the registration released.

Tearing Down a Test Cluster​

terraform destroy removes only resources tracked in its own Terraform state. After every destroy, check that terraform state list is empty and that the Hetzner Cloud Console shows no leftover servers, load balancers, floating IPs or firewalls in the project. The MicroOS snapshots are expected to remain. Do not delete leftover resources in the console before running terraform plan: deleting them first can leave stale state entries that break the next apply.

Next Steps​

  • Scaling - Configure autoscaling for production
  • Monitoring - Review monitoring options
  • Security - Review security configuration

Support​