Skip to main content
Version: Next

Deploy to Azure AKS

Deploy Lakehousecat to Azure Kubernetes Service (AKS), Microsoft Azure's managed Kubernetes service.

Overview​

Terraform reference repository

A reference Terraform configuration for this platform is available as lhc-kube-azure in the lakehousecat GitHub organisation. It provisions the cluster layer only; the application rollout below is the same either way.

Deployment Status: ⚠️ Not yet tested by Lakehousecat — community-contributed guide

Not Yet Tested

This environment has not yet been tested by Lakehousecat. The guide is provided as a reference based on standard Kubernetes deployment patterns. If you encounter environment-specific issues, please report them via portal.lakehousecat.com.

Azure Blob Storage not supported

Azure Blob Storage does not speak the S3 protocol and cannot be used as a Lakehousecat object storage backend. By default the Operator deploys SeaweedFS in-cluster (no configuration needed). Alternatively, an AKS cluster can use external S3 (e.g. AWS S3, or Google Cloud Storage via its S3-Interop API) as a hybrid setup — see Configure Storage. Azure SQL is not supported either — Lakehousecat manages its own PostgreSQL and ClickHouse instances.

Cluster administration

AKS-specific configurations — managed identity, networking, and monitoring integrations — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.

Prerequisites​

Azure Account Setup​

  • Microsoft Azure account with active subscription
  • Resource group for Lakehousecat resources
  • Appropriate Azure RBAC permissions for AKS creation

Required Tools​

  • Azure CLI (az): Azure command-line tool
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+

Install Tools (macOS)​

# Azure CLI
brew install azure-cli

# kubectl
brew install kubectl

# Helm
brew install helm

Verify installations:

az version
kubectl version --client
helm version

Authenticate with Azure​

# Login to Azure
az login

# Set default subscription
az account set --subscription YOUR_SUBSCRIPTION_ID

# Create resource group
az group create \
--name lakehousecat-rg \
--location eastus

Step 1: Provision AKS Cluster​

Create AKS Cluster​

Create an AKS cluster with production configuration:

az aks create \
--resource-group lakehousecat-rg \
--name lakehousecat-prod \
--location eastus \
--kubernetes-version 1.29 \
--node-count 3 \
--min-count 3 \
--max-count 10 \
--enable-cluster-autoscaler \
--node-vm-size Standard_D16s_v3 \
--node-osdisk-size 200 \
--network-plugin azure \
--enable-managed-identity \
--generate-ssh-keys \
--tags environment=production app=lakehousecat

VM Sizes:

  • Standard_D16s_v3: 16 vCPU, 64 GB RAM (production)
  • Standard_D8s_v3: 8 vCPU, 32 GB RAM (smaller workloads)
  • Standard_E16s_v3: 16 vCPU, 128 GB RAM (memory-intensive)
Cluster Creation Time

AKS cluster creation typically takes 5-10 minutes.

Using Azure Portal​

Alternatively, create via Azure Portal:

  1. Navigate to Kubernetes services → Create
  2. Configure basics: subscription, resource group, cluster name, region
  3. Set Kubernetes version (1.29+)
  4. Configure node pools: VM size, node count, autoscaling
  5. Review and create

Verify Cluster​

# Get cluster credentials
az aks get-credentials \
--resource-group lakehousecat-rg \
--name lakehousecat-prod

# Verify connection
kubectl cluster-info
kubectl get nodes

Step 2: Configure AKS-Specific Components​

Verify Storage Class​

AKS provides multiple storage classes:

kubectl get storageclass

# Expected output:
# default (Azure Disk - Standard HDD)
# managed-premium (Azure Disk - Premium SSD)
# azurefile (Azure Files)

Recommended for production: managed-premium (SSD-backed)

Ingress Controller​

An ingress controller is required to expose Lakehousecat via a domain. Install NGINX Ingress:

helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm repo update

helm install ingress-nginx ingress-nginx/ingress-nginx \
--namespace ingress-nginx \
--create-namespace \
--set controller.service.type=LoadBalancer \
--set controller.service.annotations."service\.beta\.kubernetes\.io/azure-load-balancer-health-probe-request-path"=/healthz

Step 3: Install Lakehousecat Operator​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

# Create the instance namespace first: the operator only gets rights in the
# namespaces listed in targetNamespaces and does not create them.
kubectl create namespace lakehousecat-prod

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait

Verify:

kubectl get pods -n lhc-operator
kubectl logs -f deployment/lhc-operator -n lhc-operator
kubectl get crd | grep lakehousecat

Step 4: Create Handshake Secret​

Get your handshake key from portal.lakehousecat.com, then create the Kubernetes secret:

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 5: Deploy Lakehousecat Instance​

Create a Lakehousecat Custom Resource with AKS-optimized configuration:

lakehousecat-aks.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "amd64"

# Premium SSD-backed persistent storage
postgresql:
persistenceSize: "100Gi"
storageClass: "managed-premium"

clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "managed-premium"

valkey:
persistence:
size: "20Gi"
storageClass: "managed-premium"

# In-cluster SeaweedFS (default). Omit this block entirely to use the same default,
# or replace it with `objectStorage: {type: s3, ...}` for external S3 — see
# "Configure Storage" in the Kubernetes Cluster guide.
objectStorage:
type: seaweedfs
volume:
persistence:
size: "100Gi"
storageClass: "managed-premium"

analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"

ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"

Apply:

kubectl apply -f lakehousecat-aks.yaml

Step 6: Monitor Deployment​

kubectl get lakehousecat -w
kubectl get pods -n lakehousecat-prod -w
kubectl logs -f deployment/lhc-operator -n lhc-operator

Step 7: Configure DNS and Access​

Get Load Balancer IP​

kubectl get svc -n lakehousecat-prod lakehousecat-ui

Configure Azure DNS​

# Create DNS zone (if not exists)
az network dns zone create \
--resource-group lakehousecat-rg \
--name company.com

# Add A record
az network dns record-set a add-record \
--resource-group lakehousecat-rg \
--zone-name company.com \
--record-set-name lakehousecat \
--ipv4-address <EXTERNAL-IP>

Troubleshooting​

Pods Pending (Insufficient Node Resources)​

# Check node pool status
az aks nodepool show \
--resource-group lakehousecat-rg \
--cluster-name lakehousecat-prod \
--name nodepool1

# Scale node pool
az aks nodepool scale \
--resource-group lakehousecat-rg \
--cluster-name lakehousecat-prod \
--name nodepool1 \
--node-count 5

Persistent Disk Mounting Issues​

kubectl describe storageclass managed-premium
kubectl get pods -n kube-system | grep csi

Load Balancer Not Provisioned​

kubectl describe svc lakehousecat-ui -n lakehousecat-prod
az network lb list --resource-group MC_lakehousecat-rg_lakehousecat-prod_eastus

Next Steps​

  • Scaling - Configure autoscaling for production
  • Monitoring - Review monitoring options
  • Security - Review security configuration

Support​