Skip to main content
Version: Next

Deploy to AWS EKS

Deploy Lakehousecat to Amazon Elastic Kubernetes Service (EKS), AWS's managed Kubernetes service. This guide covers EKS-specific cluster provisioning and configuration, followed by the standard Kubernetes Operator deployment process.

Overview​

Terraform reference repository

A reference Terraform configuration for this platform is available as lhc-kube-aws in the lakehousecat GitHub organisation. It provisions the cluster layer only; the application rollout below is the same either way.

AWS EKS is the recommended platform for production Lakehousecat deployments on AWS, providing:

  • Managed Kubernetes control plane
  • S3 for object storage (logs, backups, and data archiving)
  • Auto-scaling and high availability
RDS not supported

Lakehousecat manages its own PostgreSQL and ClickHouse instances deployed by the Operator. AWS RDS is not supported as a database backend.

Cluster administration

EKS-specific configurations — IAM policies, networking, node AMI selection, and monitoring integrations (CloudWatch, Datadog, or others) — are the responsibility of the cluster administrator. This guide covers the cluster setup required to run the Lakehousecat Operator.

Deployment Status: ✅ Tested and fully supported

Deployment Approaches​

This guide provides two approaches for EKS cluster provisioning:

  • Use case: Production deployments, managed infrastructure
  • Pros: Versionable, reproducible, production-ready
  • Cons: Requires Terraform knowledge

Option 2: eksctl (CLI-based) — Evaluation Only​

  • Use case: Quick evaluation, testing, proof-of-concept
  • Pros: Fast setup, minimal configuration
  • Cons: Not infrastructure-as-code, not suitable for production
tip

Terraform is recommended for all production deployments. Use eksctl only for initial evaluation or proof-of-concept.

Prerequisites​

AWS Account Setup​

  • AWS account with appropriate permissions
  • AWS CLI installed and configured
  • IAM permissions for EKS cluster creation

Required Tools​

For Quick Start (eksctl):

  • AWS CLI: Version 2.x
  • eksctl: EKS cluster management tool
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+

For Production (Terraform):

  • AWS CLI: Version 2.x
  • Terraform: Version 1.5+
  • kubectl: Kubernetes command-line tool
  • Helm: Version 3.8+

Install Tools (macOS)​

# AWS CLI
brew install awscli

# eksctl
brew tap weaveworks/tap
brew install weaveworks/tap/eksctl

# kubectl
brew install kubectl

# Helm
brew install helm

Verify installations:

aws --version
eksctl version
kubectl version --client
helm version

Step 1: Provision EKS Cluster​

Choose your deployment approach based on your use case:

Use for: Production deployments, managed infrastructure, reproducible environments

Infrastructure as Code

Terraform provides Infrastructure as Code (IaC) for EKS clusters, making them versionable, reproducible, and manageable. This approach is tested and recommended for production Lakehousecat deployments.

Prerequisites​

# Install Terraform (macOS)
brew install terraform

# Verify installation
terraform version

Terraform Configuration​

Create the following files in a new directory:

main.tf - Main EKS cluster configuration:

terraform {
required_version = ">= 1.5"

required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
}

provider "aws" {
region = var.region
}

# VPC Module
module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 5.0"

name = "${var.cluster_name}-vpc"
cidr = "10.0.0.0/16"

azs = ["${var.region}a", "${var.region}b", "${var.region}c"]
private_subnets = ["10.0.1.0/24", "10.0.2.0/24", "10.0.3.0/24"]
public_subnets = ["10.0.101.0/24", "10.0.102.0/24", "10.0.103.0/24"]

enable_nat_gateway = true
enable_dns_hostnames = true
enable_dns_support = true

public_subnet_tags = {
"kubernetes.io/role/elb" = "1"
}

private_subnet_tags = {
"kubernetes.io/role/internal-elb" = "1"
}

tags = {
Environment = var.environment
ManagedBy = "Terraform"
Project = "Lakehousecat"
}
}

# EKS Cluster
module "eks" {
source = "terraform-aws-modules/eks/aws"
version = "~> 20.0"

cluster_name = var.cluster_name
cluster_version = "1.34"

# Network
vpc_id = module.vpc.vpc_id
subnet_ids = module.vpc.private_subnets
control_plane_subnet_ids = module.vpc.private_subnets

# Public access for easier management (adjust for security requirements)
cluster_endpoint_public_access = true

# Enable IRSA (IAM Roles for Service Accounts)
enable_irsa = true

# Cluster Addons
cluster_addons = {
coredns = {
most_recent = true
}
kube-proxy = {
most_recent = true
}
vpc-cni = {
most_recent = true
}
aws-ebs-csi-driver = {
most_recent = true
}
}

# Node Groups
eks_managed_node_groups = {
# System node group for critical workloads
system = {
name = "${var.cluster_name}-system"
use_name_prefix = false

instance_types = ["t3.small"]
capacity_type = "ON_DEMAND"

min_size = 1
max_size = 2
desired_size = 1

labels = {
role = "system"
namespace = "kube-system"
}

taints = [{
key = "CriticalAddonsOnly"
value = "true"
effect = "NO_SCHEDULE"
}]

tags = {
NodeGroup = "system"
}
}

# Workload node group with Spot instances (cost-optimized)
workload = {
name = "${var.cluster_name}-workload"
use_name_prefix = false

# Mixed instance types for better Spot availability
instance_types = ["t3.small", "t3a.small", "t3.medium"]
capacity_type = "SPOT"

min_size = 1
max_size = 10
desired_size = 3

labels = {
role = "workload"
}

tags = {
NodeGroup = "workload"
}
}
}

# Cluster access
enable_cluster_creator_admin_permissions = true

tags = {
Environment = var.environment
ManagedBy = "Terraform"
Project = "Lakehousecat"
}
}

# IAM Role for AWS Load Balancer Controller
module "aws_load_balancer_controller_irsa" {
source = "terraform-aws-modules/iam/aws//modules/iam-role-for-service-accounts-eks"
version = "~> 5.0"

role_name = "${var.cluster_name}-aws-load-balancer-controller"

attach_load_balancer_controller_policy = true

oidc_providers = {
main = {
provider_arn = module.eks.oidc_provider_arn
namespace_service_accounts = ["kube-system:aws-load-balancer-controller"]
}
}

tags = {
Environment = var.environment
ManagedBy = "Terraform"
}
}

# IAM Role for EBS CSI Driver
module "ebs_csi_driver_irsa" {
source = "terraform-aws-modules/iam/aws//modules/iam-role-for-service-accounts-eks"
version = "~> 5.0"

role_name = "${var.cluster_name}-ebs-csi-driver"

attach_ebs_csi_policy = true

oidc_providers = {
main = {
provider_arn = module.eks.oidc_provider_arn
namespace_service_accounts = ["kube-system:ebs-csi-controller-sa"]
}
}

tags = {
Environment = var.environment
ManagedBy = "Terraform"
}
}

variables.tf - Input variables:

variable "region" {
description = "AWS region"
type = string
default = "us-east-1"
}

variable "cluster_name" {
description = "EKS cluster name"
type = string
default = "lakehousecat-prod"
}

variable "environment" {
description = "Environment name"
type = string
default = "production"
}

outputs.tf - Cluster outputs:

output "cluster_endpoint" {
description = "EKS cluster endpoint"
value = module.eks.cluster_endpoint
}

output "cluster_name" {
description = "EKS cluster name"
value = module.eks.cluster_name
}

output "cluster_certificate_authority_data" {
description = "Base64 encoded certificate data"
value = module.eks.cluster_certificate_authority_data
sensitive = true
}

output "aws_load_balancer_controller_role_arn" {
description = "IAM role ARN for AWS Load Balancer Controller"
value = module.aws_load_balancer_controller_irsa.iam_role_arn
}

output "ebs_csi_driver_role_arn" {
description = "IAM role ARN for EBS CSI Driver"
value = module.ebs_csi_driver_irsa.iam_role_arn
}

output "configure_kubectl" {
description = "Command to configure kubectl"
value = "aws eks update-kubeconfig --region ${var.region} --name ${module.eks.cluster_name}"
}

Deploy EKS Cluster​

# Initialize Terraform
terraform init

# Review planned changes
terraform plan

# Create cluster (takes 15-20 minutes)
terraform apply

Configure kubectl​

# Get kubeconfig command from Terraform output
terraform output -raw configure_kubectl | bash

# Or manually:
aws eks update-kubeconfig --region us-east-1 --name lakehousecat-prod

# Verify connection
kubectl cluster-info
kubectl get nodes

Key Features of This Setup​

  • ✅ Infrastructure as Code: Versionable, reproducible
  • ✅ Cost-Optimized: Spot instances for workload nodes (70% cost savings)
  • ✅ Separation of Concerns: System vs. workload node groups
  • ✅ IRSA Enabled: Secure IAM roles for service accounts
  • ✅ EBS CSI Driver: Included for persistent volumes
  • ✅ ALB Controller Ready: IAM role pre-configured
  • ✅ Production-Ready: Based on tested Lakehousecat deployments

Cost Optimization​

This Terraform setup uses Spot instances for the workload node group, resulting in approximately 70% cost savings compared to On-Demand instances. The system node group uses On-Demand instances to ensure critical workloads remain stable.

Estimated Monthly Cost (us-east-1):

  • System nodes (1x t3.small On-Demand): ~$15/month
  • Workload nodes (3x t3.small Spot): ~$15/month
  • EKS control plane: ~$75/month
  • Total: ~$105/month (vs. ~$300/month with On-Demand setup)

Option 2: Quick Start with eksctl (Evaluation Only)​

Use for: Evaluation, testing, proof-of-concept

Not for Production

This approach creates clusters imperatively and is not suitable for production. Use Terraform (Option 1) for production deployments.

Create an EKS cluster:

lakehousecat-cluster.yaml
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig

metadata:
name: lakehousecat-prod
region: us-east-1
version: "1.29"

iam:
withOIDC: true

managedNodeGroups:
- name: lakehousecat-nodes
instanceType: m6i.4xlarge # 16 vCPU, 64 GB RAM
desiredCapacity: 3
minSize: 3
maxSize: 10
volumeSize: 200
volumeType: gp3
privateNetworking: true
labels:
workload: lakehousecat

vpc:
cidr: 10.0.0.0/16
nat:
gateway: HighlyAvailable

Create the cluster:

eksctl create cluster -f lakehousecat-cluster.yaml
Cluster Creation Time

EKS cluster creation typically takes 15–20 minutes.

Verify:

aws eks update-kubeconfig --region us-east-1 --name lakehousecat-prod
kubectl cluster-info
kubectl get nodes

Step 2: Configure EKS-Specific Components​

Cluster administrator responsibility

The components in this section (Load Balancer Controller, EBS CSI Driver) are standard EKS cluster prerequisites. They are set up once by the cluster administrator and are not managed by the Lakehousecat Operator.

Install AWS Load Balancer Controller​

Required for ingress and LoadBalancer services:

# Create IAM policy
curl -o iam-policy.json https://raw.githubusercontent.com/kubernetes-sigs/aws-load-balancer-controller/v2.7.0/docs/install/iam_policy.json

aws iam create-policy \
--policy-name AWSLoadBalancerControllerIAMPolicy \
--policy-document file://iam-policy.json

# Create service account
eksctl create iamserviceaccount \
--cluster=lakehousecat-prod \
--namespace=kube-system \
--name=aws-load-balancer-controller \
--attach-policy-arn=arn:aws:iam::ACCOUNT_ID:policy/AWSLoadBalancerControllerIAMPolicy \
--override-existing-serviceaccounts \
--approve

# Install controller via Helm
helm repo add eks https://aws.github.io/eks-charts
helm repo update

helm install aws-load-balancer-controller eks/aws-load-balancer-controller \
-n kube-system \
--set clusterName=lakehousecat-prod \
--set serviceAccount.create=false \
--set serviceAccount.name=aws-load-balancer-controller

Configure EBS CSI Driver​

For persistent volume support:

# Create IAM policy for EBS CSI driver
aws iam create-policy \
--policy-name AmazonEKS_EBS_CSI_Driver_Policy \
--policy-document file://ebs-csi-policy.json

# Create service account
eksctl create iamserviceaccount \
--name ebs-csi-controller-sa \
--namespace kube-system \
--cluster lakehousecat-prod \
--attach-policy-arn arn:aws:iam::ACCOUNT_ID:policy/AmazonEKS_EBS_CSI_Driver_Policy \
--approve

# Install EBS CSI driver
kubectl apply -k "github.com/kubernetes-sigs/aws-ebs-csi-driver/deploy/kubernetes/overlays/stable/?ref=release-1.26"

Verify Storage Class​

Check that EBS storage class is available:

kubectl get storageclass

# Expected output:
# gp2 (default)
# gp3

Step 3: Install Lakehousecat Operator​

The operator installation is identical across all Kubernetes platforms - this is pure Kubernetes-standard configuration.

Create the Namespaces​

The operator gets rights only in the instance namespaces you list in targetNamespaces and does not create namespaces itself, so create the instance namespace before installing it:

kubectl create namespace lhc-operator
kubectl create namespace lakehousecat-prod

Install Operator via Helm​

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--set "targetNamespaces={lakehousecat-prod}" \
--wait

Verify Operator Installation​

# Check operator pod
kubectl get pods -n lhc-operator

# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

# Verify CRD installation
kubectl get crd | grep lakehousecat

Step 4: Create Handshake Secret​

Get your handshake key from portal.lakehousecat.com, then create the Kubernetes secret:

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Verify:

kubectl get secret lhc-handshake-secret -n lhc-operator

Step 5: Configure S3 Storage​

Create S3 Bucket​

# Create bucket for Lakehousecat data
aws s3 mb s3://my-company-lakehousecat-data --region us-east-1

# Enable versioning (recommended)
aws s3api put-bucket-versioning \
--bucket my-company-lakehousecat-data \
--versioning-configuration Status=Enabled

# Enable encryption
aws s3api put-bucket-encryption \
--bucket my-company-lakehousecat-data \
--server-side-encryption-configuration '{
"Rules": [{
"ApplyServerSideEncryptionByDefault": {
"SSEAlgorithm": "AES256"
}
}]
}'

Create IAM User for S3 Access​

# Create IAM user
aws iam create-user --user-name lakehousecat-s3-user

# Attach S3 policy
aws iam put-user-policy \
--user-name lakehousecat-s3-user \
--policy-name LakehousecatS3Access \
--policy-document '{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": [
"s3:ListBucket",
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject"
],
"Resource": [
"arn:aws:s3:::my-company-lakehousecat-data",
"arn:aws:s3:::my-company-lakehousecat-data/*"
]
}]
}'

# Create access keys
aws iam create-access-key --user-name lakehousecat-s3-user

Store the AccessKeyId and SecretAccessKey in a Kubernetes secret — they are never put in the CR. Create the secret in the instance namespace:

kubectl create secret generic lhc-storage-s3-secret \
-n lakehousecat-prod \
--from-literal=access-key=AKIA... \
--from-literal=secret-key=...

See Required Secrets for details.

Step 6: Deploy Lakehousecat Instance​

Create a Lakehousecat Custom Resource with EKS-optimized configuration:

lakehousecat-eks-production.yaml
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: production
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lakehousecat-prod"
adminemail: "admin@company.com"

version: "0.0.42"
architecture: "amd64"

# S3 storage — credentials via the pre-created lhc-storage-s3-secret (never in the CR)
objectStorage:
type: s3
region: "us-east-1"
bucketPrefix: "my-company-lakehousecat"
credentialsSecretName: "lhc-storage-s3-secret" # keys: access-key, secret-key

# Domain with AWS ALB
domain:
baseDomain: "lakehousecat.company.com"
tls:
enabled: true
letsEncrypt: true
email: "admin@company.com"

# Production resource allocations
analytics:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

lhc:
replicaCount: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "2"
memory: "8Gi"

ui:
replicaCount: 2
resources:
requests:
cpu: "500m"
memory: "1Gi"

# EBS-backed persistent storage (using gp3)
postgresql:
persistenceSize: "100Gi"
storageClass: "gp3"
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"

clickhouse:
shards: 2
replicaCount: 2
persistence:
size: "500Gi"
storageClass: "gp3"
resources:
requests:
cpu: "4"
memory: "32Gi"
limits:
cpu: "8"
memory: "64Gi"

valkey:
persistence:
size: "20Gi"
storageClass: "gp3"

# Airflow workers
airflow:
worker:
replicaCount: 3
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 15
resources:
requests:
cpu: "2"
memory: "8Gi"

# Monitoring — exposes /metrics endpoints on backend services
# for scraping by the customer's monitoring stack (e.g. AWS Managed Prometheus).
# The Operator does NOT deploy Prometheus or Grafana.
monitoring:
metricsEnabled: true

Apply the configuration:

kubectl apply -f lakehousecat-eks-production.yaml

Step 7: Monitor Deployment​

Watch the deployment progress:

# Check instance status
kubectl get lakehousecat -w

# Check all pods
kubectl get pods -n lakehousecat-prod -w

# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

# Describe instance for detailed status
kubectl describe lakehousecat production

Step 8: Configure DNS and Access​

Get Load Balancer DNS​

# Get ALB hostname
kubectl get svc -n lakehousecat-prod lakehousecat-ui

# Output:
# NAME TYPE EXTERNAL-IP
# lakehousecat-ui LoadBalancer a1234...us-east-1.elb.amazonaws.com

Configure Route 53​

The default setup in this guide uses ingress-nginx behind a Network Load Balancer (NLB), whose hostname resolves to fixed IPs — a plain CNAME record works for it. This is different from an AWS Application Load Balancer (ALB), covered under Advanced: AWS ALB Ingress below, whose IPs are not stable and must never be targeted with an A record.

Create a CNAME record in Route 53:

# Using AWS CLI
aws route53 change-resource-record-sets \
--hosted-zone-id Z1234567890ABC \
--change-batch '{
"Changes": [{
"Action": "CREATE",
"ResourceRecordSet": {
"Name": "lakehousecat.company.com",
"Type": "CNAME",
"TTL": 300,
"ResourceRecords": [{
"Value": "a1234...us-east-1.elb.amazonaws.com"
}]
}
}]
}'

Or use the AWS Console: Route 53 → Hosted Zones → Create Record

Targeting the zone apex (e.g. company.com without a subdomain)

A CNAME record is not allowed on a zone apex. Use a Route 53 Alias record instead (select "Alias" when creating the record and point it at the load balancer), or use a subdomain like lakehousecat.company.com, which allows a normal CNAME.

Access Lakehousecat​

Once DNS propagates (a few minutes):

https://lakehousecat.company.com

Troubleshooting​

Pods Pending (No Nodes Available)​

# Check node status
kubectl get nodes

# Check cluster autoscaler logs
kubectl logs -f deployment/cluster-autoscaler -n kube-system

# Manually scale node group
eksctl scale nodegroup \
--cluster lakehousecat-prod \
--name lakehousecat-nodes \
--nodes 5

EBS Volume Mounting Failures​

# Check EBS CSI driver
kubectl get pods -n kube-system | grep ebs-csi

# Verify IAM permissions
aws iam get-role-policy \
--role-name AmazonEKS_EBS_CSI_DriverRole \
--policy-name AmazonEKS_EBS_CSI_Driver_Policy

Load Balancer Not Created​

# Check AWS Load Balancer Controller
kubectl logs -f deployment/aws-load-balancer-controller -n kube-system

# Verify service annotations
kubectl describe svc lakehousecat-ui -n lakehousecat-prod

Certificate stuck pending after setting the DNS record late​

If you applied the CR before creating the DNS record, the domain may resolve publicly while cert-manager still fails with no such host for up to 30 minutes. This is CoreDNS caching the earlier negative (NXDOMAIN) lookup. Force a recheck:

kubectl -n kube-system rollout restart deployment coredns

See Domain and TLS in the deployment overview for the full explanation.

Advanced: AWS ALB Ingress​

Not the recommended path

nginx (the default in this guide) remains the recommended ingress for Lakehousecat on EKS. Only choose ALB mode (spec.ingress.mode: "alb") if you already run AWS WAF, or your organization mandates ALB. The nginx path enforces API rate limits via nginx.ingress.kubernetes.io/limit-* annotations; ALB has no equivalent, and ValidateIngressConfiguration surfaces a warning condition if you enable ALB mode without attaching a WAFv2 ACL:

IngressReady | True | DeployedWithWarnings
ALB mode without spec.ingress.aws.wafv2ARN: the /api/v1 routes are not rate-limited.

If you see this condition, it is not an error — it's confirming that the API is running unthrottled unless you've placed your own WAF in front of it.

NLB and ALB are easy to confuse — and the wrong DNS record breaks the instance silently​

An EKS cluster running Lakehousecat can have two different load balancers active at once, and they require different DNS record types:

TypeCreated byAddress behaviorCorrect DNS record
NLBnetworkingress-nginx (cluster layer)Hostname resolving to fixed IPsCNAME/Alias, or a plain A record on the IPs
ALBapplicationAWS Load Balancer Controller, triggered by the Operator in ALB modeHostname whose IPs change over timeOnly CNAME/Alias — never an A record

Pointing an A record at an ALB's IPs will work for hours or days and then break without warning, because AWS rotates those IPs during normal operation. Because NLB IPs are stable, this mistake is easy to make if you're used to the default nginx setup.

Symptom if you get it wrong: the CR reaches Ready, every condition is True, every Ingress has an address, DNS resolves, and the certificate is valid — but the domain still serves an error page (/error?message=backend_config_load_failed → 404). The application itself is healthy; only the DNS target is wrong. This looks like a backend problem and sends troubleshooting in the wrong direction.

Finding the right target​

Don't guess — read it from the cluster:

kubectl get ingress -n <namespace> -o wide      # ADDRESS = the ALB hostname

The address only appears once the AWS Load Balancer Controller has finished building the target groups — an empty ADDRESS for up to five minutes after deploy is normal, not a fault.

For a Route 53 Alias record, the HostedZoneId you need is the ALB's own canonical hosted zone ID (e.g. Z32O12XQLNTSW2 for eu-west-1), not the ID of your own hosted zone — using the wrong one causes Route 53 to reject the record with an unhelpful error. Look it up with:

aws elbv2 describe-load-balancers --names <alb-name> \
--query 'LoadBalancers[0].{DNS:DNSName,ZoneId:CanonicalHostedZoneId}'

As with the NLB path above, a CNAME/Alias is required on a zone apex regardless — with ALB, there's no A-record fallback option at all.

Certificates in ALB mode come from ACM, not Let's Encrypt​

The Operator does not create a cert-manager Certificate in ALB mode, even if domain.tls.letsEncrypt: true is set in the CR — you provide an ACM certificate yourself via domain.tls.certificateARN. The certificate is bound to a specific AWS account and region and cannot be shared across accounts.

Next Steps​

  • Scaling - Configure autoscaling for production
  • Monitoring - Review Lakehousecat monitoring options
  • Security - Review security configuration for your Lakehousecat instance

Support​