Skip to main content
Version: Next

Scaling the Lakehousecat Framework

The Lakehousecat framework is designed to scale efficiently across different workloads and user demands through Kubernetes Operator-based configuration. This section provides administrators with comprehensive guidance on scaling various services within the system to optimize performance, handle increased user loads, and manage resource allocation effectively.

Overview​

Scaling in Lakehousecat is managed declaratively through the Lakehousecat Operator. Administrators define desired resource allocations and replica counts in the Lakehousecat Custom Resource (CR) YAML configuration, and the operator automatically reconciles the actual state to match the desired state.

Configuration-Based Scaling

Unlike traditional UI-based scaling, Lakehousecat uses GitOps-style configuration management. All scaling changes are made by modifying the Lakehousecat CR YAML file and applying it to the Kubernetes cluster.

How Scaling Works​

The Lakehousecat Operator continuously monitors the cluster state and compares it with the desired configuration:

  1. Administrator modifies the Lakehousecat CR YAML file (e.g., increase replicaCount from 1 to 3)
  2. Operator detects the configuration change
  3. Reconciliation process updates the corresponding Kubernetes Deployments
  4. Kubernetes creates or removes pods to match the desired replica count
  5. Monitoring tracks the impact of scaling changes
# Example: Scaling the Analytics service
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: my-lakehousecat
spec:
analytics:
replicaCount: 3 # Scale from 1 to 3 replicas
resources:
requests:
cpu: "500m"
memory: "2Gi"
limits:
cpu: "2"
memory: "6Gi"

Service Categories​

The Lakehousecat framework consists of several service categories, each configurable independently:

Application Services​

  • Analytics (analytics): Data analysis and query execution engine
  • Audio (audio): Audio processing and transcription services
  • LHC (lhc): Core Lakehousecat backend API
  • LLM (llm): Large Language Model integration service
  • Semantic (semantic): Semantic search and embedding services
  • UI (ui): Web application frontend

Workflow Orchestration (Apache Airflow)​

  • API Server (airflow.api-server): Airflow webserver and API
  • Scheduler (airflow.scheduler): DAG scheduling engine
  • Worker (airflow.worker): Task execution workers
  • Triggerer (airflow.triggerer): Trigger-based task execution
  • DAG Processor (airflow.dagProcessor): DAG parsing and processing

Data Storage Services​

  • PostgreSQL (postgresql): Primary relational database
  • ClickHouse (clickhouse): Columnar analytics database
  • SeaweedFS (seaweedfs): S3-compatible object storage
  • Valkey (valkey): In-memory caching and pub/sub

Business Intelligence​

  • Apache Superset (superset): Data visualization and dashboards
  • Celery Worker (superset.celeryWorker): Async query execution

Observability​

Lakehousecat does not deploy a monitoring stack — there is nothing to scale here. Each backend service exposes a /metrics endpoint (Prometheus exposition format, toggled via monitoring.metricsEnabled) that the customer's existing monitoring tooling scrapes.

Scaling Configuration​

Horizontal Scaling (Replicas)​

Adjust the number of pod replicas for a service:

spec:
analytics:
replicaCount: 3 # Run 3 parallel instances

Best for:

  • Stateless services (analytics, llm, ui, audio)
  • High availability requirements
  • Load distribution across multiple instances

Not recommended for:

  • Stateful services with single-writer constraints (clickhouse leader, postgresql primary)

Vertical Scaling (Resources)​

Adjust CPU and memory allocations:

spec:
analytics:
resources:
requests:
cpu: "1" # Guaranteed CPU cores
memory: "4Gi" # Guaranteed memory
limits:
cpu: "4" # Maximum CPU cores
memory: "16Gi" # Maximum memory

Resource Types:

  • requests: Guaranteed resources for scheduling and baseline performance
  • limits: Maximum resources the pod can consume

Best for:

  • Resource-intensive services (clickhouse, superset, analytics)
  • Services with variable workloads
  • Memory-bound operations (semantic embeddings, large queries)

Autoscaling (HPA)​

Enable Horizontal Pod Autoscaling based on CPU/memory metrics:

spec:
analytics:
replicaCount: 2 # Initial replicas
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: 80

How it works:

  • Kubernetes HPA monitors CPU and memory utilization
  • When utilization exceeds targets, HPA increases replica count
  • When utilization drops, HPA decreases replica count (respecting minReplicas)

Best for:

  • Services with variable traffic patterns
  • Cost optimization (scale down during low usage)
  • Predictable scaling behavior

API Ingress Rate Limiting​

Scaling replicas and resources controls how much load your services can handle. Separately, the API ingress controls how much load is let through per client — this is a distinct setting, not something that grows automatically when you add replicas or upgrade your subscription tier.

By default, the nginx ingress in front of the /api/v1/... routes applies:

SettingDefaultMeaning
Requests per second60Per client IP address
Concurrent connections20Per client IP address
Burst multiplier3Allows short bursts up to 3× the requests-per-second value (180 req/s) before throttling

These are DoS-protection defaults, sized for a single client IP — not a per-tenant or per-instance budget. If most of your users reach the instance from distinct IP addresses (typical for individual remote workers), the default rarely matters: each user has their own 20-connection allowance.

It matters when many real users share one apparent source IP — most commonly a corporate NAT gateway or a site-to-site VPN bridge into the cluster's network. From the ingress's perspective, all of those users are a single client, and they compete for the same 20 connections and 60 requests/second.

Configuring spec.ingress.rateLimit​

Override the defaults, or exempt specific source ranges entirely, via the Lakehousecat CR:

spec:
ingress:
rateLimit:
requestsPerSecond: 120
connections: 50
burstMultiplier: 3
exemptCIDRs:
- "203.0.113.0/24" # e.g. your office VPN egress range
  • requestsPerSecond, connections, burstMultiplier raise the per-IP ceiling for every client — use this if your deployment expects legitimately high per-IP traffic.
  • exemptCIDRs is the more targeted fix for the NAT/VPN case above: source ranges listed here bypass rate limiting entirely, while every other IP keeps the DoS protection.

All fields are optional; omitting rateLimit entirely keeps the defaults in the table above.

nginx ingress only

This applies to deployments using the default nginx ingress (spec.ingress.mode: "nginx"). In ALB mode (AWS EKS), these annotations have no equivalent — see Advanced: AWS ALB Ingress for how rate limiting works (or doesn't) there.

When to Scale​

Consider scaling Lakehousecat services when you observe:

Performance Indicators​

  • High CPU Usage (>80% sustained): Consider vertical scaling or adding replicas
  • Memory Pressure: OOMKilled pods, frequent restarts → Increase memory limits
  • High Latency: Slow query responses → Scale analytics/superset workers
  • Queue Buildup: Airflow task queues growing → Scale workers

Capacity Planning​

  • User Growth: Increasing concurrent users → Scale ui, analytics, llm services
  • Data Volume Growth: Growing datasets → Scale clickhouse, postgresql
  • Workload Changes: New analytical workloads → Adjust airflow worker resources
  • Seasonal Peaks: Expected traffic increases → Pre-scale services

Monitoring Metrics​

# Check current resource usage
kubectl top pods -n <namespace>

# View pod scaling status
kubectl get hpa -n <namespace>

# Check deployment replica status
kubectl get deployments -n <namespace>

Scaling Examples​

Example 1: Scale Analytics for High Concurrency​

spec:
analytics:
replicaCount: 5 # Increase from 1 to 5
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "4"
memory: "16Gi"
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70

Use case: Handle 50+ concurrent analytical queries

Example 2: Scale Airflow Workers for Heavy ETL​

spec:
airflow:
worker:
replicaCount: 5 # Increase workers
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 15
targetCPUUtilizationPercentage: 75

Use case: Process large data ingestion pipelines

Example 3: Scale ClickHouse for Large Datasets​

spec:
clickhouse:
shards: 2 # Increase data sharding
replicaCount: 2 # 2 replicas per shard
resources:
requests:
cpu: "4"
memory: "32Gi"
limits:
cpu: "8"
memory: "64Gi"
persistence:
size: "500Gi" # Increase storage

Use case: Store and query TBs of analytical data

Best Practices​

Configuration Management​

# 1. Edit the Lakehousecat CR
kubectl edit lakehousecat <name> -n <namespace>

# OR apply from file
kubectl apply -f lakehousecat-scaled.yaml

# 2. Verify operator reconciliation
kubectl logs -n lakehousecat-operator-system \
deployment/lakehousecat-operator-controller-manager -f

# 3. Monitor deployment updates
kubectl get deployments -n <namespace> -w

# 4. Check pod status
kubectl get pods -n <namespace>

Gradual Scaling​

Don't scale directly from 1 to 20 replicas:

# Step 1: Scale to 3 replicas
replicaCount: 3

# Monitor for 15-30 minutes, observe metrics

# Step 2: Scale to 5 replicas
replicaCount: 5

# Continue gradual increases

Resource Right-Sizing​

Start with baseline resources and adjust based on actual usage:

# Initial configuration
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
cpu: "2"
memory: "4Gi"

# After monitoring, adjust to actual usage patterns
resources:
requests:
cpu: "1" # Observed avg usage: 0.8 cores
memory: "3Gi" # Observed avg usage: 2.5 GiB
limits:
cpu: "3"
memory: "8Gi"

Service Dependencies​

When scaling services, consider dependencies:

  • Scaling analytics → May need to scale clickhouse and postgresql
  • Scaling airflow.worker → Ensure postgresql and valkey can handle connections
  • Scaling superset → Scale superset.celeryWorker proportionally

Monitoring Scaling Changes​

Kubernetes Events​

# Watch deployment events
kubectl describe deployment <service-name> -n <namespace>

# View recent events
kubectl get events -n <namespace> --sort-by='.lastTimestamp'

Metrics​

# Resource usage
kubectl top pods -n <namespace>

# HPA status
kubectl get hpa -n <namespace>

# Deployment status
kubectl rollout status deployment/<service-name> -n <namespace>

Metrics in Your Monitoring Stack​

The /metrics endpoint on each backend service exposes per-route request counters, latency histograms, and process-level CPU/memory data. Use your existing dashboards (Grafana, Datadog, etc.) to track:

  • Pod CPU/memory usage trends (from cluster-level metrics)
  • Request rate and latency (from http_requests_total / http_request_duration_seconds on /metrics)
  • Database connection pools (instrument as custom counters if needed)

Troubleshooting​

Pods Not Scaling​

# Check operator logs
kubectl logs -n lakehousecat-operator-system \
deployment/lakehousecat-operator-controller-manager

# Verify CR configuration
kubectl get lakehousecat <name> -n <namespace> -o yaml

# Check deployment spec
kubectl get deployment <service-name> -n <namespace> -o yaml

OOMKilled Pods​

If pods are killed due to memory:

# Increase memory limits
resources:
limits:
memory: "8Gi" # Increase from 4Gi

Insufficient Resources​

If pods remain Pending:

# Check node resources
kubectl describe nodes

# Check pod events
kubectl describe pod <pod-name> -n <namespace>

Solution: Add more cluster nodes or reduce resource requests.

Next Steps​

  1. Review current resource usage: kubectl top pods -n <namespace>
  2. Identify services requiring scaling based on metrics
  3. Update Lakehousecat CR configuration
  4. Apply changes: kubectl apply -f lakehousecat.yaml
  5. Monitor the rollout and adjust as needed

For service-specific scaling guidance, refer to the individual service documentation pages.