Scaling the Lakehousecat Framework
The Lakehousecat framework is designed to scale efficiently across different workloads and user demands through Kubernetes Operator-based configuration. This section provides administrators with comprehensive guidance on scaling various services within the system to optimize performance, handle increased user loads, and manage resource allocation effectively.
Overview
Scaling in Lakehousecat is managed declaratively through the Lakehousecat Operator. Administrators define desired resource allocations and replica counts in the Lakehousecat Custom Resource (CR) YAML configuration, and the operator automatically reconciles the actual state to match the desired state.
Unlike traditional UI-based scaling, Lakehousecat uses GitOps-style configuration management. All scaling changes are made by modifying the Lakehousecat CR YAML file and applying it to the Kubernetes cluster.
How Scaling Works
The Lakehousecat Operator continuously monitors the cluster state and compares it with the desired configuration:
- Administrator modifies the Lakehousecat CR YAML file (e.g., increase
replicaCountfrom 1 to 3) - Operator detects the configuration change
- Reconciliation process updates the corresponding Kubernetes Deployments
- Kubernetes creates or removes pods to match the desired replica count
- Monitoring tracks the impact of scaling changes
# Example: Scaling the Analytics service
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: my-lakehousecat
spec:
analytics:
replicaCount: 3 # Scale from 1 to 3 replicas
resources:
requests:
cpu: "500m"
memory: "2Gi"
limits:
cpu: "2"
memory: "6Gi"
Service Categories
The Lakehousecat framework consists of several service categories, each configurable independently:
Application Services
- Analytics (
analytics): Data analysis and query execution engine - Audio (
audio): Audio processing and transcription services - LHC (
lhc): Core Lakehousecat backend API - LLM (
llm): Large Language Model integration service - Semantic (
semantic): Semantic search and embedding services - UI (
ui): Web application frontend
Workflow Orchestration (Apache Airflow)
- API Server (
airflow.api-server): Airflow webserver and API - Scheduler (
airflow.scheduler): DAG scheduling engine - Worker (
airflow.worker): Task execution workers - Triggerer (
airflow.triggerer): Trigger-based task execution - DAG Processor (
airflow.dagProcessor): DAG parsing and processing
Data Storage Services
- PostgreSQL (
postgresql): Primary relational database - ClickHouse (
clickhouse): Columnar analytics database - SeaweedFS (
seaweedfs): S3-compatible object storage - Valkey (
valkey): In-memory caching and pub/sub
Business Intelligence
- Apache Superset (
superset): Data visualization and dashboards - Celery Worker (
superset.celeryWorker): Async query execution
Observability
Lakehousecat does not deploy a monitoring stack — there is nothing to scale here. Each backend service exposes a /metrics endpoint (Prometheus exposition format, toggled via monitoring.metricsEnabled) that the customer's existing monitoring tooling scrapes.
Scaling Configuration
Horizontal Scaling (Replicas)
Adjust the number of pod replicas for a service:
spec:
analytics:
replicaCount: 3 # Run 3 parallel instances
Best for:
- Stateless services (analytics, llm, ui, audio)
- High availability requirements
- Load distribution across multiple instances
Not recommended for:
- Stateful services with single-writer constraints (clickhouse leader, postgresql primary)
Vertical Scaling (Resources)
Adjust CPU and memory allocations:
spec:
analytics:
resources:
requests:
cpu: "1" # Guaranteed CPU cores
memory: "4Gi" # Guaranteed memory
limits:
cpu: "4" # Maximum CPU cores
memory: "16Gi" # Maximum memory
Resource Types:
requests: Guaranteed resources for scheduling and baseline performancelimits: Maximum resources the pod can consume
Best for:
- Resource-intensive services (clickhouse, superset, analytics)
- Services with variable workloads
- Memory-bound operations (semantic embeddings, large queries)
Autoscaling (HPA)
Enable Horizontal Pod Autoscaling based on CPU/memory metrics:
spec:
analytics:
replicaCount: 2 # Initial replicas
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: 80
How it works:
- Kubernetes HPA monitors CPU and memory utilization
- When utilization exceeds targets, HPA increases replica count
- When utilization drops, HPA decreases replica count (respecting minReplicas)
Best for:
- Services with variable traffic patterns
- Cost optimization (scale down during low usage)
- Predictable scaling behavior
API Ingress Rate Limiting
Scaling replicas and resources controls how much load your services can handle. Separately, the API ingress controls how much load is let through per client — this is a distinct setting, not something that grows automatically when you add replicas or upgrade your subscription tier.
By default, the nginx ingress in front of the /api/v1/... routes applies:
| Setting | Default | Meaning |
|---|---|---|
| Requests per second | 60 | Per client IP address |
| Concurrent connections | 20 | Per client IP address |
| Burst multiplier | 3 | Allows short bursts up to 3× the requests-per-second value (180 req/s) before throttling |
These are DoS-protection defaults, sized for a single client IP — not a per-tenant or per-instance budget. If most of your users reach the instance from distinct IP addresses (typical for individual remote workers), the default rarely matters: each user has their own 20-connection allowance.
It matters when many real users share one apparent source IP — most commonly a corporate NAT gateway or a site-to-site VPN bridge into the cluster's network. From the ingress's perspective, all of those users are a single client, and they compete for the same 20 connections and 60 requests/second.
Configuring spec.ingress.rateLimit
Override the defaults, or exempt specific source ranges entirely, via the Lakehousecat CR:
spec:
ingress:
rateLimit:
requestsPerSecond: 120
connections: 50
burstMultiplier: 3
exemptCIDRs:
- "203.0.113.0/24" # e.g. your office VPN egress range
requestsPerSecond,connections,burstMultiplierraise the per-IP ceiling for every client — use this if your deployment expects legitimately high per-IP traffic.exemptCIDRsis the more targeted fix for the NAT/VPN case above: source ranges listed here bypass rate limiting entirely, while every other IP keeps the DoS protection.
All fields are optional; omitting rateLimit entirely keeps the defaults in the table above.
This applies to deployments using the default nginx ingress (spec.ingress.mode: "nginx").
In ALB mode (AWS EKS), these annotations have no equivalent — see
Advanced: AWS ALB Ingress
for how rate limiting works (or doesn't) there.
When to Scale
Consider scaling Lakehousecat services when you observe:
Performance Indicators
- High CPU Usage (>80% sustained): Consider vertical scaling or adding replicas
- Memory Pressure: OOMKilled pods, frequent restarts → Increase memory limits
- High Latency: Slow query responses → Scale analytics/superset workers
- Queue Buildup: Airflow task queues growing → Scale workers
Capacity Planning
- User Growth: Increasing concurrent users → Scale ui, analytics, llm services
- Data Volume Growth: Growing datasets → Scale clickhouse, postgresql
- Workload Changes: New analytical workloads → Adjust airflow worker resources
- Seasonal Peaks: Expected traffic increases → Pre-scale services
Monitoring Metrics
# Check current resource usage
kubectl top pods -n <namespace>
# View pod scaling status
kubectl get hpa -n <namespace>
# Check deployment replica status
kubectl get deployments -n <namespace>
Scaling Examples
Example 1: Scale Analytics for High Concurrency
spec:
analytics:
replicaCount: 5 # Increase from 1 to 5
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "4"
memory: "16Gi"
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
Use case: Handle 50+ concurrent analytical queries
Example 2: Scale Airflow Workers for Heavy ETL
spec:
airflow:
worker:
replicaCount: 5 # Increase workers
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 15
targetCPUUtilizationPercentage: 75
Use case: Process large data ingestion pipelines
Example 3: Scale ClickHouse for Large Datasets
spec:
clickhouse:
shards: 2 # Increase data sharding
replicaCount: 2 # 2 replicas per shard
resources:
requests:
cpu: "4"
memory: "32Gi"
limits:
cpu: "8"
memory: "64Gi"
persistence:
size: "500Gi" # Increase storage
Use case: Store and query TBs of analytical data
Best Practices
Configuration Management
# 1. Edit the Lakehousecat CR
kubectl edit lakehousecat <name> -n <namespace>
# OR apply from file
kubectl apply -f lakehousecat-scaled.yaml
# 2. Verify operator reconciliation
kubectl logs -n lakehousecat-operator-system \
deployment/lakehousecat-operator-controller-manager -f
# 3. Monitor deployment updates
kubectl get deployments -n <namespace> -w
# 4. Check pod status
kubectl get pods -n <namespace>
Gradual Scaling
Don't scale directly from 1 to 20 replicas:
# Step 1: Scale to 3 replicas
replicaCount: 3
# Monitor for 15-30 minutes, observe metrics
# Step 2: Scale to 5 replicas
replicaCount: 5
# Continue gradual increases
Resource Right-Sizing
Start with baseline resources and adjust based on actual usage:
# Initial configuration
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
cpu: "2"
memory: "4Gi"
# After monitoring, adjust to actual usage patterns
resources:
requests:
cpu: "1" # Observed avg usage: 0.8 cores
memory: "3Gi" # Observed avg usage: 2.5 GiB
limits:
cpu: "3"
memory: "8Gi"
Service Dependencies
When scaling services, consider dependencies:
- Scaling
analytics→ May need to scaleclickhouseandpostgresql - Scaling
airflow.worker→ Ensurepostgresqlandvalkeycan handle connections - Scaling
superset→ Scalesuperset.celeryWorkerproportionally
Monitoring Scaling Changes
Kubernetes Events
# Watch deployment events
kubectl describe deployment <service-name> -n <namespace>
# View recent events
kubectl get events -n <namespace> --sort-by='.lastTimestamp'
Metrics
# Resource usage
kubectl top pods -n <namespace>
# HPA status
kubectl get hpa -n <namespace>
# Deployment status
kubectl rollout status deployment/<service-name> -n <namespace>
Metrics in Your Monitoring Stack
The /metrics endpoint on each backend service exposes per-route request counters, latency histograms, and process-level CPU/memory data. Use your existing dashboards (Grafana, Datadog, etc.) to track:
- Pod CPU/memory usage trends (from cluster-level metrics)
- Request rate and latency (from
http_requests_total/http_request_duration_secondson/metrics) - Database connection pools (instrument as custom counters if needed)
Troubleshooting
Pods Not Scaling
# Check operator logs
kubectl logs -n lakehousecat-operator-system \
deployment/lakehousecat-operator-controller-manager
# Verify CR configuration
kubectl get lakehousecat <name> -n <namespace> -o yaml
# Check deployment spec
kubectl get deployment <service-name> -n <namespace> -o yaml
OOMKilled Pods
If pods are killed due to memory:
# Increase memory limits
resources:
limits:
memory: "8Gi" # Increase from 4Gi
Insufficient Resources
If pods remain Pending:
# Check node resources
kubectl describe nodes
# Check pod events
kubectl describe pod <pod-name> -n <namespace>
Solution: Add more cluster nodes or reduce resource requests.
Next Steps
- Review current resource usage:
kubectl top pods -n <namespace> - Identify services requiring scaling based on metrics
- Update Lakehousecat CR configuration
- Apply changes:
kubectl apply -f lakehousecat.yaml - Monitor the rollout and adjust as needed
For service-specific scaling guidance, refer to the individual service documentation pages.