Skip to main content
Version: 0.0.41

Deployment Overview

Lakehousecat is deployed on Kubernetes using a Kubernetes Operator that manages the complete lifecycle of your data platform instances. This operator-based approach provides automated deployment, configuration management, scaling, and upgrades across any Kubernetes environment.

How Lakehousecat Deployment Works​

The Lakehousecat deployment consists of two main components:

1. Lakehousecat Operator (Cluster-Wide)​

The Operator is a Kubernetes controller that:

  • Manages Lakehousecat instances across multiple namespaces
  • Validates licenses via the Lakehousecat license service (license.lakehousecat.com)
  • Synchronizes container images from AWS ECR (public and private) to your cluster
  • Orchestrates 40+ microservices (Airflow, PostgreSQL, Valkey, ClickHouse, Superset, SeaweedFS, etc.)
  • Handles upgrades, scaling, and configuration changes

Install once per cluster - The operator can manage multiple Lakehousecat instances.

2. Lakehousecat Instance (Namespace-Specific)​

Each instance is defined as a Custom Resource (CR) that specifies:

  • License credentials (handshake key)
  • Version and architecture (amd64/arm64)
  • Storage configuration (S3, SeaweedFS, etc.)
  • Domain and TLS settings
  • Resource allocations and scaling parameters

Deploy multiple instances - Each instance runs in its own namespace with isolated resources.

Operator Permissions​

The Lakehousecat Operator follows the principle of least privilege. It does not require cluster-admin rights and does not use ClusterRole or ClusterRoleBinding resources. All RBAC is namespace-scoped.

The operator requires the following Kubernetes permissions:

Resource GroupResourcesAccess
lhc.lakehousecat.comlakehousecats, lakehousecats/status, lakehousecats/finalizersFull lifecycle management of the operator's own CRD
Core ("")namespacesget, list, watch, create — required to provision instance namespaces
Core ("")secrets, configmaps, services, persistentvolumeclaims, serviceaccounts, pods, pods/log, eventsManage instance resources within namespaces
appsdeployments, statefulsets, daemonsets, replicasetsDeploy and manage application workloads
batchjobs, cronjobsOrchestrate data ingestion workflows
networking.k8s.ioingresses, networkpoliciesConfigure ingress and network policies
rbac.authorization.k8s.ioroles, rolebindingsNamespace-scoped RBAC for instance service accounts
autoscalinghorizontalpodautoscalersEnable HPA-based autoscaling
policypoddisruptionbudgetsEnsure availability during cluster maintenance
coordination.k8s.ioleasesOperator leader election
discovery.k8s.ioendpointslicesService discovery
Namespace-scoped RBAC only

The operator uses Role and RoleBinding (namespace-scoped), not ClusterRole or ClusterRoleBinding. The only cluster-level resources accessed are namespaces and leases — required for namespace provisioning and leader election.

Prerequisites​

Before deploying Lakehousecat, ensure you have:

Required Software​

  • Kubernetes: Version 1.29 or higher
  • Helm: Version 3.8 or higher
  • kubectl: Configured and connected to your cluster
  • Handshake Key: Obtained from portal.lakehousecat.com after registration (free)

Cluster Resources (Minimum)​

For Single-User Evaluation:

  • 10 CPU cores (minimum)
  • 32 GB RAM (minimum)
  • 50 GB free disk space (SSD recommended)
  • Supported OS: macOS, Linux

For Production Multi-User Deployment:

  • 16+ CPU cores
  • 64+ GB RAM
  • 500+ GB persistent storage (scales with data volume)
  • Persistent volume provisioner (for database storage)

Network Requirements​

  • Outbound HTTPS access to:
    • license.lakehousecat.com (license validation)
    • AWS ECR (public and private) — container image pull; handled automatically by the operator
  • DNS resolution for internal Kubernetes service discovery
  • Ingress controller (required for external access via domain)

Deployment Options​

Choose the deployment option that best fits your needs:

Local Evaluation​

Best for: Single-user evaluation, development, testing

Two options are available:

  • TUI Installer — Recommended. A terminal application that automates the full setup: Minikube cluster, Helm repository, operator, and instance. No Kubernetes experience required.
  • Manual Setup — For advanced users who need full control over the configuration.

Resource Requirements: 32 GB RAM (minimum), 10 CPU cores, 50 GB free disk, macOS or Linux

Kubernetes Cluster​

Best for: Production deployments, multi-user environments, scalability

Deploy Lakehousecat on managed or self-hosted Kubernetes clusters:

  • AWS EKS (Elastic Kubernetes Service)
  • Google GKE (Google Kubernetes Engine)
  • Azure AKS (Azure Kubernetes Service)
  • On-Premise - Self-managed Kubernetes clusters

→ Kubernetes Cluster Guide

Resource Requirements: 16+ CPU cores, 64+ GB RAM, 500+ GB storage

Deployment Process Overview​

Regardless of which option you choose, the deployment follows the same basic steps:

Step 1: Prepare Kubernetes Cluster​

Set up your Kubernetes environment (Minikube, EKS, GKE, AKS, or on-premise).

Step 2: Install Lakehousecat Operator​

Install the operator using Helm:

helm repo add lakehousecat https://charts.lakehousecat.com
helm repo update

helm install lhc-operator lakehousecat/lakehousecat-operator \
--namespace lhc-operator \
--create-namespace \
--wait

Step 3: Configure Handshake Secret​

Create a secret with your handshake key from the portal:

kubectl create secret generic lhc-handshake-secret \
--from-literal=handshake-key=YOUR_HANDSHAKE_KEY \
-n lhc-operator

Step 4: Deploy Lakehousecat Instance​

Create a Lakehousecat Custom Resource (CR):

apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
metadata:
name: my-instance
spec:
license:
handshakeSecretName: "lhc-handshake-secret"

customer:
namespace: "lhc-instance"
adminemail: "admin@company.com"

version: "0.0.40"
architecture: "amd64" # amd64 | arm64

storage:
bucketName: "my-data-bucket"
region: "us-east-1"
accessKeyID: "AKIA..."
secretAccessKey: "..."

Apply the configuration:

kubectl apply -f lakehousecat-instance.yaml

Step 5: Verify Deployment​

Monitor the deployment:

# Check status of your specific instance (replace my-instance and lhc-instance with your values)
kubectl get lakehousecat my-instance -n lhc-instance

# View operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

# Check pods in instance namespace
kubectl get pods -n lhc-instance

Verifying a Successful Deployment​

Check all instances across all namespaces:

kubectl get lakehousecats -A

Phase: Ready means the Operator has reconciled all components and all pods are online and healthy. If the phase shows Pending, the Operator is still reconciling — check the CR description for details:

kubectl describe lakehousecat <instance-name> -n <namespace>

Check individual pod status:

kubectl -n <instance-namespace> get pods

All pods should show Running status with all containers ready. The full set of expected pods in a healthy deployment:

PodComponent
airflow-api-serverAirflow API
airflow-dag-processorDAG processing
airflow-pgbouncerAirflow DB connection pool
airflow-schedulerAirflow scheduler
airflow-statsdAirflow metrics
airflow-triggererAirflow triggerer
airflow-workerAirflow task worker
analyticsAnalytics service
audioAudio transcription service
clickhouse-shard0-0ClickHouse analytics DB
clickhouse-zookeeper-0ClickHouse coordination
lhcLHC Core backend
lhc-registryInternal registry
llmLLM integration service
seaweedfs-s3Object storage
postgresql-primary-0Primary PostgreSQL
postgresql-read-0PostgreSQL read replica
valkey-primaryValkey primary
valkey-replicaValkey replica
semanticSemantic processing service
supersetSuperset web application
superset-workerSuperset Celery worker
uiLakehousecat web UI

First start timing: On initial deployment, expect 30–60 minutes before all pods reach Running state, depending on your cluster's available resources and image pull speed from AWS ECR.

Step 6: Access Lakehousecat​

Once deployed, access Lakehousecat via:

Port-forward (for testing):

kubectl port-forward svc/lhc -n lhc-instance 42021:42021
# Access at http://localhost:42021

Ingress (for production): Configure domain and TLS in the Lakehousecat CR (see cluster guides for details).

Subscription Tiers​

Lakehousecat offers multiple subscription tiers to match your needs:

TierUsersData SourcesUse Case
FREE110Evaluation, development
STANDARDUp to 2550Small teams
PREMIUMUp to 100UnlimitedMedium organizations
ENTERPRISE300+UnlimitedLarge enterprises

A handshake key is required for all tiers, including FREE. Register at portal.lakehousecat.com — no paid subscription is required to get started.

Multi-Cloud and On-Premise Support​

The Lakehousecat Operator is designed to run on any Kubernetes cluster, regardless of where it's hosted:

  • ✅ AWS EKS - Tested and supported
  • ⚠️ Google GKE - Not yet tested by Lakehousecat
  • ⚠️ Azure AKS - Not yet tested by Lakehousecat
  • ✅ On-Premise - Self-managed Kubernetes clusters
  • ✅ Minikube - Local development and evaluation

The operator abstracts cloud-specific differences, providing a consistent deployment experience across all environments.

What Gets Deployed​

When you create a Lakehousecat instance, the operator automatically deploys:

Application Services​

  • UI - SvelteKit web application
  • LHC Core - Backend API and business logic
  • Analytics - Query execution and chart generation
  • LLM - Large language model integration
  • Semantic - Natural language processing
  • Audio - Audio transcription services

Data Storage​

  • PostgreSQL - Primary relational database
  • ClickHouse - Columnar analytics database
  • Valkey - In-memory cache and messaging
  • SeaweedFS - S3-compatible object storage

Workflow Orchestration​

  • Apache Airflow - ETL pipeline orchestration
    • Scheduler, Workers, Webserver, Triggerer, DAG Processor

Business Intelligence​

  • Apache Superset - Data visualization and dashboards
    • Web application, Celery workers

Optional Monitoring​

  • Prometheus - Metrics collection
  • Grafana - Metrics visualization

All services are configured, networked, and managed automatically by the operator.

Upgrade and Maintenance​

Upgrades and scaling are managed through the Operator after initial deployment.

  • For upgrade instructions (backup recommendations, Helm commands, verification): see Upgrade
  • For scaling configuration: see Scaling
Do not upgrade components independently

Do not upgrade Airflow, ClickHouse, or Superset via Helm or any other means outside of the Lakehousecat Operator. Independent upgrades break compatibility and are not supported. Component version updates are delivered as part of Lakehousecat releases.

Next Steps​

Choose your deployment path:

  1. TUI Installer - Fastest path: automated local evaluation setup
  2. Manual Setup - Advanced local evaluation with full configuration control
  3. Kubernetes Cluster - Deploy to AWS EKS, Google GKE, Azure AKS, or on-premise

Support​