Skip to main content
Version: Next

Storage Planning

Lakehousecat stores its persistent data on Kubernetes PersistentVolumeClaims (PVCs) — one set per stateful component (PostgreSQL, ClickHouse, the object store, Valkey, Superset). Two properties of those PVCs are chosen once, in the Custom Resource, before the first deployment:

  • the storage class (storageClass) — which provisioner backs the volume
  • the volume size (persistenceSize / persistence.size)

This page explains why both are a pre-deployment decision and how to set them. It is about the persistent data volumes. The separate, node-local scratch space a data load needs while it runs is covered in Load Storage Sizing.

The defaults are fine for the normal case

Every field below has a working default, and a standard deployment runs well on them. You only need to think about this page if you already know you will be loading large data volumes, or if your cluster has no usable default storage class. Capacity planning for large workloads is the operator's responsibility — Lakehousecat does not auto-resize volumes or grow them on demand.

Why this is decided up front​

Changing a PVC after it exists is constrained by Kubernetes and by the storage backend, not by Lakehousecat:

ChangeFeasibility
Grow a volumePossible only if the storage class has allowVolumeExpansion: true and the provisioner supports online resize. Many on-premise and some cloud classes do not.
Shrink a volumeNot supported by Kubernetes on any common provisioner.
Switch storage classNo in-place path. Requires provisioning a new PVC on the target class and migrating the data (e.g. a database dump/restore or a volume clone), with downtime for that component.

So a storage class or size that turns out to be wrong is a migration, not an edit. Picking them deliberately before the first apply avoids that.

What to set, and where​

All storage settings live on the Lakehousecat Custom Resource under spec.

Global default storage class​

spec:
platform:
storageClass: "gp3" # empty string = use the cluster's default StorageClass

spec.platform.storageClass applies to every component that does not set its own. Leave it empty to use whatever the cluster marks as the default StorageClass. Check what is available with:

kubectl get storageclass

Per-component overrides​

Any component can override the global class, and each carries its own size. The two that dominate total storage — and the two most likely to need attention for a large deployment — are PostgreSQL and ClickHouse.

spec:
# PostgreSQL — operational metadata, sessions, model definitions.
# NOTE: PostgreSQL uses a flat `persistenceSize`, not a `persistence:` block.
postgresql:
persistenceSize: "50Gi" # default: 50Gi
storageClass: "gp3" # optional; falls back to platform.storageClass

# ClickHouse — the analytical warehouse. This is where loaded datasource
# data lives; size it against the data you intend to load.
clickhouse:
persistence:
enabled: true
size: "500Gi"
storageClass: "gp3"

# Object storage — only when using the in-cluster SeaweedFS default
# (spec.objectStorage.type: seaweedfs). Holds uploaded files and RAG artifacts.
objectStorage:
volume:
persistence:
size: "500Gi"
storageClass: "gp3"

# Supporting components — defaults are almost always sufficient.
valkey:
persistence:
size: "8Gi"
superset:
persistence:
size: "8Gi"
External object storage skips its PVC

If spec.objectStorage.type is s3 (an external bucket), the objectStorage.volume block does not apply — capacity is managed by the S3 provider, not by a PVC.

lhc_staging growth guard​

spec.postgresql.stagingDbSizeThresholdGi (default 20) is not a size setting — it is a warning threshold. If the lhc_staging database on the shared PostgreSQL volume grows past it, a built-in Airflow job turns red to flag it. If you deliberately expect large volumes through the staging write path, raise this together with postgresql.persistenceSize rather than in isolation.

Sizing guidance​

  • ClickHouse carries the analytical warehouse. Its footprint tracks the total size of the datasource data you load, after columnar compression (typically a fraction of the raw source size, but plan against the source size for headroom). This is the field to raise for a data-heavy deployment.
  • PostgreSQL holds operational metadata — sessions, model and datasource definitions, users. It grows slowly with usage; the 50Gi default suits most instances.
  • Object storage holds uploaded files and unstructured-data artifacts. Size it to the volume of files your users will upload.
  • The minimum production baseline (16 CPU / 64 GB RAM / 500 GB storage) is in Kubernetes Cluster Setup — Prerequisites.

Storage class recommendations by platform​

Provider-specific classes (for example gp3 on EKS, pd-ssd / premium-rwo on GKE, managed-premium on AKS, or local-path / Longhorn / Ceph on-premise) and how to verify them are covered in the platform guides:

Prefer an SSD-backed class with allowVolumeExpansion: true for the database volumes — it keeps the "grow later" option open even though shrinking and class changes remain migrations.