Backup
Responsibilities Overview
Backup and recovery in a Lakehousecat deployment involves two distinct layers of responsibility:
| Layer | Responsible party | Scope |
|---|---|---|
| Application-level database backups | Lakehousecat (via Job Definitions) | PostgreSQL and ClickHouse dumps stored in SeaweedFS |
| Kubernetes Secrets backup | Lakehousecat (via the BACKUP_SECRETS Job Definition), once you supply an encryption key | Namespace credential and encryption-key secrets, age-encrypted |
| Encryption key custody | You | Keeping the age private key safe and off-cluster — without it the secret backup cannot be read |
| Persistent Volume backups | Cluster / infrastructure administrator | PVCs for PostgreSQL, ClickHouse, SeaweedFS, Valkey |
Application-Level Backups (Built-in Job Definitions)
Lakehousecat ships with built-in Job Definitions for regular database and object storage backups. These run as Airflow-based jobs inside the cluster:
| Job Definition | What it backs up | Destination |
|---|---|---|
BACKUP_SECRETS | Namespace credential and encryption-key secrets, age-encrypted to your public key. Runs first — see Kubernetes Secrets Backup | SeaweedFS (secrets/ prefix), then mirrored out by BACKUP_OBJECT_STORAGE |
BACKUP_POSTGRES | Full dump of the PostgreSQL instance — includes the application databases lhc, lhc_visualization, and lhc_staging | SeaweedFS (lhc-backend bucket) |
BACKUP_CLICKHOUSE | Full dump of the ClickHouse instance — includes all loaded data source tables and semantic views | SeaweedFS (lhc-backend bucket) |
BACKUP_OBJECT_STORAGE | Mirrors the internal SeaweedFS object storage to an external, customer-configured target | External S3 or Google Cloud Storage (requires spec.backup.externalObjectStorage on the CR) |
These jobs can be triggered two ways:
- On demand via the CLI —
lhc backuptriggers all four job types, in the order above, and waits for completion (use--skip-object-storageif the external target is not configured yet). Seelhc backup/lhc restorefor the full command reference and Backup & Restore for the end-to-end use case. - On a schedule via the UI — navigate to Operations → Job Definitions to configure a recurring schedule (e.g., daily at 2:00 AM) for
BACKUP_POSTGRES,BACKUP_CLICKHOUSE, andBACKUP_OBJECT_STORAGE.BACKUP_SECRETSis on-demand only and is not scheduled — secrets change rarely, so run it fromlhc backupwhenever you take a full backup.
Both paths trigger the same underlying Job Definitions — the CLI is a convenience wrapper, not a separate backup mechanism.
lhc_jobs) is not part of BACKUP_POSTGRESBACKUP_POSTGRES/RESTORE_POSTGRES run themselves as Airflow-orchestrated jobs. Including
lhc_jobs — the database that holds Airflow's own scheduler and task state — would mean a restore
could overwrite the very state that the restore job itself is running under.
This is intentional, not a gap: the definitions of your scheduled jobs live in lhc (backed up
normally) and are recreated automatically after a restore. What is not carried over is Airflow's
run history for those jobs (which run executed when, past task logs) — no customer data is at
risk.
If you also back up PostgreSQL at the Persistent Volume level (see Persistent Volume Backup
below), lhc_jobs is already included in that snapshot, since it captures the whole PostgreSQL
data directory — no further action needed in that case.
If you want to back up lhc_jobs independently of a PVC snapshot, run pg_dump manually against
it, using the same credentials the built-in jobs use (POSTGRES_ADMIN_PASSWORD /
POSTGRESQL_HOST from the lhc-airflow-variables Secret in your namespace):
PGPASSWORD=$(kubectl -n <namespace> get secret lhc-airflow-variables -o jsonpath='{.data.POSTGRES_ADMIN_PASSWORD}' | base64 -d) \
pg_dump -h <postgresql-host> -U postgres -d lhc_jobs -F c -b -v -f lhc_jobs.dump
Restore it the same way, against a database named lhc_jobs:
PGPASSWORD=$(kubectl -n <namespace> get secret lhc-airflow-variables -o jsonpath='{.data.POSTGRES_ADMIN_PASSWORD}' | base64 -d) \
pg_restore -h <postgresql-host> -U postgres -d lhc_jobs --clean --if-exists -v lhc_jobs.dump
After a manual restore, restart the Airflow pods (scheduler, worker, triggerer) so they pick up the restored state instead of continuing with in-memory state from before the restore.
These backups are stored inside SeaweedFS, which itself runs on a Persistent Volume provisioned by the Kubernetes cluster. The backup data is only as durable as the underlying storage. If the SeaweedFS PVC is lost, the backups are lost with it.
A cluster-level backup strategy for the SeaweedFS PVC — or replication to external object storage — is required for true off-site durability.
Persistent Volume Backup (Cluster Administrator Responsibility)
The Lakehousecat Operator operates with namespace-scoped RBAC permissions and cannot manage storage classes, Persistent Volumes, or cluster-level infrastructure. The Operator provisions PVCs based on the Helm chart configuration, but the storage class, volume provisioner, and backup mechanism are provided by the Kubernetes cluster.
The following components use Persistent Volumes that must be backed up independently:
| Component | Default PVC size | Notes |
|---|---|---|
| PostgreSQL (primary) | 50 GiB | Stores all application and Airflow metadata |
| ClickHouse | 20 GiB | Stores all loaded data source data |
| SeaweedFS | 50 GiB | Stores files, logs, application backups, and user documents |
| Valkey | 8 GiB | Session and queue state — not backed up by application jobs |
Cluster administrators must formulate and implement a PVC backup strategy using the tools available in their environment — for example:
- Velero (cluster-agnostic PVC snapshots)
- AWS EBS Snapshots / EKS Backup
- Google Cloud Persistent Disk Snapshots / GKE Backup
- Azure Disk Snapshots / AKS Backup
- On-premise: storage class native snapshots or volume replication
The backup cadence and retention policy are determined by the cluster administrator based on business requirements.
Kubernetes Secrets Backup — Critical
All sensitive data in Lakehousecat — across the backend API, Airflow, and Superset — is encrypted using a symmetric encryption key (ENCRYPT_KEY). This key is stored only as a Kubernetes Secret in the application namespace.
If the namespace is deleted, or if the secrets are lost, encrypted data in the PostgreSQL and ClickHouse backups cannot be decrypted, even if the backup files themselves are intact.
Lakehousecat backs up the namespace secrets for you, through the BACKUP_SECRETS Job Definition — but only once you have given it a key to encrypt them with. The backup captures every credential and encryption-key secret in the namespace (lhc-backend-secret with ENCRYPT_KEY, database passwords, JWT and API-key secrets, provider API keys, OAuth credentials, and so on).
One-time setup: generate an age key pair
BACKUP_SECRETS encrypts the archive with age before it ever leaves the cluster. Generate the key pair on your own machine — never in the cluster:
age-keygen -o lhc-secret-backup-key.txt
This writes a file containing a public key (age1…) and a private key (AGE-SECRET-KEY-1…).
-
Put the public key into the Custom Resource:
spec:
backup:
secretsAgePublicKey: "age1...your-public-key..." # pragma: allowlist secret -
Store the private key somewhere safe and outside this instance — a password manager or a secrets vault your organization already trusts. You need it only to restore.
You can set or change the public key on a running instance; the Operator applies it on its next reconcile, and the next secret backup is encrypted to the new key. Backups taken before the change remain encrypted to the old key — keep the old private key for as long as you keep those backups.
The secret backup is encrypted to your public key and can only be read with the matching private key. Lakehousecat does not store that private key anywhere. If you lose it, every secret backup you have ever taken becomes permanently unreadable — and with it the encrypted contents of your database backups. Treat it with the same care as the root credentials of your most important system.
If secretsAgePublicKey is not set, the BACKUP_SECRETS job fails on purpose rather than silently producing nothing — and lhc backup then aborts before any database backup runs. Without the secrets, the PostgreSQL backup cannot be decrypted after a cluster loss, while the ClickHouse backup still could: a backup where one half is readable and the other is not looks complete but does not hold up in an emergency, so no partial backup is taken.
Running the secret backup
BACKUP_SECRETS runs automatically as the first step of lhc backup:
lhc backup
The encrypted archive (secrets.tar.age) is written to the internal object storage under secrets/<timestamp>/ and secrets/latest/, and BACKUP_OBJECT_STORAGE then mirrors it — together with a plain-text restore_secrets.py helper used at restore time — out to your external backup target.
On its own, BACKUP_SECRETS writes into this instance's internal object storage — the same cluster the secrets exist to outlive. Only BACKUP_OBJECT_STORAGE (which needs spec.backup.externalObjectStorage configured) carries it off-site. A secret backup that stays inside the cluster does not protect you against losing the cluster.
To leave secrets out of a backup deliberately — for example on an evaluation instance you never intend to restore after a cluster loss — pass lhc backup --skip-secrets. The database backups then run, and the command warns that the resulting backup contains no secrets.
What Is Not Backed Up
| Component | Reason |
|---|---|
| Valkey | In-memory session and queue state. Valkey is ephemeral by design. Loss of Valkey state requires users to re-login; no data loss occurs in persistent stores. |
| SeaweedFS storage class / PVC | Outside the Operator's scope — cluster administrator responsibility |
Backup Checklist
For a complete and restorable backup, the following must all be covered:
- age key pair generated, public key set on
spec.backup.secretsAgePublicKey, private key stored securely off-cluster -
BACKUP_SECRETSrunning successfully as part oflhc backup(the command completes instead of aborting with "no backup was taken") -
BACKUP_POSTGRESJob Definition scheduled and running successfully (or triggered vialhc backup) -
BACKUP_CLICKHOUSEJob Definition scheduled and running successfully (or triggered vialhc backup) -
BACKUP_OBJECT_STORAGEconfigured and running — this is what carries the secret and database backups off-site - SeaweedFS PVC backed up at the cluster level (snapshots or replication)
- PostgreSQL PVC backed up at the cluster level
- ClickHouse PVC backed up at the cluster level
- Backup restore procedure tested in a non-production environment