Skip to main content
Version: Next

Backup

Responsibilities Overview​

Backup and recovery in a Lakehousecat deployment involves two distinct layers of responsibility:

LayerResponsible partyScope
Application-level database backupsLakehousecat (via Job Definitions)PostgreSQL and ClickHouse dumps stored in SeaweedFS
Kubernetes Secrets backupLakehousecat (via the BACKUP_SECRETS Job Definition), once you supply an encryption keyNamespace credential and encryption-key secrets, age-encrypted
Encryption key custodyYouKeeping the age private key safe and off-cluster — without it the secret backup cannot be read
Persistent Volume backupsCluster / infrastructure administratorPVCs for PostgreSQL, ClickHouse, SeaweedFS, Valkey

Application-Level Backups (Built-in Job Definitions)​

Lakehousecat ships with built-in Job Definitions for regular database and object storage backups. These run as Airflow-based jobs inside the cluster:

Job DefinitionWhat it backs upDestination
BACKUP_SECRETSNamespace credential and encryption-key secrets, age-encrypted to your public key. Runs first — see Kubernetes Secrets BackupSeaweedFS (secrets/ prefix), then mirrored out by BACKUP_OBJECT_STORAGE
BACKUP_POSTGRESFull dump of the PostgreSQL instance — includes the application databases lhc, lhc_visualization, and lhc_stagingSeaweedFS (lhc-backend bucket)
BACKUP_CLICKHOUSEFull dump of the ClickHouse instance — includes all loaded data source tables and semantic viewsSeaweedFS (lhc-backend bucket)
BACKUP_OBJECT_STORAGEMirrors the internal SeaweedFS object storage to an external, customer-configured targetExternal S3 or Google Cloud Storage (requires spec.backup.externalObjectStorage on the CR)

These jobs can be triggered two ways:

  • On demand via the CLI — lhc backup triggers all four job types, in the order above, and waits for completion (use --skip-object-storage if the external target is not configured yet). See lhc backup / lhc restore for the full command reference and Backup & Restore for the end-to-end use case.
  • On a schedule via the UI — navigate to Operations → Job Definitions to configure a recurring schedule (e.g., daily at 2:00 AM) for BACKUP_POSTGRES, BACKUP_CLICKHOUSE, and BACKUP_OBJECT_STORAGE. BACKUP_SECRETS is on-demand only and is not scheduled — secrets change rarely, so run it from lhc backup whenever you take a full backup.

Both paths trigger the same underlying Job Definitions — the CLI is a convenience wrapper, not a separate backup mechanism.

Airflow's own metadata database (lhc_jobs) is not part of BACKUP_POSTGRES

BACKUP_POSTGRES/RESTORE_POSTGRES run themselves as Airflow-orchestrated jobs. Including lhc_jobs — the database that holds Airflow's own scheduler and task state — would mean a restore could overwrite the very state that the restore job itself is running under.

This is intentional, not a gap: the definitions of your scheduled jobs live in lhc (backed up normally) and are recreated automatically after a restore. What is not carried over is Airflow's run history for those jobs (which run executed when, past task logs) — no customer data is at risk.

If you also back up PostgreSQL at the Persistent Volume level (see Persistent Volume Backup below), lhc_jobs is already included in that snapshot, since it captures the whole PostgreSQL data directory — no further action needed in that case.

If you want to back up lhc_jobs independently of a PVC snapshot, run pg_dump manually against it, using the same credentials the built-in jobs use (POSTGRES_ADMIN_PASSWORD / POSTGRESQL_HOST from the lhc-airflow-variables Secret in your namespace):

PGPASSWORD=$(kubectl -n <namespace> get secret lhc-airflow-variables -o jsonpath='{.data.POSTGRES_ADMIN_PASSWORD}' | base64 -d) \
pg_dump -h <postgresql-host> -U postgres -d lhc_jobs -F c -b -v -f lhc_jobs.dump

Restore it the same way, against a database named lhc_jobs:

PGPASSWORD=$(kubectl -n <namespace> get secret lhc-airflow-variables -o jsonpath='{.data.POSTGRES_ADMIN_PASSWORD}' | base64 -d) \
pg_restore -h <postgresql-host> -U postgres -d lhc_jobs --clean --if-exists -v lhc_jobs.dump

After a manual restore, restart the Airflow pods (scheduler, worker, triggerer) so they pick up the restored state instead of continuing with in-memory state from before the restore.

important

These backups are stored inside SeaweedFS, which itself runs on a Persistent Volume provisioned by the Kubernetes cluster. The backup data is only as durable as the underlying storage. If the SeaweedFS PVC is lost, the backups are lost with it.

A cluster-level backup strategy for the SeaweedFS PVC — or replication to external object storage — is required for true off-site durability.


Persistent Volume Backup (Cluster Administrator Responsibility)​

The Lakehousecat Operator operates with namespace-scoped RBAC permissions and cannot manage storage classes, Persistent Volumes, or cluster-level infrastructure. The Operator provisions PVCs based on the Helm chart configuration, but the storage class, volume provisioner, and backup mechanism are provided by the Kubernetes cluster.

The following components use Persistent Volumes that must be backed up independently:

ComponentDefault PVC sizeNotes
PostgreSQL (primary)50 GiBStores all application and Airflow metadata
ClickHouse20 GiBStores all loaded data source data
SeaweedFS50 GiBStores files, logs, application backups, and user documents
Valkey8 GiBSession and queue state — not backed up by application jobs

Cluster administrators must formulate and implement a PVC backup strategy using the tools available in their environment — for example:

  • Velero (cluster-agnostic PVC snapshots)
  • AWS EBS Snapshots / EKS Backup
  • Google Cloud Persistent Disk Snapshots / GKE Backup
  • Azure Disk Snapshots / AKS Backup
  • On-premise: storage class native snapshots or volume replication

The backup cadence and retention policy are determined by the cluster administrator based on business requirements.


Kubernetes Secrets Backup — Critical​

Losing the encryption key makes database backups unrestorable

All sensitive data in Lakehousecat — across the backend API, Airflow, and Superset — is encrypted using a symmetric encryption key (ENCRYPT_KEY). This key is stored only as a Kubernetes Secret in the application namespace.

If the namespace is deleted, or if the secrets are lost, encrypted data in the PostgreSQL and ClickHouse backups cannot be decrypted, even if the backup files themselves are intact.

Lakehousecat backs up the namespace secrets for you, through the BACKUP_SECRETS Job Definition — but only once you have given it a key to encrypt them with. The backup captures every credential and encryption-key secret in the namespace (lhc-backend-secret with ENCRYPT_KEY, database passwords, JWT and API-key secrets, provider API keys, OAuth credentials, and so on).

One-time setup: generate an age key pair​

BACKUP_SECRETS encrypts the archive with age before it ever leaves the cluster. Generate the key pair on your own machine — never in the cluster:

age-keygen -o lhc-secret-backup-key.txt

This writes a file containing a public key (age1…) and a private key (AGE-SECRET-KEY-1…).

  1. Put the public key into the Custom Resource:

    spec:
    backup:
    secretsAgePublicKey: "age1...your-public-key..." # pragma: allowlist secret
  2. Store the private key somewhere safe and outside this instance — a password manager or a secrets vault your organization already trusts. You need it only to restore.

You can set or change the public key on a running instance; the Operator applies it on its next reconcile, and the next secret backup is encrypted to the new key. Backups taken before the change remain encrypted to the old key — keep the old private key for as long as you keep those backups.

The private key is yours to keep — there is no copy in the cluster

The secret backup is encrypted to your public key and can only be read with the matching private key. Lakehousecat does not store that private key anywhere. If you lose it, every secret backup you have ever taken becomes permanently unreadable — and with it the encrypted contents of your database backups. Treat it with the same care as the root credentials of your most important system.

If secretsAgePublicKey is not set, the BACKUP_SECRETS job fails on purpose rather than silently producing nothing — and lhc backup then aborts before any database backup runs. Without the secrets, the PostgreSQL backup cannot be decrypted after a cluster loss, while the ClickHouse backup still could: a backup where one half is readable and the other is not looks complete but does not hold up in an emergency, so no partial backup is taken.

Running the secret backup​

BACKUP_SECRETS runs automatically as the first step of lhc backup:

lhc backup

The encrypted archive (secrets.tar.age) is written to the internal object storage under secrets/<timestamp>/ and secrets/latest/, and BACKUP_OBJECT_STORAGE then mirrors it — together with a plain-text restore_secrets.py helper used at restore time — out to your external backup target.

Run the object storage mirror, or the secret backup never leaves the cluster

On its own, BACKUP_SECRETS writes into this instance's internal object storage — the same cluster the secrets exist to outlive. Only BACKUP_OBJECT_STORAGE (which needs spec.backup.externalObjectStorage configured) carries it off-site. A secret backup that stays inside the cluster does not protect you against losing the cluster.

To leave secrets out of a backup deliberately — for example on an evaluation instance you never intend to restore after a cluster loss — pass lhc backup --skip-secrets. The database backups then run, and the command warns that the resulting backup contains no secrets.


What Is Not Backed Up​

ComponentReason
ValkeyIn-memory session and queue state. Valkey is ephemeral by design. Loss of Valkey state requires users to re-login; no data loss occurs in persistent stores.
SeaweedFS storage class / PVCOutside the Operator's scope — cluster administrator responsibility

Backup Checklist​

For a complete and restorable backup, the following must all be covered:

  • age key pair generated, public key set on spec.backup.secretsAgePublicKey, private key stored securely off-cluster
  • BACKUP_SECRETS running successfully as part of lhc backup (the command completes instead of aborting with "no backup was taken")
  • BACKUP_POSTGRES Job Definition scheduled and running successfully (or triggered via lhc backup)
  • BACKUP_CLICKHOUSE Job Definition scheduled and running successfully (or triggered via lhc backup)
  • BACKUP_OBJECT_STORAGE configured and running — this is what carries the secret and database backups off-site
  • SeaweedFS PVC backed up at the cluster level (snapshots or replication)
  • PostgreSQL PVC backed up at the cluster level
  • ClickHouse PVC backed up at the cluster level
  • Backup restore procedure tested in a non-production environment