Restore
Restoring a Lakehousecat instance requires coordinating multiple components. The correct restore order and the availability of encryption keys are critical — a restore attempted without the original secrets will fail even if all backup files are intact.
All sensitive application data is encrypted with a symmetric key stored in the Kubernetes namespace secrets. Restoring the database backups without the original lhc-backend-secret will not restore a working system — the data will be present but unreadable.
Always restore secrets first — before applying the Custom Resource, and before restoring database content.
Restore Scope
As with backup, restore responsibilities are split between the application layer and the cluster/infrastructure layer:
| Layer | Responsible party |
|---|---|
| Kubernetes Secrets restore | You, with the age private key — a documented manual step, not covered by lhc restore (see Step 1) |
| Application-level restore (PostgreSQL, ClickHouse, object storage) | Cluster / infrastructure administrator, via lhc restore |
| Persistent Volume restore | Cluster / infrastructure administrator using PVC snapshots |
The secret backup runs as an Airflow job inside the cluster. The restore cannot: in a disaster-recovery scenario the secrets are the precondition for Airflow — and the whole instance — starting at all. So restoring them is a short manual step you run from your own machine, before the instance exists. It is a deliberate design decision, not a missing feature.
Restore Order
Follow this sequence to restore a Lakehousecat instance:
Step 1 — Restore Kubernetes Secrets
Only needed when rebuilding the instance in a fresh namespace. If you are rolling back a still-running instance (its namespace and secrets are intact), skip to Step 4.
This must happen before the Lakehousecat Custom Resource is applied for the first time. If the CR is applied first, the Operator's first reconcile generates fresh secrets — including a new ENCRYPT_KEY — and the restore becomes silently ineffective.
You need, on your own machine:
age,python3, andkubectl(pointed at the target cluster)- the age private key you kept when you set up the secret backup (see Backup)
- from your external backup target, both files from the same
secrets/<timestamp>/(orsecrets/latest/) folder:secrets.tar.age— the encrypted archiverestore_secrets.py— the helper that rewrites the secrets for the target namespace (shipped in plain text alongside every archive)
Create the target namespace, then decrypt and apply:
kubectl create namespace <target-namespace>
age -d -i lhc-secret-backup-key.txt secrets.tar.age \
| python3 restore_secrets.py --namespace <target-namespace> \
| kubectl apply -f -
restore_secrets.py applies only the credential and encryption-key secrets the Operator preserves on first sight, and rewrites their namespace (and one Helm ownership annotation — without which the first Airflow chart install would fail with invalid ownership metadata).
Verify the critical secret is present:
kubectl -n <target-namespace> get secret lhc-backend-secret
When you apply the Custom Resource (Step 3 below), check the Operator log. For each secret it must report PRESERVING existing credentials (or SKIPPING (immutable)) — not generating new. If you see generating new for lhc-backend-secret, the CR reached the cluster before the secrets and the restore did not take; delete the namespace and start over from this step.
Step 2 — Restore Persistent Volumes (if needed)
If the underlying storage was lost (e.g., PVC deleted, node failure, storage corruption), restore PVC snapshots for:
- PostgreSQL primary (
postgresql-primary) - ClickHouse (
clickhouse-shard0) - SeaweedFS (
seaweedfs)
Use the snapshot or volume restore mechanism provided by your cluster infrastructure (Velero, EBS snapshot restore, GKE Backup restore, etc.).
If PVCs are intact and only the application data within the database needs to be restored (e.g., accidental data deletion), skip this step and proceed directly to Step 4.
Step 3 — Apply the Custom Resource and wait for the instance
With the secrets in place (Step 1), apply the Lakehousecat Custom Resource so the Operator builds the instance:
kubectl apply -f lakehousecat.yaml
Watch the Operator log and confirm it reports PRESERVING existing credentials / SKIPPING (immutable) for the secrets — not generating new (see the note in Step 1). Wait until all pods reach Running. Airflow must be up before the next step, because lhc restore runs through it.
If you are restoring in place on an instance that is already running (PVCs intact, no namespace rebuild), the CR is already applied — skip this step.
Step 4 — Restore PostgreSQL, ClickHouse, and Object Storage via the CLI
Restoring the application-level backups (PostgreSQL, ClickHouse, and — if configured — the
external object storage mirror) is done with a single CLI command, lhc restore. It triggers the
built-in RESTORE_POSTGRES / RESTORE_CLICKHOUSE / RESTORE_OBJECT_STORAGE Job Definitions and
waits for them to complete — no manual pg_restore or port-forwarding required.
lhc restore --yes \
--postgres-timestamp latest \
--clickhouse-backup-path 20260711210000
- PostgreSQL defaults to the
latestbackup (--postgres-timestamp). - ClickHouse has no
latestalias — pass the exact backup timestamp folder via--clickhouse-backup-path, or skip it with--skip-clickhouse. - Add
--skip-object-storageif no external object storage target is configured. --yesis required — the command refuses to run without it, since restore overwrites current data.
The PostgreSQL instance contains the application databases lhc, lhc_visualization, and
lhc_staging — all are restored together. It also contains lhc_jobs, Airflow's own scheduler
and task metadata database — this one is deliberately not part of RESTORE_POSTGRES, since
RESTORE_POSTGRES itself runs as an Airflow job and would otherwise overwrite the state it is
currently running under. See Backup — Airflow's own metadata database
for why, and for the manual restore procedure if you back up lhc_jobs independently of a PVC
snapshot. This has no effect on your job definitions themselves (in lhc, restored normally and
recreated automatically if missing) — only on Airflow's run history for those jobs.
ClickHouse holds all loaded data source content and semantic views.
lhc restore needs no scaled-down deployments: it runs as a job through lhc and Airflow, and
it keeps Superset out of its metadata database while that database is replaced. Don't scale
deployments by hand, as that leaves them out of line with the Lakehousecat resource. Run the
restore in a maintenance window, since changes made while it runs are overwritten. See
lhc backup / lhc restore for the full flag reference
and Backup & Restore for the complete,
worked-through sequence.
Two lhc restore bugs were found and fixed in August 2026 — a foreign key on Airflow's
jobrun table blocking the lhc database restore, and ClickHouse's own INFORMATION_SCHEMA
system databases being included in the ClickHouse restore. Both are handled automatically now;
no customer workaround is needed. If you hit either error message against an older image build,
upgrade to the current release.
Step 5 — Verify
After restoring secrets and databases:
- Verify all pods are
Running - Log in to the application and confirm data is accessible
- Check the Job Runs section in Operations to confirm Airflow is functional
- Confirm a data source connection still works and a chart with history renders — these exercise the encrypted content directly, so a key mismatch shows up here first
Valkey
Valkey does not require a restore step. Valkey holds in-memory session state and task queues. After a restart, users will need to log in again, but no persistent data is lost — all durable state resides in PostgreSQL, ClickHouse, and SeaweedFS.
SeaweedFS — The Restore Dependency
SeaweedFS stores the PostgreSQL and ClickHouse backup files produced by the built-in Job Definitions. If SeaweedFS itself is lost and was not backed up at the PVC level, the application-level backups are also lost.
This is why cluster-level PVC backup of the SeaweedFS volume is a critical part of any complete backup strategy. See Backup for the full checklist.
Restore Checklist
- age private key and both
secrets.tar.age+restore_secrets.pyretrieved from the backup target - Namespace secrets decrypted and applied before the Custom Resource
- Operator log confirms
PRESERVING existing credentials(notgenerating new) forlhc-backend-secret - PVC snapshots restored (if underlying storage was lost)
- Custom Resource applied, all pods
Running, Airflow up - Maintenance window: no users working on the instance during
lhc restore -
lhc restore --yescompleted successfully for PostgreSQL, ClickHouse, and (if configured) object storage - All pods running and healthy
- Application login and data access verified
- Backup Job Definitions re-enabled and scheduled