Skip to main content
Version: Next

Restore

Restoring a Lakehousecat instance requires coordinating multiple components. The correct restore order and the availability of encryption keys are critical — a restore attempted without the original secrets will fail even if all backup files are intact.

Encryption keys are required for all restores

All sensitive application data is encrypted with a symmetric key stored in the Kubernetes namespace secrets. Restoring the database backups without the original lhc-backend-secret will not restore a working system — the data will be present but unreadable.

Always restore secrets first — before applying the Custom Resource, and before restoring database content.


Restore Scope​

As with backup, restore responsibilities are split between the application layer and the cluster/infrastructure layer:

LayerResponsible party
Kubernetes Secrets restoreYou, with the age private key — a documented manual step, not covered by lhc restore (see Step 1)
Application-level restore (PostgreSQL, ClickHouse, object storage)Cluster / infrastructure administrator, via lhc restore
Persistent Volume restoreCluster / infrastructure administrator using PVC snapshots
Why is the secret restore manual when the backup is automatic?

The secret backup runs as an Airflow job inside the cluster. The restore cannot: in a disaster-recovery scenario the secrets are the precondition for Airflow — and the whole instance — starting at all. So restoring them is a short manual step you run from your own machine, before the instance exists. It is a deliberate design decision, not a missing feature.


Restore Order​

Follow this sequence to restore a Lakehousecat instance:

Step 1 — Restore Kubernetes Secrets​

Only needed when rebuilding the instance in a fresh namespace. If you are rolling back a still-running instance (its namespace and secrets are intact), skip to Step 4.

This must happen before the Lakehousecat Custom Resource is applied for the first time. If the CR is applied first, the Operator's first reconcile generates fresh secrets — including a new ENCRYPT_KEY — and the restore becomes silently ineffective.

You need, on your own machine:

  • age, python3, and kubectl (pointed at the target cluster)
  • the age private key you kept when you set up the secret backup (see Backup)
  • from your external backup target, both files from the same secrets/<timestamp>/ (or secrets/latest/) folder:
    • secrets.tar.age — the encrypted archive
    • restore_secrets.py — the helper that rewrites the secrets for the target namespace (shipped in plain text alongside every archive)

Create the target namespace, then decrypt and apply:

kubectl create namespace <target-namespace>

age -d -i lhc-secret-backup-key.txt secrets.tar.age \
| python3 restore_secrets.py --namespace <target-namespace> \
| kubectl apply -f -

restore_secrets.py applies only the credential and encryption-key secrets the Operator preserves on first sight, and rewrites their namespace (and one Helm ownership annotation — without which the first Airflow chart install would fail with invalid ownership metadata).

Verify the critical secret is present:

kubectl -n <target-namespace> get secret lhc-backend-secret
Confirm the Operator preserved the restored secrets

When you apply the Custom Resource (Step 3 below), check the Operator log. For each secret it must report PRESERVING existing credentials (or SKIPPING (immutable)) — not generating new. If you see generating new for lhc-backend-secret, the CR reached the cluster before the secrets and the restore did not take; delete the namespace and start over from this step.

Step 2 — Restore Persistent Volumes (if needed)​

If the underlying storage was lost (e.g., PVC deleted, node failure, storage corruption), restore PVC snapshots for:

  • PostgreSQL primary (postgresql-primary)
  • ClickHouse (clickhouse-shard0)
  • SeaweedFS (seaweedfs)

Use the snapshot or volume restore mechanism provided by your cluster infrastructure (Velero, EBS snapshot restore, GKE Backup restore, etc.).

If PVCs are intact and only the application data within the database needs to be restored (e.g., accidental data deletion), skip this step and proceed directly to Step 4.

Step 3 — Apply the Custom Resource and wait for the instance​

With the secrets in place (Step 1), apply the Lakehousecat Custom Resource so the Operator builds the instance:

kubectl apply -f lakehousecat.yaml

Watch the Operator log and confirm it reports PRESERVING existing credentials / SKIPPING (immutable) for the secrets — not generating new (see the note in Step 1). Wait until all pods reach Running. Airflow must be up before the next step, because lhc restore runs through it.

If you are restoring in place on an instance that is already running (PVCs intact, no namespace rebuild), the CR is already applied — skip this step.

Step 4 — Restore PostgreSQL, ClickHouse, and Object Storage via the CLI​

Restoring the application-level backups (PostgreSQL, ClickHouse, and — if configured — the external object storage mirror) is done with a single CLI command, lhc restore. It triggers the built-in RESTORE_POSTGRES / RESTORE_CLICKHOUSE / RESTORE_OBJECT_STORAGE Job Definitions and waits for them to complete — no manual pg_restore or port-forwarding required.

lhc restore --yes \
--postgres-timestamp latest \
--clickhouse-backup-path 20260711210000
  • PostgreSQL defaults to the latest backup (--postgres-timestamp).
  • ClickHouse has no latest alias — pass the exact backup timestamp folder via --clickhouse-backup-path, or skip it with --skip-clickhouse.
  • Add --skip-object-storage if no external object storage target is configured.
  • --yes is required — the command refuses to run without it, since restore overwrites current data.

The PostgreSQL instance contains the application databases lhc, lhc_visualization, and lhc_staging — all are restored together. It also contains lhc_jobs, Airflow's own scheduler and task metadata database — this one is deliberately not part of RESTORE_POSTGRES, since RESTORE_POSTGRES itself runs as an Airflow job and would otherwise overwrite the state it is currently running under. See Backup — Airflow's own metadata database for why, and for the manual restore procedure if you back up lhc_jobs independently of a PVC snapshot. This has no effect on your job definitions themselves (in lhc, restored normally and recreated automatically if missing) — only on Airflow's run history for those jobs.

ClickHouse holds all loaded data source content and semantic views.

lhc restore needs no scaled-down deployments: it runs as a job through lhc and Airflow, and it keeps Superset out of its metadata database while that database is replaced. Don't scale deployments by hand, as that leaves them out of line with the Lakehousecat resource. Run the restore in a maintenance window, since changes made while it runs are overwritten. See lhc backup / lhc restore for the full flag reference and Backup & Restore for the complete, worked-through sequence.

Two restore fixes you may hear about

Two lhc restore bugs were found and fixed in August 2026 — a foreign key on Airflow's jobrun table blocking the lhc database restore, and ClickHouse's own INFORMATION_SCHEMA system databases being included in the ClickHouse restore. Both are handled automatically now; no customer workaround is needed. If you hit either error message against an older image build, upgrade to the current release.

Step 5 — Verify​

After restoring secrets and databases:

  1. Verify all pods are Running
  2. Log in to the application and confirm data is accessible
  3. Check the Job Runs section in Operations to confirm Airflow is functional
  4. Confirm a data source connection still works and a chart with history renders — these exercise the encrypted content directly, so a key mismatch shows up here first

Valkey​

Valkey does not require a restore step. Valkey holds in-memory session state and task queues. After a restart, users will need to log in again, but no persistent data is lost — all durable state resides in PostgreSQL, ClickHouse, and SeaweedFS.


SeaweedFS — The Restore Dependency​

SeaweedFS stores the PostgreSQL and ClickHouse backup files produced by the built-in Job Definitions. If SeaweedFS itself is lost and was not backed up at the PVC level, the application-level backups are also lost.

This is why cluster-level PVC backup of the SeaweedFS volume is a critical part of any complete backup strategy. See Backup for the full checklist.


Restore Checklist​

  • age private key and both secrets.tar.age + restore_secrets.py retrieved from the backup target
  • Namespace secrets decrypted and applied before the Custom Resource
  • Operator log confirms PRESERVING existing credentials (not generating new) for lhc-backend-secret
  • PVC snapshots restored (if underlying storage was lost)
  • Custom Resource applied, all pods Running, Airflow up
  • Maintenance window: no users working on the instance during lhc restore
  • lhc restore --yes completed successfully for PostgreSQL, ClickHouse, and (if configured) object storage
  • All pods running and healthy
  • Application login and data access verified
  • Backup Job Definitions re-enabled and scheduled