Skip to main content
Version: 0.0.41

Restore

Restoring a Lakehousecat instance requires coordinating multiple components. The correct restore order and the availability of encryption keys are critical — a restore attempted without the original secrets will fail even if all backup files are intact.

Encryption keys are required for all restores

All sensitive application data is encrypted with a symmetric key stored in the Kubernetes namespace secrets. Restoring the database backups without the original secrets (lhc-encryption-key, lhc-backend-secret) will not restore a working system — the data will be present but unreadable.

Always restore secrets before restoring database content.


Restore Scope​

As with backup, restore responsibilities are split between the application layer and the cluster/infrastructure layer:

LayerResponsible party
Application database restore (PostgreSQL, ClickHouse)Cluster / infrastructure administrator using SeaweedFS-stored backups
Persistent Volume restoreCluster / infrastructure administrator using PVC snapshots
Kubernetes Secrets restoreCluster / infrastructure administrator

Restore Order​

Follow this sequence to restore a Lakehousecat instance:

Step 1 — Restore Kubernetes Secrets​

Before starting any services or applying the Custom Resource, restore the namespace secrets from the secure external backup:

kubectl apply -f lhc-secrets-backup.yaml

Verify the critical secrets are present in the namespace:

kubectl get secrets -n <namespace> | grep -E "lhc-backend-secret|lhc-encryption-key|lhc-api-key|lhc-jwt-secret"
important

If the secrets are not restored first, the Operator will generate new random encryption keys on startup. The new keys will not match the encrypted data in the database backups, making the backup unrestorable.

Step 2 — Restore Persistent Volumes (if needed)​

If the underlying storage was lost (e.g., PVC deleted, node failure, storage corruption), restore PVC snapshots for:

  • PostgreSQL primary (postgresql-primary)
  • ClickHouse (clickhouse-shard0)
  • SeaweedFS (seaweedfs)

Use the snapshot or volume restore mechanism provided by your cluster infrastructure (Velero, EBS snapshot restore, GKE Backup restore, etc.).

If PVCs are intact and only the application data within the database needs to be restored (e.g., accidental data deletion), skip this step and proceed directly to Step 3.

Step 3 — Restore PostgreSQL from SeaweedFS Backup​

Retrieve the latest PostgreSQL backup from SeaweedFS (lhc-backend bucket) and restore it to the PostgreSQL instance:

# Port-forward to PostgreSQL
kubectl port-forward -n <namespace> svc/postgresql-primary 5432:5432

# Restore the backup (adjust filename to the backup you want to restore)
pg_restore -h localhost -U postgres -d <database> <backup-file>

The PostgreSQL instance contains all application databases:

  • lhc — core application data
  • lhc_jobs — runtime state
  • lhc_visualization — Superset metadata (charts, dashboards)
  • lhc_staging — staging data
  • Airflow metadata database

Step 4 — Restore ClickHouse from SeaweedFS Backup​

Retrieve the latest ClickHouse backup from SeaweedFS and restore it to the ClickHouse instance. ClickHouse stores all loaded data source content and semantic views.

# Port-forward to ClickHouse HTTP interface
kubectl port-forward -n <namespace> svc/clickhouse-shard0 8123:8123

Use the ClickHouse backup restore tooling appropriate for the backup format used.

Step 5 — Verify and Restart Services​

After restoring secrets and databases:

  1. Apply or re-apply the Lakehousecat Custom Resource to allow the Operator to reconcile the deployment state
  2. Verify all pods reach Running status
  3. Log in to the application and confirm data is accessible
  4. Check the Job Runs section in Operations to confirm Airflow is functional

Valkey​

Valkey does not require a restore step. Valkey holds in-memory session state and task queues. After a restart, users will need to log in again, but no persistent data is lost — all durable state resides in PostgreSQL, ClickHouse, and SeaweedFS.


SeaweedFS — The Restore Dependency​

SeaweedFS stores the PostgreSQL and ClickHouse backup files produced by the built-in Job Definitions. If SeaweedFS itself is lost and was not backed up at the PVC level, the application-level backups are also lost.

This is why cluster-level PVC backup of the SeaweedFS volume is a critical part of any complete backup strategy. See Backup for the full checklist.


Restore Checklist​

  • Namespace secrets restored from secure off-cluster backup
  • Critical secrets verified present (lhc-encryption-key, lhc-backend-secret, etc.)
  • PVC snapshots restored (if underlying storage was lost)
  • PostgreSQL restored from SeaweedFS backup
  • ClickHouse restored from SeaweedFS backup
  • Operator Custom Resource re-applied
  • All pods running and healthy
  • Application login and data access verified
  • Backup Job Definitions re-enabled and scheduled