Restore
Restoring a Lakehousecat instance requires coordinating multiple components. The correct restore order and the availability of encryption keys are critical — a restore attempted without the original secrets will fail even if all backup files are intact.
All sensitive application data is encrypted with a symmetric key stored in the Kubernetes namespace secrets. Restoring the database backups without the original secrets (lhc-encryption-key, lhc-backend-secret) will not restore a working system — the data will be present but unreadable.
Always restore secrets before restoring database content.
Restore Scope
As with backup, restore responsibilities are split between the application layer and the cluster/infrastructure layer:
| Layer | Responsible party |
|---|---|
| Application database restore (PostgreSQL, ClickHouse) | Cluster / infrastructure administrator using SeaweedFS-stored backups |
| Persistent Volume restore | Cluster / infrastructure administrator using PVC snapshots |
| Kubernetes Secrets restore | Cluster / infrastructure administrator |
Restore Order
Follow this sequence to restore a Lakehousecat instance:
Step 1 — Restore Kubernetes Secrets
Before starting any services or applying the Custom Resource, restore the namespace secrets from the secure external backup:
kubectl apply -f lhc-secrets-backup.yaml
Verify the critical secrets are present in the namespace:
kubectl get secrets -n <namespace> | grep -E "lhc-backend-secret|lhc-encryption-key|lhc-api-key|lhc-jwt-secret"
If the secrets are not restored first, the Operator will generate new random encryption keys on startup. The new keys will not match the encrypted data in the database backups, making the backup unrestorable.
Step 2 — Restore Persistent Volumes (if needed)
If the underlying storage was lost (e.g., PVC deleted, node failure, storage corruption), restore PVC snapshots for:
- PostgreSQL primary (
postgresql-primary) - ClickHouse (
clickhouse-shard0) - SeaweedFS (
seaweedfs)
Use the snapshot or volume restore mechanism provided by your cluster infrastructure (Velero, EBS snapshot restore, GKE Backup restore, etc.).
If PVCs are intact and only the application data within the database needs to be restored (e.g., accidental data deletion), skip this step and proceed directly to Step 3.
Step 3 — Restore PostgreSQL from SeaweedFS Backup
Retrieve the latest PostgreSQL backup from SeaweedFS (lhc-backend bucket) and restore it to the PostgreSQL instance:
# Port-forward to PostgreSQL
kubectl port-forward -n <namespace> svc/postgresql-primary 5432:5432
# Restore the backup (adjust filename to the backup you want to restore)
pg_restore -h localhost -U postgres -d <database> <backup-file>
The PostgreSQL instance contains all application databases:
lhc— core application datalhc_jobs— runtime statelhc_visualization— Superset metadata (charts, dashboards)lhc_staging— staging data- Airflow metadata database
Step 4 — Restore ClickHouse from SeaweedFS Backup
Retrieve the latest ClickHouse backup from SeaweedFS and restore it to the ClickHouse instance. ClickHouse stores all loaded data source content and semantic views.
# Port-forward to ClickHouse HTTP interface
kubectl port-forward -n <namespace> svc/clickhouse-shard0 8123:8123
Use the ClickHouse backup restore tooling appropriate for the backup format used.
Step 5 — Verify and Restart Services
After restoring secrets and databases:
- Apply or re-apply the Lakehousecat Custom Resource to allow the Operator to reconcile the deployment state
- Verify all pods reach
Runningstatus - Log in to the application and confirm data is accessible
- Check the Job Runs section in Operations to confirm Airflow is functional
Valkey
Valkey does not require a restore step. Valkey holds in-memory session state and task queues. After a restart, users will need to log in again, but no persistent data is lost — all durable state resides in PostgreSQL, ClickHouse, and SeaweedFS.
SeaweedFS — The Restore Dependency
SeaweedFS stores the PostgreSQL and ClickHouse backup files produced by the built-in Job Definitions. If SeaweedFS itself is lost and was not backed up at the PVC level, the application-level backups are also lost.
This is why cluster-level PVC backup of the SeaweedFS volume is a critical part of any complete backup strategy. See Backup for the full checklist.
Restore Checklist
- Namespace secrets restored from secure off-cluster backup
- Critical secrets verified present (
lhc-encryption-key,lhc-backend-secret, etc.) - PVC snapshots restored (if underlying storage was lost)
- PostgreSQL restored from SeaweedFS backup
- ClickHouse restored from SeaweedFS backup
- Operator Custom Resource re-applied
- All pods running and healthy
- Application login and data access verified
- Backup Job Definitions re-enabled and scheduled