Skip to main content
Version: Next

Backup & Restore

This guide walks through drawing a full backup before a Lakehousecat upgrade, and restoring it if the upgrade needs to be rolled back. See the backup / restore command reference for all available flags.

One-Time Setup: Secret Backup Key​

lhc backup backs up the namespace secrets (including the ENCRYPT_KEY that decrypts your database backups), age-encrypted. This needs an age key pair, generated on your own machine:

age-keygen -o lhc-secret-backup-key.txt

Put the public key (age1…) on spec.backup.secretsAgePublicKey in the Custom Resource, and keep the private key somewhere safe and off-cluster — you need it to restore, and there is no copy anywhere else. Full detail in Backup — Kubernetes Secrets Backup.

Before an Upgrade: Take a Full Backup​

lhc backup

If the external object storage target is not configured yet (see below), skip that part of the backup:

lhc backup --skip-object-storage

The command waits for all triggered jobs to finish and prints a summary table. If you lose the connection while it is running, check the status of the triggered runs separately:

lhc jobs runs --select id,status,job_type

Configuring an External Object Storage Target (optional)​

Backing up object storage to an external, customer-controlled target (S3 or Google Cloud Storage) requires this to be configured once by an administrator on the Lakehousecat instance (via the deployment configuration), together with a Kubernetes secret holding the target's credentials. Without this configuration, lhc backup/lhc restore still cover Postgres and ClickHouse — only the object storage step is unavailable.

After a Failed Upgrade: Restore​

If you are rebuilding the instance from scratch (new namespace), restore the namespace secrets before applying the Custom Resource — lhc restore does not do this. See Restore — Step 1. For a rollback on a still-running instance the secrets are already present; continue here.

Don't scale any deployment down for the restore. lhc restore runs as a job through the lhc service and Airflow, so both have to stay up. The restore keeps Superset out of its own database while that database is replaced, so Superset needs no preparation either. Scaling deployments by hand would also leave them out of line with the Lakehousecat resource. Run the restore in a maintenance window, though: changes users make while it runs are overwritten.

ClickHouse backups have no "latest" alias, so determine the backup timestamp to restore first:

lhc jobs runs --select id,status,job_type

Then run the restore:

lhc restore --yes \
--postgres-timestamp latest \
--clickhouse-backup-path 20260711210000

To restore only some components, skip the rest:

lhc restore --yes --skip-clickhouse --skip-object-storage