Recovery runbooks
These are for the day you need them, read by someone who may never have run
safegrd before. Follow the steps in order. Every command here is run, in
this order, by an automated test on each change to SafeGrd: the
first runbook on a new host enrolled into the same project, the
others on a machine that holds nothing but the key file and read access to the bucket.
Which one you follow depends on the key, which the console states on each host. With a SafeGrd-managed key, start at the host is gone. With a customer-managed key, start at before you start.
These steps need a CLI released after v0.0.2: earlier ones cannot find a snapshot without
its node id. Check with safegrd --version, and install the latest if yours is
older.
With a SafeGrd-managed key: the host is gone
This is the default. Each host's key is sealed at SafeGrd and released only to your organization's enrolled hosts, so a new host can read what the old one wrote, also after the old host is removed from the console. You need a machine that can reach the database or directory you are restoring into, and a login to the console. No key file and no bucket credentials.
- On the new machine, install and enrol it:
curl -fsSL https://safegrd.dev/install.sh | shAnswer
Yto connect it, then approve the sign-in in a browser on any device. When the organization has more than one project, enrolment asks which: choose the old host's project, so the new host uses the same storage. Without a terminal it joins the default project; runsafegrd enroll --project SLUGinstead (the slug is on the console's Projects page). - Find the snapshot to restore:
safegrd list, or the console's Snapshots page, where each row names the surface and when it was taken. Copy its ID. - Create a new, empty database and restore into it:
safegrd restore --snapshot SNAPSHOT_ID --target "postgres://user:pass@host:5432/recovered"For a directory or a mailbox, use
--target-dir ./restored, as in Runbook 2 and Runbook 3. The restore prints Decryption key: SG:…, SafeGrd-managed key of org … (fetched, not stored): the one key this snapshot was sealed to was released to this host for this restore and not written to disk. It checks every table against the digest and row count recorded at backup time before it reports success. - Check what was released, and to which host: the console's Storage page, under Key access record. The row names this host, the key it was given and the snapshot.
- Protect the restored database from this host: in the console, Protect a surface, then run the command it shows on this machine.
- Retire the old host: open its row's Edit and choose Retire. Its snapshots stay in storage, and SafeGrd keeps its key, so they still restore. Until you retire it, the console marks it Silent after three missed check-ins (at least 15 minutes) and alerts the organization's owners.
The same steps work when the project's backups go to your own bucket set on the console's Storage page: the enrolled host is given the bucket's credentials the same way.
With a customer-managed key: before you start
- The private key file (it starts
AGE-SECRET-KEY-1). It is usually~/.safegrd/keys/daemon.keyon the host that took the backups. Keep a copy stored offline: for a host that keeps its own key, it is the one file that decrypts these backups. - Read access to the bucket: its name, its region, the endpoint if it is not AWS, and an access key that can list and read it.
- The
safegrdCLI on the machine you are restoring to:curl -fsSL https://safegrd.dev/install.sh | SAFEGRD_NO_SETUP=1 sh.SAFEGRD_NO_SETUP=1installs the binary and stops: enrolling this machine would generate a new key, and a restore needs the old one.
You do not need SafeGrd, the old host or its config. If the host that took the backups still works, skip the recovery config below and run the same commands there.
Set up the recovery machine
- Copy the key file here and lock it down:
chmod 600 ./safegrd.key
- Find the bucket's
prefix. A wrong one finds no snapshots at all. The old host'sconfig.yamldoes not have it when the bucket was set in the console, so read it from the bucket: every snapshot is stored asPREFIX/NODE_ID/SNAPSHOT_ID.safegrd, and the prefix is everything before the node id (safegrd/snapshotsunless someone changed it).aws s3 ls s3://your-bucket/ --recursive | grep '\.safegrd$' | head -3Add--endpoint-url https://your-s3-endpointwhen it is not AWS. Write the prefix down now, while you can also read it on the console's Storage page. - Write
recovery.yaml. Leave outendpointon AWS. Setforce_path_style: trueonly for MinIO.safegrd reads a config only its owner can read:storage:type: "s3"bucket: "your-bucket"region: "us-east-1"endpoint: "https://your-s3-endpoint"prefix: "safegrd/snapshots"encryption:key_path: "./safegrd.key"chmod 600 recovery.yaml - Give it the bucket credentials through the environment, so no secret sits in the file:
export AWS_ACCESS_KEY_ID=...export AWS_SECRET_ACCESS_KEY=...
- List what is there. Each row names the snapshot, the node it was filed under, what it
holds and when it can be deleted. Incremental file backups are listed in a second table,
by month. You do not need to know the node id: restore finds it.
safegrd --config recovery.yaml list
- Prove the snapshot you picked decrypts and is whole before you touch anything. Nothing
is written anywhere:
safegrd --config recovery.yaml verify --snapshot SNAPSHOT_ID --dry-runIf this fails, stop. Try the next-newest snapshot, and see Troubleshooting.
Runbook 1: the database is gone
- Create a new, empty database to restore into. Do not restore over a
database that still holds data you may need: restore into a new one, check it, then
switch your application over.
createdb recovered
- Restore into it:
safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target "postgres://user:pass@host:5432/recovered"The restore checks the snapshot's digest before it reports success, and says so if it could only check it against the sidecar (with SafeGrd unreachable, that is expected). It creates the roles the schema names that the new cluster lacks, with their passwords, and prints each one; run it as a superuser so every owner and grant comes back (roles, owners and grants).
- Check the tables and row counts it printed against what you expect, then run a query your application depends on.
- Point your application at
recovered. Once it is serving, take a fresh backup from it:safegrd backup --database-url ...
Runbook 2: a directory tree is gone
- Restore into a new directory, not over the old path:
safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target-dir ./restored
- Read the last lines it prints. It says what it restored and what it cannot restore: ownership comes back only when you run the restore as root, and hard links, extended attributes and ACLs are not restored (the full list).
- If you need ownership back, run the same command with
sudo, pointing at the same config and key. - Check a few files you know, then move the directory into place.
A restore never writes outside --target-dir. If the target already has a
symlink where the backup has a file or directory, the restore stops and names it, rather
than follow it.
Runbook 3: a mailbox is gone
- Restore the messages to disk:
safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target-dir ./mail
- Each mailbox folder becomes a directory and each message a standard
.emlfile, readable only by you. Any mail client can open them. A Gmail mailbox comes back as oneAll Maildirectory, with each message's labels in.safegrd-email-manifest.json(Gmail labels). - To put them back on a server, import the directory with your mail system's own import
tool, or drag the
.emlfiles into the mailbox from a desktop client connected over IMAP.
Runbook 4: restore from an export
safegrd export --to-dir writes a single-archive snapshot as NODE/SNAPSHOT_ID.safegrd
with its metadata beside it in SNAPSHOT_ID.meta.json, and an incremental backup (the
default) as its month of packs under NODE/repo/. No bucket is needed to restore
from either, only the key. The snapshot ID is the one safegrd backup printed, or the
name of a file under snapshots/ in the export.
- Restore from the export directory, or from one file in it:
The key comes fromsafegrd restore --from ./export --snapshot SNAPSHOT_ID --target-dir ./restoredsafegrd restore --from ./export/NODE/SNAPSHOT_ID.safegrd --target-dir ./restored
--key-path,--private-keyor the config, as in the runbooks above. The digest is checked against the remote server's record when the host can reach it, and against the.meta.jsonbeside the file when it cannot. - Use
--targetinstead of--target-dirfor a database snapshot, exactly as in Runbook 1.
Without SafeGrd
A .safegrd file is zstd-compressed data encrypted to your Age key, so
age and zstd open it:
For a files or email snapshot, payload is a tar archive: tar -xf payload
gives back the directory tree, or the .eml files and their manifest. A PostgreSQL
snapshot is a tar archive too: the schema as SQL from pg_dump, then each table's
rows in PostgreSQL's binary COPY format. pg_restore does not read it,
and psql does: safegrd restore --snapshot ID --to-sql DIR writes it as
files with a load.sql that runs them in order, and psql -f load.sql
loads them into an empty database with no SafeGrd involved. This route checks nothing: compare
sha256sum payload with sha256_checksum in the .meta.json.
This route is for single-archive snapshots: MySQL, MongoDB and SQLite databases, mailboxes,
and PostgreSQL databases or files backed up with format: tar. An incremental
backup is a set of packs, indexes and directory listings rather than one archive. safegrd export --to-dir copies a
whole month of it, and safegrd restore --from reads the copy with the key and
no bucket.
Runbook 5: one file, as it was on a given day
For a directory tree backed up incrementally, which is the default for files. The whole tree is Runbook 2.
- List the kept versions of the file. Paths are relative to
/, with no leading slash:safegrd --config recovery.yaml find etc/nginx/nginx.confEach version is one content, numbered from the oldest, with the first and last snapshot that holds it. A directory lists everything below it, andfind --deleted var/www/uploads/lists only files the newest backup no longer holds. - Restore the version whose dates bracket the day you want, into a new directory:
safegrd --config recovery.yaml restore --path etc/nginx/nginx.conf --version 2 --target-dir ./restoredOr restore everything under a directory as it was in one snapshot, with the id from
findorlist:safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --path 'var/www/uploads/**' --target-dir ./restoredThe file comes back under the target at the path the whole tree would give it (from the root for a surface of one directory, from/otherwise), with the permissions and times it had in that snapshot. Only that file's chunks are downloaded, and every byte is checked against the digest recorded at backup time. - Compare
sha256sumof the restored file with the SHA-256findprinted, then copy it into place.
Rehearse it before you need it
Rehearse a restore every quarter, and before a compliance audit, following this page:
- Take a machine that is not your production host: a laptop, a fresh VM.
- Put only the key file and the read-only bucket credentials on it, and follow Set up the recovery machine.
- Run Runbook 1 into a scratch database, Runbook 2 into a scratch directory, and Runbook 5 for one file from a month ago.
- Write down the date, the snapshot id and how long it took. That time is your real recovery time. A Fire Drill proves each snapshot restores, and this proves that you can do it. The console’s restore time is measured from drills, some of which may run on SafeGrd’s runner. This one is your own hardware.
- Delete the scratch copies, and remove the key file from the machine if it is not kept secure.