Docs / Recovery runbooks

Recovery runbooks

These are for the day you need them, read by someone who may never have run safegrd before. Follow the steps in order. Every command here is run, in this order, by an automated test on each change to SafeGrd: the first runbook on a new host enrolled into the same project, the others on a machine that holds nothing but the key file and read access to the bucket.

Which one you follow depends on the key, which the console states on each host. With a SafeGrd-managed key, start at the host is gone. With a customer-managed key, start at before you start.

These steps need a CLI released after v0.0.2: earlier ones cannot find a snapshot without its node id. Check with safegrd --version, and install the latest if yours is older.

With a SafeGrd-managed key: the host is gone

This is the default. Each host's key is sealed at SafeGrd and released only to your organization's enrolled hosts, so a new host can read what the old one wrote, also after the old host is removed from the console. You need a machine that can reach the database or directory you are restoring into, and a login to the console. No key file and no bucket credentials.

  1. On the new machine, install and enrol it:
    curl -fsSL https://safegrd.dev/install.sh | sh
    Answer Y to connect it, then approve the sign-in in a browser on any device. When the organization has more than one project, enrolment asks which: choose the old host's project, so the new host uses the same storage. Without a terminal it joins the default project; run safegrd enroll --project SLUG instead (the slug is on the console's Projects page).
  2. Find the snapshot to restore: safegrd list, or the console's Snapshots page, where each row names the surface and when it was taken. Copy its ID.
  3. Create a new, empty database and restore into it:
    safegrd restore --snapshot SNAPSHOT_ID --target "postgres://user:pass@host:5432/recovered"
    For a directory or a mailbox, use --target-dir ./restored, as in Runbook 2 and Runbook 3. The restore prints Decryption key: SG:…, SafeGrd-managed key of org … (fetched, not stored): the one key this snapshot was sealed to was released to this host for this restore and not written to disk. It checks every table against the digest and row count recorded at backup time before it reports success.
  4. Check what was released, and to which host: the console's Storage page, under Key access record. The row names this host, the key it was given and the snapshot.
  5. Protect the restored database from this host: in the console, Protect a surface, then run the command it shows on this machine.
  6. Retire the old host: open its row's Edit and choose Retire. Its snapshots stay in storage, and SafeGrd keeps its key, so they still restore. Until you retire it, the console marks it Silent after three missed check-ins (at least 15 minutes) and alerts the organization's owners.

The same steps work when the project's backups go to your own bucket set on the console's Storage page: the enrolled host is given the bucket's credentials the same way.

With a customer-managed key: before you start

  1. The private key file (it starts AGE-SECRET-KEY-1). It is usually ~/.safegrd/keys/daemon.key on the host that took the backups. Keep a copy stored offline: for a host that keeps its own key, it is the one file that decrypts these backups.
  2. Read access to the bucket: its name, its region, the endpoint if it is not AWS, and an access key that can list and read it.
  3. The safegrd CLI on the machine you are restoring to: curl -fsSL https://safegrd.dev/install.sh | SAFEGRD_NO_SETUP=1 sh. SAFEGRD_NO_SETUP=1 installs the binary and stops: enrolling this machine would generate a new key, and a restore needs the old one.

You do not need SafeGrd, the old host or its config. If the host that took the backups still works, skip the recovery config below and run the same commands there.

Set up the recovery machine

  1. Copy the key file here and lock it down:
    chmod 600 ./safegrd.key
  2. Find the bucket's prefix. A wrong one finds no snapshots at all. The old host's config.yaml does not have it when the bucket was set in the console, so read it from the bucket: every snapshot is stored as PREFIX/NODE_ID/SNAPSHOT_ID.safegrd, and the prefix is everything before the node id (safegrd/snapshots unless someone changed it).
    aws s3 ls s3://your-bucket/ --recursive | grep '\.safegrd$' | head -3
    Add --endpoint-url https://your-s3-endpoint when it is not AWS. Write the prefix down now, while you can also read it on the console's Storage page.
  3. Write recovery.yaml. Leave out endpoint on AWS. Set force_path_style: true only for MinIO.
    storage:
      type: "s3"
      bucket: "your-bucket"
      region: "us-east-1"
      endpoint: "https://your-s3-endpoint"
      prefix: "safegrd/snapshots"
    encryption:
      key_path: "./safegrd.key"
    safegrd reads a config only its owner can read:
    chmod 600 recovery.yaml
  4. Give it the bucket credentials through the environment, so no secret sits in the file:
    export AWS_ACCESS_KEY_ID=...
    export AWS_SECRET_ACCESS_KEY=...
  5. List what is there. Each row names the snapshot, the node it was filed under, what it holds and when it can be deleted. Incremental file backups are listed in a second table, by month. You do not need to know the node id: restore finds it.
    safegrd --config recovery.yaml list
  6. Prove the snapshot you picked decrypts and is whole before you touch anything. Nothing is written anywhere:
    safegrd --config recovery.yaml verify --snapshot SNAPSHOT_ID --dry-run
    If this fails, stop. Try the next-newest snapshot, and see Troubleshooting.

Runbook 1: the database is gone

  1. Create a new, empty database to restore into. Do not restore over a database that still holds data you may need: restore into a new one, check it, then switch your application over.
    createdb recovered
  2. Restore into it:
    safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target "postgres://user:pass@host:5432/recovered"
    The restore checks the snapshot's digest before it reports success, and says so if it could only check it against the sidecar (with SafeGrd unreachable, that is expected). It creates the roles the schema names that the new cluster lacks, with their passwords, and prints each one; run it as a superuser so every owner and grant comes back (roles, owners and grants).
  3. Check the tables and row counts it printed against what you expect, then run a query your application depends on.
  4. Point your application at recovered. Once it is serving, take a fresh backup from it: safegrd backup --database-url ...

Runbook 2: a directory tree is gone

  1. Restore into a new directory, not over the old path:
    safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target-dir ./restored
  2. Read the last lines it prints. It says what it restored and what it cannot restore: ownership comes back only when you run the restore as root, and hard links, extended attributes and ACLs are not restored (the full list).
  3. If you need ownership back, run the same command with sudo, pointing at the same config and key.
  4. Check a few files you know, then move the directory into place.

A restore never writes outside --target-dir. If the target already has a symlink where the backup has a file or directory, the restore stops and names it, rather than follow it.

Runbook 3: a mailbox is gone

  1. Restore the messages to disk:
    safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --target-dir ./mail
  2. Each mailbox folder becomes a directory and each message a standard .eml file, readable only by you. Any mail client can open them. A Gmail mailbox comes back as one All Mail directory, with each message's labels in .safegrd-email-manifest.json (Gmail labels).
  3. To put them back on a server, import the directory with your mail system's own import tool, or drag the .eml files into the mailbox from a desktop client connected over IMAP.

Runbook 4: restore from an export

safegrd export --to-dir writes a single-archive snapshot as NODE/SNAPSHOT_ID.safegrd with its metadata beside it in SNAPSHOT_ID.meta.json, and an incremental backup (the default) as its month of packs under NODE/repo/. No bucket is needed to restore from either, only the key. The snapshot ID is the one safegrd backup printed, or the name of a file under snapshots/ in the export.

  1. Restore from the export directory, or from one file in it:
    safegrd restore --from ./export --snapshot SNAPSHOT_ID --target-dir ./restored
    safegrd restore --from ./export/NODE/SNAPSHOT_ID.safegrd --target-dir ./restored
    The key comes from --key-path, --private-key or the config, as in the runbooks above. The digest is checked against the remote server's record when the host can reach it, and against the .meta.json beside the file when it cannot.
  2. Use --target instead of --target-dir for a database snapshot, exactly as in Runbook 1.

Without SafeGrd

A .safegrd file is zstd-compressed data encrypted to your Age key, so age and zstd open it:

age -d -i safegrd.key SNAPSHOT_ID.safegrd | zstd -d > payload

For a files or email snapshot, payload is a tar archive: tar -xf payload gives back the directory tree, or the .eml files and their manifest. A PostgreSQL snapshot is a tar archive too: the schema as SQL from pg_dump, then each table's rows in PostgreSQL's binary COPY format. pg_restore does not read it, and psql does: safegrd restore --snapshot ID --to-sql DIR writes it as files with a load.sql that runs them in order, and psql -f load.sql loads them into an empty database with no SafeGrd involved. This route checks nothing: compare sha256sum payload with sha256_checksum in the .meta.json.

This route is for single-archive snapshots: MySQL, MongoDB and SQLite databases, mailboxes, and PostgreSQL databases or files backed up with format: tar. An incremental backup is a set of packs, indexes and directory listings rather than one archive. safegrd export --to-dir copies a whole month of it, and safegrd restore --from reads the copy with the key and no bucket.

Runbook 5: one file, as it was on a given day

For a directory tree backed up incrementally, which is the default for files. The whole tree is Runbook 2.

  1. List the kept versions of the file. Paths are relative to /, with no leading slash:
    safegrd --config recovery.yaml find etc/nginx/nginx.conf
    Each version is one content, numbered from the oldest, with the first and last snapshot that holds it. A directory lists everything below it, and find --deleted var/www/uploads/ lists only files the newest backup no longer holds.
  2. Restore the version whose dates bracket the day you want, into a new directory:
    safegrd --config recovery.yaml restore --path etc/nginx/nginx.conf --version 2 --target-dir ./restored
    Or restore everything under a directory as it was in one snapshot, with the id from find or list:
    safegrd --config recovery.yaml restore --snapshot SNAPSHOT_ID --path 'var/www/uploads/**' --target-dir ./restored
    The file comes back under the target at the path the whole tree would give it (from the root for a surface of one directory, from / otherwise), with the permissions and times it had in that snapshot. Only that file's chunks are downloaded, and every byte is checked against the digest recorded at backup time.
  3. Compare sha256sum of the restored file with the SHA-256 find printed, then copy it into place.

Rehearse it before you need it

Rehearse a restore every quarter, and before a compliance audit, following this page:

  1. Take a machine that is not your production host: a laptop, a fresh VM.
  2. Put only the key file and the read-only bucket credentials on it, and follow Set up the recovery machine.
  3. Run Runbook 1 into a scratch database, Runbook 2 into a scratch directory, and Runbook 5 for one file from a month ago.
  4. Write down the date, the snapshot id and how long it took. That time is your real recovery time. A Fire Drill proves each snapshot restores, and this proves that you can do it. The console’s restore time is measured from drills, some of which may run on SafeGrd’s runner. This one is your own hardware.
  5. Delete the scratch copies, and remove the key file from the machine if it is not kept secure.