Docs / Restore and Fire Drills

Restore and Fire Drills

SafeGrd restores each surface's newest backup on a schedule, checks what came back, and signs a dated record of the result.

Restore, for real

# into a PostgreSQL target
safegrd restore --snapshot snap-1b094c0af8 --target "$DATABASE_URL"
# files or mail, onto disk
safegrd restore --snapshot snap-94dec39a5e --target-dir ./recovered

The ciphertext is pulled from storage, decrypted on this machine with your private key, decompressed and written to the target. The key is resolved from key_path, from SAFEGRD_PRIVATE_KEY, or from --key-path / --private-key on a host that is not the one that took the backup.

Fire Drills

A Fire Drill restores a snapshot somewhere disposable, counts what came back and records the result. It runs in your environment against your key; the control plane receives the report, never the data.

# in-memory dry restore, no database needed
safegrd verify --snapshot snap-1b094c0af8
# full drill into an ephemeral database
safegrd verify --snapshot snap-1b094c0af8 --sandbox-target "$SANDBOX_URL"

The default dry run decrypts and replays the stream in memory and asserts table counts, row counts, column counts and extensions without a running PostgreSQL anywhere. It checks the data but does not load the schema; only a sandbox drill does that. Because it needs no database, you can run it in CI or on a laptop. Pass --sandbox-target without --dry-run for the full drill, which actually loads the data into a throwaway database and checks it there. The remote server records sandbox drills on every paid plan; on the free plan it records the in-memory drill, and refuses a sandbox one with the reason.

Every plan runs scheduled Fire Drills. The free plan drills once a month, in memory, and the paid plans drill weekly or daily and record sandbox drills (pricing).

The Fire Drills tab: the drill record, then each surface's last verified drill, recovery point and restore time. The Fire Drills tab: the drill record, then each surface's last verified drill, recovery point and restore time.
The Fire Drills tab. History lists every drill with what was restored and its certificate, and Badge and public page holds the badge.

A full drill of a demo database, as safegrd verify prints it:

Fire Drill: restoring snap-20261003-160211-06e021 into the sandbox database
   Snapshot ID:    snap-20261003-160211-06e021
   Sandbox:        postgres://postgres:xxxxx@localhost:33776/drill_sandbox?sslmode=disable

Fire Drill Passed
   Verification ID: verif-4075149c
   Duration:        318ms
   Tables restored: 4
   Rows restored:   45404
   Certificate:     cert_sg_5589e8b170277fd955261211d9db3188e70566e1a5f7d8cc043edc3fddd18263

Assertions:
   [PASS] DigestIntegrity (Expected: da3ca6f50beebda676ad61c158deaa458d0c3c9dfca1715ce3a84c52d8af463b, Actual: da3ca6f50beebda676ad61c158deaa458d0c3c9dfca1715ce3a84c52d8af463b)
   [PASS] RemoteServerRecord (Expected: da3ca6f50beebda676ad61c158deaa458d0c3c9dfca1715ce3a84c52d8af463b, Actual: da3ca6f50beebda676ad61c158deaa458d0c3c9dfca1715ce3a84c52d8af463b)
   [PASS] TableCountMatch (Expected: 4 tables, Actual: 4 tables)
   [PASS] RowCountMatch (Expected: 45404 rows, Actual: 45404 rows)
   [PASS] ExtensionBootCheck (Expected: [], Actual: [])
   [PASS] Schema Captured By pg_dump (Expected: schema from pg_dump, Actual: pg_dump 18.6)

Drills and backups run on SafeGrd

A surface with no daemon beside it, such as one backed up from a GitHub Actions workflow, can have SafeGrd run its drills. Set Drill on SafeGrd on the surface in the console. From then on, each time a drill is due, SafeGrd starts a machine for that drill alone, restores the newest snapshot there, checks it as the host would, signs the record, and destroys the machine. The record is the same one a host would send, and the Fire Drills page shows that it ran on SafeGrd.

It is offered when SafeGrd holds the organization’s key, since the machine decrypts the snapshot, and for backups in hosted storage. Every surface type drills there: PostgreSQL, MySQL, MariaDB, MongoDB and SQLite into a throwaway database of the same engine, files and mailboxes read back in memory. The key is released to a token minted for that one job, which reaches that snapshot and nothing else, and every release is in the key access log as the drill runner. The backup goes from storage to that machine and nowhere else.

A database with no server of your own beside it, such as Supabase, Neon or RDS, can be backed up by SafeGrd as well. For Supabase, Pick a Supabase project signs you in there and lists your projects with the session pooler's address filled in; you type the project's database password, which Supabase does not give out. Under Surfaces, Back up on SafeGrd takes a name, the engine (PostgreSQL, MySQL or MongoDB), the connection string, how often and how long to keep each backup. SafeGrd seals the connection string and starts a machine for each backup, which connects to the database, dumps it, encrypts it to the organization’s key and writes it to hosted storage. A PostgreSQL backup uploads only what changed since the last one: the machine reads a cache of what the repository holds, which SafeGrd keeps encrypted between backups, to a key it holds sealed and gives only to that surface’s backup machines. The drill of that backup follows on the plan’s cadence. The machine that takes the backup is given the connection string and never the key that decrypts backups, and the machine that drills it is given that key and never the connection string. The first backup starts within a few minutes of adding the surface, so a wrong string is found out then. The database has to accept connections from the internet, and a read-only role is enough. Where SafeGrd's machines connect from fixed addresses, the console shows them beside Back up on SafeGrd, for a database whose firewall lists addresses. A PostgreSQL backup whose manifest names a Supabase extension drills on Supabase's own image, since its dump creates extensions the plain image does not have.

Each plan sets how much of this SafeGrd runs for an organization over 30 days, counted in GB of snapshot drilled or backed up, and the largest snapshot a sandbox drill takes:

PlanRuns on SafeGrdLargest snapshotAt once
Free2 GB a monthIn memory, within the plan’s storage1
Starter80 GB a month5 GB1
Growth250 GB a month10 GB2
Scale800 GB a month25 GB3

A drill or a backup is charged its snapshot’s size in whole GB, and the machine is given a fixed slice of time per GB, so a restore that stalls ends at its deadline and costs no more than it was charged. A backup is charged an estimate until it has run, then its real size. Past the plan’s amount the next drill waits until enough of the 30 days has rolled off, and the console says when. Drills on your own hosts are never counted. A snapshot larger than the plan’s sandbox limit is refused before anything is downloaded, with the alternative named: drill it on a host with drill.sandbox_url.

Recovery point and restore time

The console shows four figures for each surface, measured from its snapshots and drills. The first three are what a security questionnaire calls the recovery point objective (RPO), and restore time is the recovery time objective (RTO). They are measured from this surface’s record. SafeGrd sets no target and guarantees none.

FigureHow it is measured
VerifiedThe time since the newest snapshot that a drill restored, and how deep that drill went. The console leads with this one, because it counts only backups that came back.
Latest backupThe time since the newest snapshot that was written. A snapshot Threat Shield flagged still counts, since the data is there, and the console marks it.
Worst gap, 30 daysThe longest time between two backups in the last 30 days, counted from the last backup before the window, and including the time since the newest one. A surface backed up for less than 30 days is measured from its first backup, and the console says so. It is shown against the schedule. A gap over twice the schedule is shown in red, the point where the overdue alert fires.
Restore timeThe median (p50) and the 95th percentile (p95) of how long this surface’s passed drills of the last 90 days took, with how many drills that is, the usual size of the snapshots they restored, and where they ran.

A drill is timed from before the download, through decryption, to the end of the restore. Restore time counts only drills that wrote the data out: a sandbox drill, which loads a database, and a restore of a repository, which writes the files. An in-memory drill is left out. For a database it never loads the data or builds an index, and for a files or mail archive it reads the archive without writing a file, so its time is shorter than any real restore. SQLite is the exception: its in-memory drill writes the database file out and opens it, which is the whole of a SQLite restore, so it counts.

A surface with no drill that counts says why in place of a number: the free plan drills in memory, the host could not run a sandbox drill (and the console names the reason), or no drill has passed yet.

A drill that ran on SafeGrd’s runner is counted separately from one on your host. The runner’s machine is not yours, so its time can differ from a restore on your own hardware. For that figure, time a restore yourself: Rehearse it before you need it.

The same figures are in GET /api/v1/orgs/{org_id}/recovery and in the Evidence Pack, each with the date it was measured.

Verifying from another machine

A drill does not have to run where the backup was taken. Point the verifier at the sink directly and give it the key:

safegrd verify --snapshot snap-1b094c0af8 \
  --s3-bucket my-worm-bucket --s3-endpoint https://s3.example.com \
  --s3-access-key "$AWS_ACCESS_KEY_ID" --s3-secret-key env:AWS_SECRET_ACCESS_KEY \
  --key-path ./daemon.key

This is the disaster-recovery case: a separate machine, with read access to the bucket and your key, rebuilding the data from nothing. The recovery runbooks take you through it step by step, and testing that a backup restores explains the same checks with plain pg_restore and mysql.