Daemon
Automating backups removes the risk of missed manual runs. The daemon reads one config file, evaluates which surfaces are due, takes a per-surface lock, runs backups, and reports host health to the console so you are alerted if a host becomes unresponsive.
Run it
daemon status --json provides structured output showing detailed execution
timestamps and status across all configured surfaces.
Install it as a service
The service runs the same binary with the config you installed it from (the one
--config names, or ~/.safegrd/config.yaml), so it backs up
exactly what safegrd daemon run does in your shell. It reads the config when it
starts: after adding a surface, by hand or with safegrd claim
(details), run safegrd daemon restart. Its state and locks live
in ~/.safegrd of the user who ran install: /root/.safegrd
for a system service installed with sudo. A system service sees the
filesystem read-only except that directory and a local sink, which
install creates for it if they do not exist yet.
A user service stops when you log out unless you run
loginctl enable-linger $USER once. safegrd daemon restart restarts
it (add --user for a user service). safegrd daemon uninstall
removes the unit file and tells you the command that stops a service still running. Run
the service one way or the other, not both: two daemons writing the same sink as
different users leave directories one of them cannot write to.
Schedules
interval is how often the daemon wakes to look for work, not how
often a backup runs. Each surface's schedule decides that. A jitter of
about ten per cent is applied so a fleet of hosts does not stampede the same bucket
on the minute.
| Form | Meaning |
|---|---|
| @hourly, @daily, @weekly | The usual shorthands. |
| 6h, 90m | A plain duration between runs. |
| 2d | A whole number of days. |
The minimum interval is one hour. With compliance-mode Object Lock enabled, each snapshot
remains immutable until retention expires, so frequent runs increase storage usage
(for example, a 30-minute interval with 30-day retention retains 1,440 snapshots).
safegrd config validate enforces this one-hour minimum. Standard interval notation
is supported; cron-style syntax is not accepted.
An owner or admin can also set a surface's schedule and retention from the console, with
Edit on its row under Nodes. The daemon picks them up at its next check-in and runs them in
place of the config's, until they are cleared in the console; the config file on the host
is not changed. The daemon says so when they change and once each time it starts, and
safegrd daemon status marks such a schedule (console). The console
shows a setting as waiting until the daemon reports that it runs it.
The daemon block
| Key | What it does |
|---|---|
| interval | Poll interval. Default 5m. |
| state_dir | State, locks and logs. Never put this on a network filesystem: the in-progress lock depends on working flock semantics. |
| metrics_addr | Accepted, not yet acted on: there is no metrics endpoint. |
Surfaces run one at a time. A failed backup is tried 3 times in all, 5 minutes apart, then the daemon waits for the surface's next scheduled run. Logs are plain text on stdout and stderr, which systemd and launchd keep.
Credentials the daemon needs
A daemon cannot be prompted for a password. Each surface's credential block says
where its secret comes from instead of carrying it: from: env reads the
environment variable in name, from: command runs your own secret
manager's run and reads the value from its output, and from: file
reads path. None of these sends a secret to SafeGrd. from: safegrd
is the other choice: SafeGrd holds it, encrypted, and gives it only to this host when the
surface backs up.
- For a database the secret is the whole connection URL; for a mailbox, the password. A
mailbox that names no credential uses
SAFEGRD_EMAIL_PASSWORD. - What the block names and cannot read is an error for that surface: an unset variable, a failing command or a missing file never falls back to anything else.
- The command runs under
/bin/shas the daemon's user and has 30 seconds. If it fails, that surface fails with the command's error output and the others carry on. It never falls back to an environment variable. - A service does not inherit your login shell. Check with
safegrd doctor, which resolves every surface's credential the way the daemon will. Then check again as the service's user: a secret manager signed in for you may not be for it. - The daemon requires configuration file permissions to be restricted to the owner (0600) to prevent unauthorized users from modifying commands or inspecting credentials.
Before and after a backup
A surface can run a command of yours on either side of each backup, for example to flush a cache or pause a writer and resume it afterwards:
pre_backupmust succeed. If it fails, that backup is not taken and the surface counts a failure, with the hook's own output in the daemon's log.post_backupruns after every attempt, so what the first hook paused is resumed even when the backup failed. It getsSAFEGRD_BACKUP_STATUS(successorfailed) andSAFEGRD_SNAPSHOT_ID.- Both run under
/bin/shas the daemon's user, with 10 minutes each. They are part of the config file, which is one more reason it must be readable only by its owner.
What the console sees
Each surface in the config is its own node in the console, registered under the host you enrolled the first time the daemon sees it. The host's token speaks for all of its surfaces, so there is one credential per host, and rotating it covers them all. A new surface counts against your plan like any other.
Every tick, the daemon checks in for each surface. That is how the console tells the daemon has stopped from this one surface is failing:
- Silent: the daemon has missed three check-ins (and at least fifteen minutes). Nothing is backing that host up.
- Overdue: a surface has not backed up for twice its schedule. This
applies to hosts run from cron with
daemon run --onceas well, which do not count as daemons. - Failing: the last attempt failed, with the error the daemon saw.
- Unreachable: on three ticks in a row, the daemon could not open the surface’s database or reach its sink. The row says since when and the error.
Checks between backups
On each tick, before it checks in, the daemon opens each PostgreSQL, MySQL and MongoDB
surface’s database and runs SELECT 1 (or the engine’s ping), and
asks the sink whether it answers: a HeadBucket on your own bucket, or a look
at the local directory. It reads nothing else and writes nothing. Each check gives up
after 10 seconds. A database that goes away at 09:00 is in the console a few ticks later,
not at the next scheduled backup.
Some surfaces are not checked. A surface whose credential SafeGrd holds is checked only
when it backs up, because the daemon fetches that credential for a backup or a drill and
for nothing else. The same goes for one whose credential comes from
credential.from: command, which would otherwise run that command every tick.
Hosted storage and a bucket whose key SafeGrd holds are not checked either. A surface
with nothing to check is shown as its last backup left it.
Back up now, under a surface’s Details in the console, signals the daemon to trigger an immediate backup during its next check-in. To avoid creating redundant immutable objects in retention-locked storage, manual requests within an hour of the previous backup are skipped.
Alerts
Set a Slack, Discord or plain webhook URL under Settings → Alerts in the console, on any plan. Paid plans also email the organization's owners and admins. You hear about failed backups, failed Fire Drills, a drill that cannot run for want of the key, silent hosts, overdue and unreachable surfaces, and each recovery, once each way. Alerts lists every event and what a message contains.
Fire Drills on an unattended host
On a plan with Fire Drills, the control plane tells the daemon when a surface is due for one. The daemon restores the latest snapshot, checks it against the digests recorded when it was taken, and reports the result to your attestation record.
For a Postgres surface, give the drill a scratch database and it restores the rows into it for real and counts them there. Without one, the snapshot is replayed in memory:
The sandbox must be a database used for nothing else. A drill restores into it, so the
daemon refuses one that holds any table, refuses one that is the database the surface
backs up, and empties it again after every drill. Create it once
(createdb app_drill) and give the daemon a role that owns it.
On a plan with sandbox drills, a Postgres surface with no drill.sandbox_url
gets a scratch database from the host. When the host has the PostgreSQL server installed
(initdb and postgres, of the source database's major version or
newer), the daemon creates a throwaway cluster under its state directory, restores the
snapshot into it, and deletes the cluster when the drill ends. Docker is not needed.
- The cluster listens on a Unix socket only, in a directory only the daemon's user can open.
- It uses under 200 MB of memory whatever the database's size. Restored rows go to disk.
- It needs about as much disk as the database. The daemon checks first, and does not start one that would leave less than 5% of the filesystem free.
- PostgreSQL does not run as root, so a daemon running as root runs the cluster as the user who owns its state directory.
Look at a failed drill's sandbox
With drill.keep_failed_sandbox: true, a drill that restored the snapshot and
then failed a check leaves the sandbox as it was for 24 hours. The daemon prints where it
is, with the password left out:
A sandbox_url database is left alone until then: the next drill is held back,
and the first one after the 24 hours empties it. A throwaway cluster keeps running on its
socket until the daemon removes it. Everything stays on the host. A restore that fails is
rolled back, so it leaves nothing to keep, and the error is in the daemon's log and the
console.
On Debian and Ubuntu, install the server with apt install postgresql-16
(postgresql-client has no server). The package also starts a server of its
own, which the daemon does not use: systemctl disable --now postgresql stops
it. On macOS, brew install postgresql@16. When the daemon cannot start a
cluster (no server, one older than the database, too little disk, or an extension the
snapshot uses that the server lacks), the drill runs in memory and the daemon's log says
what to install.
The console shows why a drill was not the one your plan includes, under the surface: drilled in memory for want of a sandbox, or not run because the host lacks the disk. A drill that is not run is not a failed backup. The daemon tries again at the next drill, and you are alerted once.
Disk the daemon uses on the host
Database and file backups stream to storage without a copy on the host. Four things do write to the host's disk, and the daemon checks for room before each, leaving 5% of the filesystem free for everything else on it:
- A drill of an incremental file surface restores the whole tree under the state directory, then deletes it. Without room it is not run.
- A local sandbox drill needs about the database's size, as above.
- A SQLite backup copies the database to
TMPDIRfirst. Without room the backup fails and says so. PointTMPDIRat a larger disk. - A local sink (
storage.type: local) holds every backup. A backup stops before the disk would drop below 5% free, and fails with what to free.
A drill needs the private key. The daemon runs one only where the key already is (a
key_path, SAFEGRD_PRIVATE_KEY, or a SafeGrd-managed key). It never
copies the key to a host. Every surface in an organization is sealed to one key, so a key
on every host would let one compromised host read every other host's backups. A host
without the key reports Drill blocked: no key, and you are alerted. You
can then run safegrd verify wherever you keep the key.
A backup never depends on the control plane
If the control plane is unreachable, the daemon keeps running scheduled backups and writing
them to your storage. Each snapshot is in your bucket and restores, and
safegrd list shows it. The daemon prints that the record was not sent, keeps it,
and sends it on the first tick the control plane answers, so the console lists the backup
then. safegrd daemon status shows how many records are waiting. A report the
control plane refuses (an expired token, say) is not sent again: it is printed to stderr and
shown as not recorded. A backup run by hand with safegrd backup
keeps no state, so its record is not sent again; --surface goes through the
daemon's state and is.