Docs / Daemon

Daemon

Automating backups removes the risk of missed manual runs. The daemon reads one config file, evaluates which surfaces are due, takes a per-surface lock, runs backups, and reports host health to the console so you are alerted if a host becomes unresponsive.

Run it

# resident daemon
safegrd daemon run
# one reconciliation pass and exit, for cron or a Kubernetes Job
safegrd daemon run --once
# what every surface is doing, and when each last succeeded
safegrd daemon status

daemon status --json provides structured output showing detailed execution timestamps and status across all configured surfaces.

Install it as a service

# systemd on Linux, launchd on macOS
safegrd daemon install
# as a user agent rather than a system service
safegrd daemon install --user
# print the unit and install nothing, to review or commit it
safegrd daemon install --print

The service runs the same binary with the config you installed it from (the one --config names, or ~/.safegrd/config.yaml), so it backs up exactly what safegrd daemon run does in your shell. It reads the config when it starts: after adding a surface, by hand or with safegrd claim (details), run safegrd daemon restart. Its state and locks live in ~/.safegrd of the user who ran install: /root/.safegrd for a system service installed with sudo. A system service sees the filesystem read-only except that directory and a local sink, which install creates for it if they do not exist yet.

A user service stops when you log out unless you run loginctl enable-linger $USER once. safegrd daemon restart restarts it (add --user for a user service). safegrd daemon uninstall removes the unit file and tells you the command that stops a service still running. Run the service one way or the other, not both: two daemons writing the same sink as different users leave directories one of them cannot write to.

Schedules

interval is how often the daemon wakes to look for work, not how often a backup runs. Each surface's schedule decides that. A jitter of about ten per cent is applied so a fleet of hosts does not stampede the same bucket on the minute.

FormMeaning
@hourly, @daily, @weeklyThe usual shorthands.
6h, 90mA plain duration between runs.
2dA whole number of days.

The minimum interval is one hour. With compliance-mode Object Lock enabled, each snapshot remains immutable until retention expires, so frequent runs increase storage usage (for example, a 30-minute interval with 30-day retention retains 1,440 snapshots). safegrd config validate enforces this one-hour minimum. Standard interval notation is supported; cron-style syntax is not accepted.

An owner or admin can also set a surface's schedule and retention from the console, with Edit on its row under Nodes. The daemon picks them up at its next check-in and runs them in place of the config's, until they are cleared in the console; the config file on the host is not changed. The daemon says so when they change and once each time it starts, and safegrd daemon status marks such a schedule (console). The console shows a setting as waiting until the daemon reports that it runs it.

The daemon block

KeyWhat it does
intervalPoll interval. Default 5m.
state_dirState, locks and logs. Never put this on a network filesystem: the in-progress lock depends on working flock semantics.
metrics_addrAccepted, not yet acted on: there is no metrics endpoint.

Surfaces run one at a time. A failed backup is tried 3 times in all, 5 minutes apart, then the daemon waits for the surface's next scheduled run. Logs are plain text on stdout and stderr, which systemd and launchd keep.

Credentials the daemon needs

A daemon cannot be prompted for a password. Each surface's credential block says where its secret comes from instead of carrying it: from: env reads the environment variable in name, from: command runs your own secret manager's run and reads the value from its output, and from: file reads path. None of these sends a secret to SafeGrd. from: safegrd is the other choice: SafeGrd holds it, encrypted, and gives it only to this host when the surface backs up.

Before and after a backup

A surface can run a command of yours on either side of each backup, for example to flush a cache or pause a writer and resume it afterwards:

pre_backup: "redis-cli BGSAVE && sleep 5"
post_backup: "systemctl reload app"

What the console sees

Each surface in the config is its own node in the console, registered under the host you enrolled the first time the daemon sees it. The host's token speaks for all of its surfaces, so there is one credential per host, and rotating it covers them all. A new surface counts against your plan like any other.

Every tick, the daemon checks in for each surface. That is how the console tells the daemon has stopped from this one surface is failing:

Checks between backups

On each tick, before it checks in, the daemon opens each PostgreSQL, MySQL and MongoDB surface’s database and runs SELECT 1 (or the engine’s ping), and asks the sink whether it answers: a HeadBucket on your own bucket, or a look at the local directory. It reads nothing else and writes nothing. Each check gives up after 10 seconds. A database that goes away at 09:00 is in the console a few ticks later, not at the next scheduled backup.

Some surfaces are not checked. A surface whose credential SafeGrd holds is checked only when it backs up, because the daemon fetches that credential for a backup or a drill and for nothing else. The same goes for one whose credential comes from credential.from: command, which would otherwise run that command every tick. Hosted storage and a bucket whose key SafeGrd holds are not checked either. A surface with nothing to check is shown as its last backup left it.

Back up now, under a surface’s Details in the console, signals the daemon to trigger an immediate backup during its next check-in. To avoid creating redundant immutable objects in retention-locked storage, manual requests within an hour of the previous backup are skipped.

Alerts

Set a Slack, Discord or plain webhook URL under Settings → Alerts in the console, on any plan. Paid plans also email the organization's owners and admins. You hear about failed backups, failed Fire Drills, a drill that cannot run for want of the key, silent hosts, overdue and unreachable surfaces, and each recovery, once each way. Alerts lists every event and what a message contains.

Fire Drills on an unattended host

On a plan with Fire Drills, the control plane tells the daemon when a surface is due for one. The daemon restores the latest snapshot, checks it against the digests recorded when it was taken, and reports the result to your attestation record.

For a Postgres surface, give the drill a scratch database and it restores the rows into it for real and counts them there. Without one, the snapshot is replayed in memory:

  - id: "app"
    type: "postgres"
    credential:
      from: env
      name: "APP_DATABASE_URL"
    drill:
      sandbox_url_env: "APP_DRILL_DATABASE_URL"

The sandbox must be a database used for nothing else. A drill restores into it, so the daemon refuses one that holds any table, refuses one that is the database the surface backs up, and empties it again after every drill. Create it once (createdb app_drill) and give the daemon a role that owns it.

On a plan with sandbox drills, a Postgres surface with no drill.sandbox_url gets a scratch database from the host. When the host has the PostgreSQL server installed (initdb and postgres, of the source database's major version or newer), the daemon creates a throwaway cluster under its state directory, restores the snapshot into it, and deletes the cluster when the drill ends. Docker is not needed.

Look at a failed drill's sandbox

With drill.keep_failed_sandbox: true, a drill that restored the snapshot and then failed a check leaves the sandbox as it was for 24 hours. The daemon prints where it is, with the password left out:

Warning: Surface app: sandbox kept at postgres://drill:xxxxx@db.internal/app_drill until 2026-10-06T09:25:25Z; the first drill after that empties it.

A sandbox_url database is left alone until then: the next drill is held back, and the first one after the 24 hours empties it. A throwaway cluster keeps running on its socket until the daemon removes it. Everything stays on the host. A restore that fails is rolled back, so it leaves nothing to keep, and the error is in the daemon's log and the console.

On Debian and Ubuntu, install the server with apt install postgresql-16 (postgresql-client has no server). The package also starts a server of its own, which the daemon does not use: systemctl disable --now postgresql stops it. On macOS, brew install postgresql@16. When the daemon cannot start a cluster (no server, one older than the database, too little disk, or an extension the snapshot uses that the server lacks), the drill runs in memory and the daemon's log says what to install.

The console shows why a drill was not the one your plan includes, under the surface: drilled in memory for want of a sandbox, or not run because the host lacks the disk. A drill that is not run is not a failed backup. The daemon tries again at the next drill, and you are alerted once.

Disk the daemon uses on the host

Database and file backups stream to storage without a copy on the host. Four things do write to the host's disk, and the daemon checks for room before each, leaving 5% of the filesystem free for everything else on it:

A drill needs the private key. The daemon runs one only where the key already is (a key_path, SAFEGRD_PRIVATE_KEY, or a SafeGrd-managed key). It never copies the key to a host. Every surface in an organization is sealed to one key, so a key on every host would let one compromised host read every other host's backups. A host without the key reports Drill blocked: no key, and you are alerted. You can then run safegrd verify wherever you keep the key.

A backup never depends on the control plane

If the control plane is unreachable, the daemon keeps running scheduled backups and writing them to your storage. Each snapshot is in your bucket and restores, and safegrd list shows it. The daemon prints that the record was not sent, keeps it, and sends it on the first tick the control plane answers, so the console lists the backup then. safegrd daemon status shows how many records are waiting. A report the control plane refuses (an expired token, say) is not sent again: it is printed to stderr and shown as not recorded. A backup run by hand with safegrd backup keeps no state, so its record is not sent again; --surface goes through the daemon's state and is.