Alerts
The control plane monitors reports from enrolled hosts to notify you when action is needed. Alerts originate centrally from the control plane rather than from individual hosts, ensuring notifications are still dispatched if a host becomes entirely unreachable.
Where alerts go
| Channel | Plans | How |
|---|---|---|
| Webhook | Every plan, the free floor included | Paste a URL under Settings → Alerts in the console, then press Send a test. A Slack incoming-webhook URL gets Slack's message format and a Discord webhook URL gets Discord's; any other URL gets the JSON below. |
| Paid plans, off until you turn it on | An owner or admin turns on Email alerts to owners and admins under Settings → Alerts. Alerts then go to every owner and admin of the organization, at the address they signed in with. Email points you at the console rather than logging every event: at most one email per kind of alert for a surface in a day, and one per kind across the organization in an hour. The webhook and the console get every event. |
Hosts do not send alerts. An alert: block in a host's
config.yaml is not read, and safegrd config validate warns if one
is set.
The webhook must be a public address. The control plane refuses a URL that points at a private or loopback network, both when you save it and again when it connects, and it does not follow redirects.
What fires
Each condition sends one alert when it starts and one when it clears, not one per check, so a surface that stays broken does not flood the channel.
| Event | When |
|---|---|
| backup_failed | A backup reported a failure, or the daemon starts failing to back a surface up. The daemon tries it 3 times, 5 minutes apart, then waits for the next scheduled backup. |
| backup_recovered | The daemon backs that surface up again. |
| drill_failed | A Fire Drill ran and the snapshot did not restore whole, or the
drill stopped before restoring: the host could not open storage, drill.sandbox_url
was refused, or the daemon stopped during the drill. |
| drill_blocked | A Fire Drill is due and the host cannot run it: it does not hold the private key, or it lacks the disk to restore the snapshot. Posted to the webhook and shown in the console, not emailed. |
| silent | A daemon has missed three check-ins, and at least 15 minutes have passed. Nothing is backing its surfaces up. |
| overdue | A surface has gone twice its schedule without a backup, whether or not its host is checking in. |
| unreachable | On three ticks in a row, the daemon could not open a surface’s database or reach its sink (checks between backups). The alert says since when and the error. The next backup fails unless it answers by then. |
| reachable_again | The daemon opened that surface’s database and reached its sink again. |
| anomaly | Threat Shield flagged a snapshot: a drop in rows (>20%), files (25%) or messages (30%), a dropped table, an empty snapshot, or a missing extension. The alert names the last known-good snapshot to restore from (Threat Shield). |
| hosted_storage_80 | Hosted storage is 80% full of locked backups. On a paid plan backups continue past 100% and the excess is billed per GB-month. On Free or a trial, new backups are refused at 100%. |
| hosted_storage_over | A paid plan’s hosted storage passed its included storage. Backups continue, and the average held above it each billing period is billed per GB-month (pricing). |
| hosted_storage_full | Hosted storage refused a backup: Free or a trial at its included storage, a lapsed plan past its margin, or the server’s per-organization cap. New backups are refused until locks expire or the organization moves to a larger plan or its own bucket; existing backups stay restorable. |
| recovered | A silent host is checking in again, or an overdue surface is backing up again. |
| test | You pressed Send a test. |
Surfaces in an archived project do not alert.
What a webhook receives
A URL that is neither Slack nor Discord receives one JSON object per alert, sent with
Content-Type: application/json and User-Agent: SafeGrd-Alerts/1:
Slack receives {"text": "*title*\ntext"} and Discord
{"content": "**title**\ntext"}. In keeping with SafeGrd's zero-custody model,
alerts contain only operational metadata and never include snapshot payload data.
A send that fails is tried again after 1, 5 and 15 minutes, then after 1, 3 and 6 hours. An answer of 4xx,
other than 408 or 429, is not retried. An alert can arrive more than once, for example when SafeGrd restarts
during a send; id is the same on every attempt, so a receiver that must act once can skip an
id it has seen.