Docs / Surfaces / Files

Files and directory trees

# a directory tree, incremental
safegrd backup --files /srv/uploads
# with exclusions
safegrd backup --files /srv/app --exclude '*.tmp,node_modules/*'
# one archive per backup instead
safegrd backup --files /srv/app --format tar

A symlink is captured as a symlink rather than followed, so a link pointing outside the tree cannot pull the rest of the filesystem into your backup. FIFOs, sockets and device nodes are left out, and the backup lists each one.

Incremental backups

Files are cut into chunks of about 1 MiB by their content, so a one-byte edit stores one chunk and an insert does not shift the chunks after it. Chunks are compressed, encrypted on the host and gathered into packs of about 32 MiB, which are the objects in storage. A file whose size, times and inode have not changed since the last run is not read again. Every seventh run reads every file anyway, and --rescan does it now.

Runs are grouped into monthly epochs. The first run of each UTC month uploads every file once. Every run after it uploads only the chunks the month does not already hold, plus a few kilobytes of metadata. Nothing is shared between months, so each object is locked once, when it is written, and no lock is ever extended. A new epoch also starts when a host's cache is gone (a rebuilt host), when retention is increased, and on --new-epoch.

# a VM's root, in ~/.safegrd/config.yaml, run by the daemon
surfaces:
  - id: vm-root
    type: files
    roots: ["/"]
    excludes: ["/home/*/.cache"]
    schedule: "@daily"

A root of / stays on one filesystem and leaves out /proc, /sys, /dev, /run, /tmp, /var/tmp, swap files and SafeGrd's own cache. A surface may list several roots, none inside another.

Exclusions are the surface's excludes, or Leave out in the console. Each run says how many entries they left out, and which patterns did it.

Back up a surface now

safegrd backup --surface vm-root

A surface backs up on its schedule. backup --surface takes one now, whatever the schedule, into the same history the daemon writes. backup --files backs up a directory that is not a surface, as a separate history.

Find and restore one file

safegrd find etc/nginx/nginx.conf
safegrd restore --path etc/nginx/nginx.conf --target-dir ./out
safegrd restore --path etc/nginx/nginx.conf --version 2 --target-dir ./out
safegrd restore --snapshot <id> --path 'var/www/**' --target-dir ./out
safegrd find --deleted var/www/uploads/

find lists every kept version of a file. A version is one content: it is first seen in the snapshot that first held those bytes and last seen in the last one that did, so a file unchanged across months is one version. A change of permissions alone is not a new one; restoring a version applies that snapshot's permissions.

etc/nginx/nginx.conf (vm-root)
  #   FIRST SEEN         LAST SEEN          SIZE      SHA-256        SNAPSHOTS
  3   2026-10-02 02:00   (current)          2.1 KiB   9f1c2a7b0d3e   2
  2   2026-09-14 02:00   2026-10-01 02:00   2.0 KiB   4b8e11c0aa52   18 (2 epochs)
  1   2026-09-01 02:00   2026-09-13 02:00   1.9 KiB   77d0e3f5b2c1   13

Versions are numbered from the oldest and printed newest first, and the number is what restore --version takes. With one --path and no --version or --snapshot, restore takes the newest version. Each snapshot is a point in time you can restore to, so the surface's schedule decides how far apart they are: a daily surface gives one point per day. To get a file as it was on a given day, pick the version whose first-seen and last-seen dates bracket that day, or pass that day's snapshot id with --path. A directory pattern selects everything below it, and *, ? and ** match. --deleted lists only files the newest snapshot no longer holds, with the last snapshot that held each, and --json prints the same histories for scripts.

Paths given to find and --path are relative to /. A snapshot of one directory restores that directory's contents into the target, as an archive of it does; a snapshot of / or of several roots restores each entry at its path from /, under the target. Restoring one file reads that file's chunks and the snapshot's metadata, not the rest of the tree. find and every restore run where the private key is, and decrypt on that host. Runbook 5 walks through it on a recovery machine.

On a host in an organization with other hosts, find searches the surfaces this host has the key for and says how many others it did not search. --surface <id> searches one surface, and with a SafeGrd-managed key it fetches that surface's key for the search.

What is locked, and for how long

Each epoch has two lock dates, fixed when it opens. Let T be the longest of retention_days, keep_daily and keep_weekly in days.

A snapshot is kept at most as long as the objects it uses, so the date a backup prints after Immutable until is a date every byte it needs is locked to. Once the later objects expire, the month's first snapshot, the monthly copy, still restores from its own objects.

What it holds

Storage holds the kept monthly copies, about two months of changes, and, at the turn of each month, one more full upload while last month's objects are still locked. A 200 GB tree with 2% daily change and 30-day retention holds about 500 to 650 GB on average, where daily archives would hold about 6 TB. On hosted storage the console warns a week before the month turns if that upload will take you past 80% of your plan.

What the storage's holder can see

File names, contents, directory structure and which content repeats are encrypted on your host. Whoever holds the storage (you, or SafeGrd on hosted storage) can see the size of each pack and how many packs each run wrote, which says how much changed. A single-archive backup shows only each archive's size.

Checking a repository

safegrd check --all --read-data

check confirms every pack a snapshot names is in storage and agrees with its index, every directory listing decodes, and the content digest recomputed from them matches the one recorded at backup time. --read-data also opens every chunk.

Listing, copying and pruning

safegrd list shows incremental snapshots in their own table: the month (epoch) each belongs to, whether its run opened the month or followed it, the size of the whole tree it holds and the bytes its run uploaded. safegrd export copies a repository a whole month at a time, since a snapshot needs every object of its month, and restore --from reads the copy. safegrd prune in your own bucket deletes a month's expired objects once every snapshot that could use them is past its retention, and never from the month holding the surface's newest snapshot.

Restoring an incremental snapshot needs CLI v0.0.10 or later. An older CLI does not see these snapshots; safegrd --version says which you have.

What a restore brings back

As of the CLI release after v0.0.2. Earlier releases restore every directory as 0755 and do not refuse a symlink in the target; safegrd --version says which you have.

PropertyRestored?
ContentsYes. Every regular file is checked against the digest sealed at backup time.
PermissionsYes, for files and directories, whatever the restoring shell's umask. A 0700 directory comes back 0700.
Modification timesYes, for files, directories and symlinks. Restoring a one-directory snapshot into a new target gives the target that directory's mode and time.
SymlinksYes, as symlinks pointing where they pointed.
OwnershipOnly when safegrd restore runs as root. Otherwise files belong to whoever ran it, and the restore says so. Snapshots taken before 24 September 2026 did not record owners.
Hard linksNo. Each linked name comes back as its own copy, so a tree full of hard links restores larger than it was.
Extended attributes, ACLs, SELinux labels, capabilitiesNo. A tree that depends on them (an SELinux-labelled web root, a binary granted setcap) needs them reapplied after a restore.
setuid, setgid and sticky bitsNo. Only the permission bits are restored.
Sparse filesContents yes; sparseness no. Holes are written out in full.
FIFOs, sockets, device nodesNot backed up. The backup lists each one it leaves out.

When it finishes, safegrd restore prints what it restored and what it could not: ownership when it did not run as root, and the properties above that are never restored, so you know what to reapply.

A restore writes everything into a staging directory inside the target first and checks every file against its SHA-256. Only when all of them match are they moved into place; if one does not, nothing is left in the target and the restore names the file. The target must be empty or absent.

What a Fire Drill checks

An incremental snapshot is restored whole into a scratch directory on the host, every file checked against its SHA-256. The drill then reads the restored files back, recomputes the snapshot's content digest from them, and compares it with the digest the remote server recorded at backup time. The scratch directory is removed afterwards.

A single archive is checked in memory: every entry matches the manifest sealed inside the archive (names, types, symlink targets, the file count and byte volume, and a SHA-256 for every file). Nothing is written to disk.

To rehearse your team's own restore procedure, see the recovery runbooks.

One archive per backup

--format tar, or format: tar on a surface, stores each backup as one encrypted tar with a manifest sealed inside it, so SafeGrd's servers never see your file names. It takes one root. Every backup uploads the whole tree, and a restore reads the whole archive. Snapshots taken this way restore the same way whatever the default. On hosted storage, the console can download one as a single file.