RPO and RTO for a database backup
Two numbers describe how well a backup protects a database. The recovery point objective (RPO) is how much recent data you can afford to lose. The recovery time objective (RTO) is how long you can afford to be down while you restore. A security questionnaire, an auditor or a customer's contract will ask for both.
Both are objectives: targets you choose. What matters is how they compare with what you measure, the recovery point and the restore time your backups actually give you.
Recovery point (RPO)
If the database were lost now, you would restore the newest backup. Everything written since that backup was taken is gone. The recovery point is the age of that backup, so it is the amount of data, measured in time, that a failure costs you.
With a backup every 24 hours, the recovery point is anywhere from a few minutes to 24 hours, depending on when the failure lands. An RPO of 24 hours is a promise about the worst case, so the number to watch is the longest gap between two backups. One skipped night turns a 24-hour gap into 48 hours.
The longest gap is the recovery point you achieved
A schedule says what should happen. The gaps between the backups that were actually written say what did. Take every backup in the last 30 days, in order, and find the longest stretch between two of them, including the time since the newest one. That stretch is the most data a failure would have cost you at the worst moment of the month. If it is longer than your RPO, the target was missed.
What sets it for a database
- A logical dump (
pg_dump,mysqldump,mongodump) restores to the moment the dump began. The recovery point can be no better than the time between dumps. - Point-in-time recovery archives the write-ahead log (WAL in PostgreSQL, the binary log in MySQL) between full backups and replays it up to a chosen moment. The recovery point drops to seconds or minutes, at the cost of keeping and testing the log archive. pg_dump, pg_basebackup or pgBackRest compares the two approaches.
- A backup that was taken but cannot be restored does not count. If the newest good backup is a week old, the recovery point is a week, however many backups were written since.
Restore time (RTO)
The restore time is how long it takes from deciding to restore until the database is back and holding the data. For a database backup it is the sum of:
- getting a machine and a database server to restore into;
- downloading the backup;
- decrypting and decompressing it;
- loading it: inserting the rows and rebuilding every index and constraint, which is usually the longest step for a logical dump;
- checking that what came back is complete;
- pointing the application at it.
The only way to know this number is to time a real restore of a real backup, at its real size, on the kind of machine you would use. A restore that reads the backup without loading it into a database skips the slowest step, so its time is shorter than any real restore and tells you nothing about RTO. Time it again as the database grows: a restore that took 10 minutes at 5 GB will not take 10 minutes at 50 GB.
Choosing the targets
Start from what a loss costs. Ask the people who own the data two questions: how many hours of writes could we re-enter or live without, and how many hours can the application be down. The answers are your RPO and RTO. Then check they are affordable:
- An RPO of a day is met by a nightly dump. An RPO of an hour needs hourly backups or a log archive. An RPO of minutes needs point-in-time recovery or a replica.
- An RTO shorter than your measured restore time needs a different approach: a smaller database to restore first, a physical backup instead of a dump, or a standby that is already loaded.
Lowering each one
To lower the recovery point:
- Back up more often, and alert when a backup is late. A missed backup is what makes the longest gap long.
- Add point-in-time recovery when the time between backups is still too long.
- Test the backups, so that the newest backup is also a good one.
To lower the restore time:
- Keep the backup in the region where you would restore it.
- Restore in parallel (
pg_restore --jobs) on a machine with fast disks. - Write the restore down as a runbook and rehearse it, so nobody spends the outage rediscovering the steps.
How SafeGrd measures them
The SafeGrd console shows four figures for each surface, under Fire Drills, measured from its own backups and drills. SafeGrd sets no target: the figures are there to compare with yours.
- Verified: the age of the newest backup a drill restored, and how deep that drill went. This is the recovery point counting only backups that came back.
- Latest backup: the age of the newest backup written.
- Worst gap, 30 days: the longest gap between backups over the last 30 days, as above, shown against the surface's schedule. A gap more than twice the schedule is shown in red, which is when the overdue alert fires.
- Restore time: the median and 95th percentile of how long the surface's passed drills took over the last 90 days, counting only drills that loaded the data into a database or wrote the files out.
Recovery point and restore time defines each figure exactly. The runbooks cover timing a restore on your own hardware, and pricing lists how often, and how deep, each plan drills.