Backup Strategy & Restore Testing
A backup strategy defines which data and configuration must survive a failure, where protected copies live, and how a service will be recovered. A green backup job proves that a job completed; it does not prove that people can resume work.
TL;DR
- Set recovery objectives with the service owner before choosing schedules.
- Protect application data, configuration, keys, and recovery tooling.
- Separate backup administration and failure domains from production.
- Measure a complete restore, including validation and dependencies.
Quick Example
Illustrative planning record; tailor the values and acceptance criteria to the service.
Core Concepts
RPO and RTO
Recovery point objective (RPO) describes the acceptable data-loss interval. Recovery time objective (RTO) describes the target time to restore service. A backup every 15 minutes does not guarantee a 15-minute recovery point if jobs fail or copies arrive late. Include access, provisioning, transfer, replay, and validation in timing.
Consistency and Dependencies
A crash-consistent copy resembles storage after abrupt power loss. An application-consistent copy coordinates the application so the backup is usable according to its recovery procedure. Related databases and files may need a shared recovery point; restoring them independently can break references.
Copies and Isolation
The traditional 3-2-1 rule uses three copies, two media types, and one off-site copy. Off-site protects against some location failures, while offline or immutable copies and separate administration address different risks. Encryption also requires a recoverable key.
Run a Meaningful Restore Exercise
- Select a service and failure scenario, including loss of its normal administration system.
- Retrieve the approved backup, catalog, keys, and instructions through the emergency access path.
- Restore into an isolated environment that cannot send real notifications or payments.
- Validate data integrity and application behavior with the service owner.
- Record the actual recovery point, elapsed time, missing prerequisites, and follow-up owners.
Test representative older restore points as well as the newest copy. Retention only helps if the required backup chain and keys still exist.
Comparison
Best Practices
Alert on Protection Gaps
Track the age of the newest usable recovery point, not only the last job result. Alert when a workload is missing from the policy or copy lag exceeds its objective.
Bound Retention Deliberately
Define retention and deletion with data owners. Test immutability settings on nonproduction data before applying them, because locked retention can constrain removal.
Common Mistakes
Bad: Keep all backups writable by the same compromised production administrator.
Correct: Separate administrative authority and test protected-copy recovery.
Bad: Restore the database but omit the keys and configuration needed to open it.
Correct: Exercise the full application dependency chain.
FAQ
Is replication a backup?
It is not sufficient by itself. Replication can propagate corruption or deletion; historical, protected recovery points solve a different problem.
How often should restores be tested?
Set frequency from criticality, change rate, and recovery risk. Re-test after material changes to applications, keys, storage, or recovery tooling.
Does SaaS need a backup assessment?
Yes. Establish which deleted or corrupted records the vendor can recover, for how long, and whether independent protection is required.