Multi-region and disaster recovery
Every disaster recovery conversation starts by extracting two numbers from the business: RTO, how long until service is restored, and RPO, how much data you can afford to lose. Those two numbers set the budget, and the correct first response to "we cannot lose any data and cannot be down" is to ask for them, then show the cost curve.
The most useful thing to say in this area is that untested failover is fiction, and that failback is harder than failover. The hard part of a region evacuation is not the mechanics, it is deciding to do it.
What this chapter covers
- [done] The multi-region write path
- [done] Active-active conflict resolution
- [done] Cell-based architecture
- [done] RTO and RPO, extracted and priced
- [done] The DR ladder and global routing
- [done] Declaring failover, the runbook, failback and split-brain
- [done] Data residency and the dependency audit
- [done] Backup hygiene
Source: §32.