Multi-Region & Disaster Recovery
Read a little, play a little. No scary maths, and no rush.
One datacentre is a bet
Your application runs in one region. On Thursday that region has a bad afternoon — a control-plane outage, a network misconfiguration, a power event. Requests fail, and for as long as the region is down your product does not exist. This is not a hypothetical; a single-region design is a bet that one region never has a bad day.
The fix costs money and adds latency to every request, so the decision starts with numbers, not architecture.
The two promises: RPO and RTO
Before designing anything, get two numbers agreed with whoever owns the business:
- RPO (Recovery Point Objective): how much data can you afford to lose? "Five minutes" means backups and replication must be at most five minutes behind. RPO 0 requires synchronous replication to a second region.
- RTO (Recovery Time Objective): how long can you be down? "One hour" means you do not need a hot standby — a tested restore from backup can qualify.
These are commitments. A design that promises RPO 0 and RTO 1 minute is a much more expensive design than one promising RPO 15 minutes and RTO 4 hours, and both may be correct. Interviewers reward the reasoning more than the tool choice.
Two shapes, two tradeoffs
Active-passive: one region serves everything; the other stands by, maybe warm. Cheap, simple, and the failover is a real event with real downtime — DNS or traffic-manager changes take minutes, and untested failovers fail exactly when you need them. Minimise RTO with warm standby, not cold.
Active-active: several regions serve live traffic simultaneously. There is no failover wait, so the RTO can be near zero. The cost is permanent: every write now has to be reconciled between regions, and "eventually consistent" leaks into places a product manager will not enjoy — two users in different regions can briefly see different balances.
Replication lag is the honest cost
The thing that breaks designs here is that replicas are not instant. Under load, a secondary can fall seconds to minutes behind. So during a failover you are choosing: serve slightly stale data from the survivor, or refuse writes until it catches up. Payments, bookings and inventory usually take the second option; feeds and dashboards do not. State this trade-off out loud rather than claiming a seamless failover.
Test it, or it does not exist
An untested failover is a hypothesis. Schedule regular game days — trigger the failover in a staging account, time how long it actually takes, and confirm the runbook the on-call engineer will read at 3am is accurate. The gap between "documented RTO" and "measured RTO" is where incidents go to become outages.
Remember this
- Decide RPO and RTO with the business before picking an architecture.
- Active-active removes downtime but makes every write eventually consistent.
- Active-passive is cheaper and simpler, and its failover must be exercised to be real.
- Replica lag means failover trades stale reads against refusing writes.
- Measure your failover time. An untested RTO is a guess.
Check your understanding
2 questions · correct answers earn XP once each
My notes
Saved in this browser. Highlight a line above and save it, or write it in your own words.
Nothing saved yet. Your highlights will live here.
References
Finished reading?
Ticking it here also ticks the chapter in the sidebar, the section count and your streak — it is all one number.
Related chapters
Spotted a mistake or want a topic covered? Report an issue