Interview Setup
Interview Prompt
Design a backup and disaster recovery system for 500 TB of production data with hourly incremental backups (25 TB/day change rate), 15-minute RTO, and ~5 PB total backup storage across hot/warm/cold tiers.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| RPO and RTO per service tier: payments vs analytics? | Payments: RPO 0, RTO 5 min. Analytics: RPO 1 hour, RTO 4 hours. Drives cost. |
| Backup to same region or cross-region by default? |
|
| Restore testing frequency: automated or manual DR drills? |
|
| Deduplication and compression for 5 PB storage? | 500 TB full x weekly + 25 TB/day incremental without dedup exceeds 10 PB quickly. |
Scope
In scope
- RPO/RTO definitions
- Incremental backups
- Cross-region replication
- Failover orchestration
- DR drills
- Capacity estimation with shown math
Out of scope (state explicitly)
- Detailed frontend/UI pixel implementation
- Org structure, staffing, and hiring plan
Functional Requirements
Start by asking your interviewer to define RPO and RTO, because every backup tier flows from those two numbers. Confirm which data classes are critical (databases, object storage, config) and whether automated failover is in scope.
- Full backups: Complete copy of all data (databases, object storage, config)
- Incremental backups: Only changes since last backup
- Point-in-time recovery (PITR): Restore to any second within retention window
- Cross-region replication: Backups stored in geographically separate region
- Backup scheduling: Configurable policies (hourly incremental, daily full, weekly archive)
- Restore testing: Automated periodic restore verification
- Multi-tier storage: Recent on fast storage, older on cold/archive
- RPO and RTO enforcement
- Disaster recovery runbook: Automated failover to DR region
Non-Functional Requirements
RPO and RTO are not merely bullet points; they represent the core operational contract your interviewer will stress-test. A 1-minute RPO rules out daily full backups alone, while a 15-minute RTO pushes you past backup-and-restore toward pilot light or warm standby architectures.
- RPO: < 1 minute for critical data
- RTO: < 15 minutes for critical services
- Durability: 11 nines for backup data
- Consistency: Backup must represent a consistent point-in-time snapshot
- Encryption: All backups encrypted; key management via KMS
- Compliance: Retention policies per regulation (GDPR)
Capacity Estimations
Backup storage cost often exceeds production, so calculate incremental vs full backup math and factor in cross-region replication egress bandwidth. Retention tiers (hot to cold to glacier) maintain 11-nines durability without exhausting operational budgets.
| Metric | Calculation | Value |
|---|---|---|
| Total production data | Given | 500 TB |
| Daily change rate | Given | 5% → 25 TB/day incremental |
| Full backup size | Given | 500 TB (compressed: ~200 TB) |
| Full backup frequency | Given | Weekly |
| Incremental backup frequency | Given | Hourly |
| Retention | Given | 30 days hot, 1 year warm, 7 years cold |
| Total backup storage | Given | ~5 PB (with retention + dedup) |
| Restore throughput needed (RTO=15min) | Given | 556 GB/sec |
Architecture Diagram
Walk the diagram top-down: production workloads emit change data, backup agents capture full and incremental snapshots, and object storage in a separate region holds the durable copies. The restore orchestrator replays WAL segments for point-in-time recovery when RPO demands sub-hour precision.
DR strategy tiers map directly to cost: backup-and-restore is the most economical with RTO measured in hours, whereas warm standby costs significantly more while achieving a 15-minute RTO. The tier table illustrates these trade-offs before diving into PITR mechanics.
Cross-region replication is non-negotiable for Tier 1 data because a regional disaster that destroys same-region backups defeats the purpose of disaster recovery.
In the room
Ask for RPO and RTO per service tier before drawing components, because payments at RPO 0 require synchronous replication whereas analytics at RPO 1 hour can leverage hourly incrementals.
DR Strategy Tiers
| Strategy | RPO | RTO | Cost | Description |
|---|---|---|---|---|
| Backup & Restore | Hours | Hours | $ | Restore from S3 on demand |
| Pilot Light | Minutes | 30 min | $$ | Minimal infra, scale on failover |
| Warm Standby | Seconds | 15 min | $$$ | Reduced capacity, always running |
| Active-Active | 0 | ~0 | $$$$ | Both regions serve traffic |
Component Deep Dives
Point-in-time recovery (PITR) is a core deep-dive topic: combining base backups with continuous WAL replay delivers second-level precision. Pair it with backup consistency guarantees to defend restore correctness under concurrent writes.
PITR (Point-in-Time Recovery) for PostgreSQL
Base backups combine with continuous write-ahead log replay to provide second-level recovery precision without storing every intermediate state.
Base backups are captured via pg_basebackup alongside continuous WAL archiving to S3. To restore, the system deploys the latest base snapshot taken prior to the target time, then replays WAL segments up to the exact transaction timestamp.
Backup Consistency
Backup consistency determines whether your restore is trustworthy: crash-consistent snapshots are fast but require WAL replay, while application-quiesced backups are slower but ensure clean states.
Consistency options range from database-native serializable dumps and filesystem-level frozen LVM snapshots to cloud block volume snapshots and application-level write quiescing.
Backup Layout (S3)
S3 key layout structures restore throughput and lifecycle tiering, placing hot backups on fast tiers for rapid restore while transitioning older archives to cold storage for long-term compliance.
s3://backups/
+-- postgresql/
¦ +-- 2026-03-14/
¦ ¦ +-- full/base.tar.gz.enc (weekly full backup)
¦ ¦ +-- wal/000000010000001A000000FF (continuous WAL archive)
¦ ¦ +-- manifest.json
+-- cassandra/
¦ +-- keyspace_orders/sstable-*.gz
+-- redis/
¦ +-- rdb-snapshot.rdb.enc
+-- config/
+-- etcd-snapshot.dbAPI Design
Backup control-plane APIs are operator-facing rather than user-facing, governing triggers, statuses, restores, and manual failovers. Emphasize that restoration is a complex operation rehearsed quarterly rather than a casual one-click task.
POST /api/backups/trigger → Trigger manual backup {source, type}
GET /api/backups → List backups with status
GET /api/backups/{id}/status → Backup job status
POST /api/restore → Initiate restore {backup_id, target}
POST /api/dr/failover → Initiate DR failover (manual trigger)
GET /api/dr/status → DR region health, replication lagCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
PostgreSQL (Backup Metadata: Control Plane)
CREATE TABLE backup_jobs (
backup_id UUID PRIMARY KEY,
source_system TEXT NOT NULL,
backup_type TEXT NOT NULL,
status TEXT DEFAULT 'running',
storage_path TEXT,
size_bytes BIGINT,
checksum TEXT,
started_at TIMESTAMPTZ DEFAULT NOW(),
completed_at TIMESTAMPTZ,
retention_tier TEXT DEFAULT 'hot',
expires_at TIMESTAMPTZ,
encrypted BOOLEAN DEFAULT TRUE
);
CREATE TABLE restore_jobs (
restore_id UUID PRIMARY KEY,
backup_id UUID REFERENCES backup_jobs(backup_id),
target_system TEXT NOT NULL,
restore_point TIMESTAMPTZ,
status TEXT DEFAULT 'running',
validated BOOLEAN DEFAULT FALSE
);Fault Tolerance
Automated Failover Runbook
Trigger: Primary region health check fails for > 60 seconds Automated sequence: T=0s: Health check failure detected (Route53 health checker) T=5s: Alert PagerDuty + Slack channel T=10s: Fence old primary (revoke DNS, block writes via security group) T=15s: Promote DR PostgreSQL replica to primary (pg_promote(), < 5s) T=20s: Update service discovery (Consul/etcd) → DR endpoints T=25s: Scale DR auto-scaling groups to full capacity T=30s: DNS failover: Route53 switches to DR region T=60s: Traffic flowing to DR region T=120s: Run smoke tests T=180s: Declare DR active Total RTO: ~3 minutes (automated) vs 30+ minutes (manual) CRITICAL: Fence old primary BEFORE promoting DR If old primary comes back online → split-brain → data corruption
Restore Testing
"A backup that has never been tested is not a backup" Weekly automated restore test: 1. Pick random recent backup 2. Restore to isolated environment (separate VPC) 3. Run validation checks: a. Row counts match expected (within 0.1%) b. Checksums of critical tables match c. Application can connect and query d. Measure actual restore time vs RTO target 4. Report results → dashboard + alert on failure 5. Tear down isolated environment
Immutable Backups (Ransomware Protection)
Solution: S3 Object Lock (WORM: Write Once Read Many) COMPLIANCE mode: NOBODY can delete (not even root/admin) for 30 days Combined with: - Cross-account replication: backups in separate AWS account - MFA delete: require MFA token to delete any backup - Separate KMS keys: backup encryption keys in different account
Crypto-Shredding (GDPR Right to Erasure from Backups)
Problem: User requests data deletion, but data exists in 90 days of backups Crypto-shredding: 1. Each user's data encrypted with per-user key before backup 2. User deletion request → delete their encryption key from KMS 3. Backup data still exists but is unreadable (key is gone) 4. Effectively deleted without modifying backup files 5. Immutable backups + GDPR compliance = both satisfied
Additional Considerations
- Cost optimization: Dedup + compression + lifecycle policies reduce storage 5-10x
- Chaos engineering: Regularly test DR failover (GameDay exercises)
- Multi-database consistency: Coordinated snapshot across systems (or accept small inconsistency window)
- Backup bandwidth: Use incremental + local snapshot + async replication
- Monitoring: Track backup success rate, size trends, replication lag, last successful restore test date
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- RPO/RTO per tier (8 min)
- Full weekly + hourly incremental (9 min)
- 3-2-1 backup rule (8 min)
- Start by defining RPO (maximum acceptable data loss) and RTO (maximum acceptable downtime), because every architectural choice flows from these two numbers.
- Apply the 3-2-1 backup rule (3 copies, 2 media types, 1 offsite): immutable and object-lock backups protect against ransomware that encrypts live systems.
- Incremental backups and log archiving for databases: schedule full snapshots weekly, incrementals daily, and continuous log shipping for point-in-time recovery.
- Active-passive DR with documented failover runbooks: execute DNS and load balancer switches, promote replicas, and validate data integrity before accepting traffic.
- Quarterly restore tests are non-negotiable, because a backup never tested is a backup you cannot trust during an actual outage.
- Crypto-shredding for GDPR: encrypt sensitive data with per-user keys and destroy keys upon erasure requests, rendering backup records cryptographically unreadable without rewriting storage archives.
- Common pitfall: backing up without ever performing a full restore test, discovering corrupted or incomplete backups only during an active disaster.
Engineering Trade-offs
RPO vs RTO: The Fundamental Trade-off
RPO 0 (zero data loss): Requires: synchronous replication to DR site Sync replication: primary WAITS for DR to confirm write Latency cost: +20-100ms per write Use for: Financial transactions, ledger systems, payment data RPO 1 minute: Requires: async replication + WAL shipping every minute Primary writes freely → ships WAL logs to DR every 60 seconds Use for: E-commerce orders, user data RPO 1 hour: Requires: hourly backups to S3 (snapshot + incremental) Cheapest approach: no replication, just scheduled backups Use for: Analytics, non-critical data RTO → cost relationship: RTO = 4 hours: ~$500/month (just S3 storage) RTO = 15 minutes: ~$10,000/month (duplicate infrastructure) RTO = 30 seconds: ~$50,000/month (full infra in 2+ regions)
Split-Brain Prevention During Failover
1. Fencing (STONITH, Shoot The Other Node In The Head): Before promoting DR: revoke IAM write permissions, block traffic, shut DB Only AFTER fencing succeeds → promote DR 2. Witness/Quorum: Deploy "witness" node in third region. Failover requires 2-of-3 agreement. 3. Lease-based leadership: Primary holds lease in distributed lock, expires every 30s If primary can't renew → safe to promote 4. Epoch-based writes: Every write includes epoch number. On failover: new primary increments epoch. Old primary's epoch < current: writes rejected.
Failback: Returning to Primary After DR
Step 1: Rebuild the primary as a replica. Step 2: Validate integrity through row counts and checksums. Step 3: Execute planned failover during a maintenance window. Step 4: Resume post-failback monitoring. Failback duration typically requires 2 to 12 hours depending on total data volume. Rushing failback without full synchronization risks losing transactions processed on the DR site during the outage.
Active-Active Multi-Region
Instead of active-passive, run BOTH regions as active simultaneously. Why it's hard: Conflict resolution: Last-write-wins (LWW), application-level merge, CRDTs, region-owned data Referential integrity: route user + all related entities to same region When active-active makes sense: ✓ Read-heavy workloads (global CDN-like caching) ✓ Partition-able data (each user assigned to home region) ✓ CRDT-compatible operations (counters, sets, append-only logs) When active-passive is better: ✓ Strong consistency required (financial, inventory) ✓ Complex transactions spanning multiple entities ✓ Write-heavy workloads ✓ Simpler operations
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.