System Design Problem

Design a Backup and Disaster Recovery System

Commonly Asked By:AWSGoogleMicrosoftDropbox

Interview Setup

Interview Prompt

Design a backup and disaster recovery system for 500 TB of production data with hourly incremental backups (25 TB/day change rate), 15-minute RTO, and ~5 PB total backup storage across hot/warm/cold tiers.

Clarifying Questions (ask before designing)

QuestionWhy it matters
RPO and RTO per service tier: payments vs analytics?Payments: RPO 0, RTO 5 min. Analytics: RPO 1 hour, RTO 4 hours. Drives cost.
Backup to same region or cross-region by default?
  • Regional failure destroys same-region backups
  • cross-region adds egress cost.
Restore testing frequency: automated or manual DR drills?
  • Untested backups are Schrödinger's backups
  • quarterly drills catch silent corruption.
Deduplication and compression for 5 PB storage?500 TB full x weekly + 25 TB/day incremental without dedup exceeds 10 PB quickly.

Scope

In scope

  • RPO/RTO definitions
  • Incremental backups
  • Cross-region replication
  • Failover orchestration
  • DR drills
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Detailed frontend/UI pixel implementation
  • Org structure, staffing, and hiring plan

Functional Requirements

Start by asking your interviewer to define RPO and RTO, because every backup tier flows from those two numbers. Confirm which data classes are critical (databases, object storage, config) and whether automated failover is in scope.

  • Full backups: Complete copy of all data (databases, object storage, config)
  • Incremental backups: Only changes since last backup
  • Point-in-time recovery (PITR): Restore to any second within retention window
  • Cross-region replication: Backups stored in geographically separate region
  • Backup scheduling: Configurable policies (hourly incremental, daily full, weekly archive)
  • Restore testing: Automated periodic restore verification
  • Multi-tier storage: Recent on fast storage, older on cold/archive
  • RPO and RTO enforcement
  • Disaster recovery runbook: Automated failover to DR region

Non-Functional Requirements

RPO and RTO are not merely bullet points; they represent the core operational contract your interviewer will stress-test. A 1-minute RPO rules out daily full backups alone, while a 15-minute RTO pushes you past backup-and-restore toward pilot light or warm standby architectures.

  • RPO: < 1 minute for critical data
  • RTO: < 15 minutes for critical services
  • Durability: 11 nines for backup data
  • Consistency: Backup must represent a consistent point-in-time snapshot
  • Encryption: All backups encrypted; key management via KMS
  • Compliance: Retention policies per regulation (GDPR)

Capacity Estimations

Backup storage cost often exceeds production, so calculate incremental vs full backup math and factor in cross-region replication egress bandwidth. Retention tiers (hot to cold to glacier) maintain 11-nines durability without exhausting operational budgets.

MetricCalculationValue
Total production dataGiven500 TB
Daily change rateGiven5% → 25 TB/day incremental
Full backup sizeGiven500 TB (compressed: ~200 TB)
Full backup frequencyGivenWeekly
Incremental backup frequencyGivenHourly
RetentionGiven30 days hot, 1 year warm, 7 years cold
Total backup storageGiven~5 PB (with retention + dedup)
Restore throughput needed (RTO=15min)Given556 GB/sec

Architecture Diagram

Walk the diagram top-down: production workloads emit change data, backup agents capture full and incremental snapshots, and object storage in a separate region holds the durable copies. The restore orchestrator replays WAL segments for point-in-time recovery when RPO demands sub-hour precision.

DR strategy tiers map directly to cost: backup-and-restore is the most economical with RTO measured in hours, whereas warm standby costs significantly more while achieving a 15-minute RTO. The tier table illustrates these trade-offs before diving into PITR mechanics.

Cross-region replication is non-negotiable for Tier 1 data because a regional disaster that destroys same-region backups defeats the purpose of disaster recovery.

Loading...

In the room

Ask for RPO and RTO per service tier before drawing components, because payments at RPO 0 require synchronous replication whereas analytics at RPO 1 hour can leverage hourly incrementals.

DR Strategy Tiers

StrategyRPORTOCostDescription
Backup & RestoreHoursHours$Restore from S3 on demand
Pilot LightMinutes30 min$$Minimal infra, scale on failover
Warm StandbySeconds15 min$$$Reduced capacity, always running
Active-Active0~0$$$$Both regions serve traffic

Component Deep Dives

Point-in-time recovery (PITR) is a core deep-dive topic: combining base backups with continuous WAL replay delivers second-level precision. Pair it with backup consistency guarantees to defend restore correctness under concurrent writes.

PITR (Point-in-Time Recovery) for PostgreSQL

Base backups combine with continuous write-ahead log replay to provide second-level recovery precision without storing every intermediate state.

Base backups are captured via pg_basebackup alongside continuous WAL archiving to S3. To restore, the system deploys the latest base snapshot taken prior to the target time, then replays WAL segments up to the exact transaction timestamp.

Backup Consistency

Backup consistency determines whether your restore is trustworthy: crash-consistent snapshots are fast but require WAL replay, while application-quiesced backups are slower but ensure clean states.

Consistency options range from database-native serializable dumps and filesystem-level frozen LVM snapshots to cloud block volume snapshots and application-level write quiescing.

Backup Layout (S3)

S3 key layout structures restore throughput and lifecycle tiering, placing hot backups on fast tiers for rapid restore while transitioning older archives to cold storage for long-term compliance.

s3://backups/
+-- postgresql/
¦   +-- 2026-03-14/
¦   ¦   +-- full/base.tar.gz.enc           (weekly full backup)
¦   ¦   +-- wal/000000010000001A000000FF   (continuous WAL archive)
¦   ¦   +-- manifest.json
+-- cassandra/
¦   +-- keyspace_orders/sstable-*.gz
+-- redis/
¦   +-- rdb-snapshot.rdb.enc
+-- config/
    +-- etcd-snapshot.db

API Design

Backup control-plane APIs are operator-facing rather than user-facing, governing triggers, statuses, restores, and manual failovers. Emphasize that restoration is a complex operation rehearsed quarterly rather than a casual one-click task.

HTTP
POST /api/backups/trigger        → Trigger manual backup {source, type}
GET  /api/backups                → List backups with status
GET  /api/backups/{id}/status    → Backup job status
POST /api/restore                → Initiate restore {backup_id, target}
POST /api/dr/failover            → Initiate DR failover (manual trigger)
GET  /api/dr/status              → DR region health, replication lag

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

PostgreSQL (Backup Metadata: Control Plane)

SQL
CREATE TABLE backup_jobs (
    backup_id      UUID PRIMARY KEY,
    source_system  TEXT NOT NULL,
    backup_type    TEXT NOT NULL,
    status         TEXT DEFAULT 'running',
    storage_path   TEXT,
    size_bytes     BIGINT,
    checksum       TEXT,
    started_at     TIMESTAMPTZ DEFAULT NOW(),
    completed_at   TIMESTAMPTZ,
    retention_tier TEXT DEFAULT 'hot',
    expires_at     TIMESTAMPTZ,
    encrypted      BOOLEAN DEFAULT TRUE
);

CREATE TABLE restore_jobs (
    restore_id     UUID PRIMARY KEY,
    backup_id      UUID REFERENCES backup_jobs(backup_id),
    target_system  TEXT NOT NULL,
    restore_point  TIMESTAMPTZ,
    status         TEXT DEFAULT 'running',
    validated      BOOLEAN DEFAULT FALSE
);

Fault Tolerance

Automated Failover Runbook

Trigger: Primary region health check fails for > 60 seconds

Automated sequence:
  T=0s:   Health check failure detected (Route53 health checker)
  T=5s:   Alert PagerDuty + Slack channel
  T=10s:  Fence old primary (revoke DNS, block writes via security group)
  T=15s:  Promote DR PostgreSQL replica to primary (pg_promote(), < 5s)
  T=20s:  Update service discovery (Consul/etcd) → DR endpoints
  T=25s:  Scale DR auto-scaling groups to full capacity
  T=30s:  DNS failover: Route53 switches to DR region
  T=60s:  Traffic flowing to DR region
  T=120s: Run smoke tests
  T=180s: Declare DR active

Total RTO: ~3 minutes (automated) vs 30+ minutes (manual)

CRITICAL: Fence old primary BEFORE promoting DR
  If old primary comes back online → split-brain → data corruption

Restore Testing

"A backup that has never been tested is not a backup"

Weekly automated restore test:
1. Pick random recent backup
2. Restore to isolated environment (separate VPC)
3. Run validation checks:
   a. Row counts match expected (within 0.1%)
   b. Checksums of critical tables match
   c. Application can connect and query
   d. Measure actual restore time vs RTO target
4. Report results → dashboard + alert on failure
5. Tear down isolated environment

Immutable Backups (Ransomware Protection)

Solution: S3 Object Lock (WORM: Write Once Read Many)
  COMPLIANCE mode: NOBODY can delete (not even root/admin) for 30 days

  Combined with:
  - Cross-account replication: backups in separate AWS account
  - MFA delete: require MFA token to delete any backup
  - Separate KMS keys: backup encryption keys in different account

Crypto-Shredding (GDPR Right to Erasure from Backups)

Problem: User requests data deletion, but data exists in 90 days of backups

Crypto-shredding:
  1. Each user's data encrypted with per-user key before backup
  2. User deletion request → delete their encryption key from KMS
  3. Backup data still exists but is unreadable (key is gone)
  4. Effectively deleted without modifying backup files
  5. Immutable backups + GDPR compliance = both satisfied

Additional Considerations

  • Cost optimization: Dedup + compression + lifecycle policies reduce storage 5-10x
  • Chaos engineering: Regularly test DR failover (GameDay exercises)
  • Multi-database consistency: Coordinated snapshot across systems (or accept small inconsistency window)
  • Backup bandwidth: Use incremental + local snapshot + async replication
  • Monitoring: Track backup success rate, size trends, replication lag, last successful restore test date

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • RPO/RTO per tier (8 min)
    • Full weekly + hourly incremental (9 min)
    • 3-2-1 backup rule (8 min)
  • Start by defining RPO (maximum acceptable data loss) and RTO (maximum acceptable downtime), because every architectural choice flows from these two numbers.
  • Apply the 3-2-1 backup rule (3 copies, 2 media types, 1 offsite): immutable and object-lock backups protect against ransomware that encrypts live systems.
  • Incremental backups and log archiving for databases: schedule full snapshots weekly, incrementals daily, and continuous log shipping for point-in-time recovery.
  • Active-passive DR with documented failover runbooks: execute DNS and load balancer switches, promote replicas, and validate data integrity before accepting traffic.
  • Quarterly restore tests are non-negotiable, because a backup never tested is a backup you cannot trust during an actual outage.
  • Crypto-shredding for GDPR: encrypt sensitive data with per-user keys and destroy keys upon erasure requests, rendering backup records cryptographically unreadable without rewriting storage archives.
  • Common pitfall: backing up without ever performing a full restore test, discovering corrupted or incomplete backups only during an active disaster.

Engineering Trade-offs

RPO vs RTO: The Fundamental Trade-off

RPO 0 (zero data loss):
  Requires: synchronous replication to DR site
  Sync replication: primary WAITS for DR to confirm write
  Latency cost: +20-100ms per write
  Use for: Financial transactions, ledger systems, payment data

RPO 1 minute:
  Requires: async replication + WAL shipping every minute
  Primary writes freely → ships WAL logs to DR every 60 seconds
  Use for: E-commerce orders, user data

RPO 1 hour:
  Requires: hourly backups to S3 (snapshot + incremental)
  Cheapest approach: no replication, just scheduled backups
  Use for: Analytics, non-critical data

RTO → cost relationship:
  RTO = 4 hours: ~$500/month (just S3 storage)
  RTO = 15 minutes: ~$10,000/month (duplicate infrastructure)
  RTO = 30 seconds: ~$50,000/month (full infra in 2+ regions)

Split-Brain Prevention During Failover

1. Fencing (STONITH, Shoot The Other Node In The Head):
   Before promoting DR: revoke IAM write permissions, block traffic, shut DB
   Only AFTER fencing succeeds → promote DR

2. Witness/Quorum:
   Deploy "witness" node in third region. Failover requires 2-of-3 agreement.

3. Lease-based leadership:
   Primary holds lease in distributed lock, expires every 30s
   If primary can't renew → safe to promote

4. Epoch-based writes:
   Every write includes epoch number. On failover: new primary increments epoch.
   Old primary's epoch < current: writes rejected.

Failback: Returning to Primary After DR

Step 1: Rebuild the primary as a replica. Step 2: Validate integrity through row counts and checksums. Step 3: Execute planned failover during a maintenance window. Step 4: Resume post-failback monitoring. Failback duration typically requires 2 to 12 hours depending on total data volume. Rushing failback without full synchronization risks losing transactions processed on the DR site during the outage.

Active-Active Multi-Region

Instead of active-passive, run BOTH regions as active simultaneously.

Why it's hard:
  Conflict resolution: Last-write-wins (LWW), application-level merge, CRDTs, region-owned data
  Referential integrity: route user + all related entities to same region

When active-active makes sense:
  ✓ Read-heavy workloads (global CDN-like caching)
  ✓ Partition-able data (each user assigned to home region)
  ✓ CRDT-compatible operations (counters, sets, append-only logs)

When active-passive is better:
  ✓ Strong consistency required (financial, inventory)
  ✓ Complex transactions spanning multiple entities
  ✓ Write-heavy workloads
  ✓ Simpler operations

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...