System Design Problem

Design a Blob Storage System (like S3)

Commonly Asked By:AWSGoogleMicrosoftBackblaze

Interview Setup

Interview Prompt

Design an S3-style blob storage system storing 100+ trillion objects at exabyte scale, serving 100M+ requests/sec with 11 nines durability and strong read after write consistency.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Consistency model: strong read after write or eventual?The design chooses strong read after write so a successful object write is immediately visible to GET, HEAD, and LIST in the authoritative namespace. Eventual consistency would require clients or services to tolerate stale reads and possibly use retries or read repair.
Replication (3x full copies) or erasure coding for durability?Replication = 3x storage cost and simpler fast recovery. A Reed-Solomon design such as 6+3 uses 1.5x storage overhead and slower reconstruction. The exact coding profile is a system design choice, not an assumption about S3 internals.
Max object size and multipart threshold?This design uses a 5 TB maximum object size and a 100 MB threshold as interview assumptions. A single S3 PUT supports up to 5 GB, while multipart upload handles larger objects.
Pre-signed URLs for direct client upload: in scope?Signed data-plane URLs offload large payload transfer from the API control plane. The storage nodes still validate the signature and authorization context.

Scope

In scope

  • Object metadata store
  • Erasure coding vs replication
  • Multipart upload
  • Garbage collection
  • Pre-signed URLs
  • Consistency model (strong read after write)

Out of scope (state explicitly)

  • Client desktop/mobile app implementation
  • End-user file preview rendering for every format
  • Building raw block storage hardware

Functional Requirements

Start by confirming object size range and consistency model with your interviewer. Ask about multipart uploads, versioning, lifecycle tiers, and pre signed URL requirements.

  • Upload objects (blobs) from 1 byte to a 5 TB design limit. Single PUT operations handle up to 5 GB, whereas larger objects use multipart uploads.
  • Download objects via GET with support for range requests (partial downloads)
  • Delete objects with immediate namespace visibility change and lazy physical storage reclamation
  • List objects in a bucket with prefix filtering and pagination
  • Multipart upload: upload large objects in parts, assemble on completion
  • Versioning: maintain multiple versions of the same object key
  • Object metadata: custom key-value headers, content-type, content-encoding
  • Access control: per-bucket and per-object ACLs, IAM policies
  • Pre-signed URLs: generate time-limited URLs for upload/download without credentials
  • Lifecycle policies: auto-transition to cheaper storage tiers, auto-delete after N days

Non-Functional Requirements

Call out durability and commit ordering early. Strong read after write consistency means the current object version becomes visible only after the required durable data has been acknowledged and the metadata transaction commits. For architectural fundamentals of object stores compared to POSIX filesystems, see Storage Types: Block vs File vs Object.

  • Durability target: 99.999999999% (11 nines), achieved through replication, erasure coding, integrity checks, and continuous repair
  • Availability target: 99.99% for reads, 99.9% for writes in this design
  • Scalability: Exabytes of storage, millions of objects per bucket
  • Throughput: 100K+ requests/sec per bucket, multi-Gbps per object
  • Consistency: Strong read after write consistency
  • Low Latency: First-byte in < 100ms for most objects
  • Cost Efficiency: Tiered storage (hot, warm, cold, archive)

Capacity Estimations

Run this math before you size data nodes. Total exabytes and object count determine metadata store pressure. Replication factor then determines the raw storage multiplier.

MetricCalculationValue
Total stored objectsGiven (assumption documented in value)100+ trillion (S3-scale)
Total storageGiven (assumption documented in value)Exabytes
Requests / sec (global)From Requests / day ÷ 86400 (+ peak factor in value)100M+
Avg object sizeGiven (typical workload assumption)100 KB (highly variable)
Metadata per objectGiven~1 KB
Metadata storage100T x 1KB100 PB logical metadata before index/replication overhead
Logical payload at avg size100T × 100 KB~10 EB
3x raw capacity if all payload were hot10 EB × 3~30 EB before filesystem / parity overhead
Write throughputGiven (assumption documented in value)10M objects/sec
Full payload write bandwidth10M × 100 KB~1 TB/sec peak-equivalent
Replicated hot tier write bandwidth10M × 100 KB × 3~3 TB/sec including 3 copies
Full payload request bandwidth upper bound100M × 100 KB~10 TB/sec if every request transfers a full object
Replication factorGiven (assumption documented in value)3 (cross-AZ) + erasure coding for cold

Architecture Diagram

Walk your interviewer through the write path first. The system has an Object Service, Placement Service, Metadata Store, and replicated Data Nodes. The write path streams data to a primary node, replicates it across the selected failure domains, verifies integrity, and then commits the object version in the metadata store. A linearizable metadata commit makes the new version visible to subsequent reads.

Loading...

Component Deep Dives

In the room: write path order matters. Replicate data to three nodes and verify checksums BEFORE writing metadata for strong read after write.

Write Path (PUT Object)

Walk your interviewer through each stage on the diagram, starting with the PUT write path because the ordering between data durability and metadata visibility is the consistency guarantee your interviewer will probe. This follows all seven write steps: placement, replication, checksum verification, and metadata commit last.

  1. Client → API Gateway → Object Service (authenticates + authorizes)
  2. If object > 5GB → require multipart upload
  3. Object Service asks Placement Service: "Where to store 3 replicas?"
  4. Object Service streams data to primary data node
  5. Primary replicates to secondary nodes using the chosen replication protocol
  6. After all 3 replicas written + checksums verified: atomically commit the new object version, current-version pointer, and durable lifecycle event record
  7. Metadata commit is LAST step → the object becomes visible only after durable data exists

Read Path (GET Object)

The GET path is simpler but still metadata-first. Query the current version and data locations, then stream from the closest healthy replica.

  1. Client → API Gateway → Object Service
  2. Object Service queries Metadata Store with (bucket, key)
  3. Gets data locations: [{dn-17, vol-3, offset-4096}, ...]
  4. Selects closest/healthiest data node
  5. Streams data directly from data node to client
  6. Verifies checksum on the fly
  7. If checksum mismatch → try next replica + trigger repair

For small objects below 256 KB, cache metadata and payloads in Redis or memory when the hot object rate justifies it. Key cached reads by object version or invalidate them from the same metadata commit path so a successful overwrite or delete cannot leave stale data behind. For large objects, use range requests so clients can download different byte ranges in parallel.

Data Node: Block Format

Storage is organized as an append only volume file (log structured). The volume file contains fixed-size 64 MB blocks appended sequentially, similar to chunk abstractions in a Distributed File System (HDFS). Objects ≥ 64 MB span multiple blocks. Multiple small objects are packed into a single 64 MB block to avoid wasted space. An in-memory block index maps object_id → (volume_id, block_offset, length) for O(1) random access. Deletion marks the version or block references unreachable, and compaction later reclaims the physical space.

Event Bus Design (Kafka)

Topic: object-lifecycle-events
  Partitions: 128 (partition by hash(bucket_name, object_key))
  Events: ObjectCreated, ObjectRemoved, ObjectRestoreInitiated, StorageClassTransition
  Retention: 30 days (billing dispute + compliance replay)
  Producers: Object Service after metadata commit through a durable outbox/event log
  Consumers: Lifecycle Manager, Garbage Collector, usage metering, audit pipeline

Topic: replication-repair-events
  Partitions: 64 (partition by data_node_id)
  Events: ReplicaLost, ChecksumMismatch, RebuildStarted, RebuildCompleted
  Producers: Anti-entropy scrubber, Placement Service
  Consumers: Replication Repair workers (rebuild 3rd replica cross-AZ)

Topic: multipart-upload-events
  Events: UploadInitiated, PartUploaded, UploadCompleted, UploadAborted
  Partition key: upload_id
  Consumers: stale multipart abort job (Lifecycle Manager)

PUT path:
  validate IAM → write data nodes → commit metadata + durable event record → publish lifecycle event → 200
  Async: GC scans tombstones, lifecycle rules transition IA → Glacier, repair fixes bit-rot

Publication rule:
  If metadata commits but Kafka is temporarily unavailable, the durable event record is retried
  Object visibility does not depend on a one-shot Kafka publish succeeding

API Design

Object and Multipart APIs

Use conditional writes when concurrent clients may update the same key. A current version pointer or ETag lets the metadata store enforce optimistic concurrency.

HTTP
PUT    /{bucket}/{key}                    # Up to 5 GB in one request
GET    /{bucket}/{key}                    # Full or ranged download
HEAD   /{bucket}/{key}                    # Read current metadata
DELETE /{bucket}/{key}                    # Delete current version or add delete marker
GET    /{bucket}/{key}?versionId={v}      # Read a specific version
If-None-Match: *                           # Create only if current key is absent
If-Match: <etag>                           # Replace only if current version matches

POST   /{bucket}/{key}?uploads             # Initiate multipart upload
PUT    /{bucket}/{key}?partNumber={n}&uploadId={id}  # Upload part
GET    /{bucket}/{key}?uploadId={id}       # List uploaded parts
POST   /{bucket}/{key}?uploadId={id}       # Complete multipart upload
DELETE /{bucket}/{key}?uploadId={id}       # Abort multipart upload

PUT    /{bucket}                           # Create bucket
DELETE /{bucket}                           # Delete bucket when empty
GET    /{bucket}?list-type=2               # List objects with prefix and continuation token

POST   /api/presign
  { "method": "PUT", "bucket": "my-bucket", "key": "photo.jpg", "expires": 900 }
  Returns a short lived signed regional data endpoint

For retries, use an idempotency key for control plane operations where duplicate submission would be harmful. Multipart parts are naturally deduplicated by upload_id and part_number, while versioned PUTs may intentionally create a new version for each accepted write.

Pre-signed URL Contract

TYPESCRIPT
interface PresignRequest {
  method: "GET" | "PUT";
  bucket: string;
  key: string;
  expiresInSeconds: number;
  contentType?: string;
  checksumSha256?: string;
}

interface PresignResponse {
  url: string;
  expiresAt: string;
}

The signed request carries only the permissions and constraints needed for the transfer. The data plane validates the signature, expiry, bucket, key, and any declared content constraints before accepting the request.

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Metadata Store (FoundationDB / DynamoDB)

Primary key: (bucket_name, object_key, version_id).

JSON

{
  "bucket": "my-photos",
  "key": "2026/march/sunset.jpg",
  "version_id": "v_abc123",
  "size": 2457600,
  "etag": "etag-abc123",
  "content_type": "image/jpeg",
  "storage_class": "STANDARD",
  "checksum_sha256": "abc123...",
  "data_chunks": [
    {"object_start": 0, "length": 2457600, "node_id": "dn-17", "volume_id": "vol-3", "offset": 4096},
    {"object_start": 0, "length": 2457600, "node_id": "dn-42", "volume_id": "vol-1", "offset": 8192},
    {"object_start": 0, "length": 2457600, "node_id": "dn-63", "volume_id": "vol-7", "offset": 2048}
  ],
  "custom_metadata": { "x-amz-meta-photographer": "Alice" },
  "delete_marker": false,
  "created_at": "2026-03-14T10:00:00Z",
  "version_generation": 17
}

Object and Listing Indexes

Point lookup index:
  hash(bucket_name, object_key) -> current_version_id
  used for GET/HEAD/DELETE and updated atomically with the metadata transaction

Version index:
  (bucket_name, object_key, version_id) -> object metadata + data locations

Current namespace index:
  (bucket_name, object_key) -> current_version_id
  ordered by object_key so LIST can range scan by prefix with a continuation cursor
  delete markers are represented in the current namespace state
  range partition this index by key ranges rather than hashing so prefix scans stay efficient

Multipart state:
  (bucket_name, object_key, upload_id) -> uploaded part numbers, ETags, checksums, placement

Bucket Metadata

SQL
CREATE TABLE buckets (
    bucket_name     TEXT PRIMARY KEY,
    owner_id        UUID NOT NULL,
    region          TEXT NOT NULL,
    versioning      TEXT DEFAULT 'disabled',
    storage_class   TEXT DEFAULT 'STANDARD',
    lifecycle_rules JSONB,
    cors_config     JSONB,
    policy          JSONB,
    created_at      TIMESTAMPTZ DEFAULT NOW()
);

Fault Tolerance

How 11 Nines Durability Is Approached

The 11 nines figure is a design target. Replication, erasure coding, integrity verification, repair speed, and failure domain separation all contribute to the probability of losing an object.

  1. 3x replication across AZs: For illustration only, if one AZ has a 10-3 probability of a relevant failure over a stated period and the failures are independent, simultaneous loss of all 3 AZs is 10-9. This redundancy math is not a proof of 11 nines object durability.
  2. Checksums at every storage layer: Object-level SHA-256 and block-level checksums detect corruption. TCP and TLS protect transport integrity but do not substitute for persistent object checksums.
  3. Background scrubbing: Continuously scrub stored blocks so every object is checked at least weekly. Any checksum mismatch triggers repair from a healthy replica.
  4. Rebuild on failure: Placement Service detects failure through heartbeat, triggers re-replication, and throttles rebuild traffic at 100 MB/s per node.
  5. Erasure coding for cold data: Example design RS(10,4) can lose any 4 fragments and still recover. It has 1.4x storage overhead versus 3x full replication. This is an illustrative coding profile, not a claim about S3 internals.
  6. Cross-region replication (optional): Asynchronous replication to another region provides disaster recovery rather than immediate cross region consistency.

Strong Read-After-Write Consistency

Amazon S3 provides strong read after write consistency for object operations and LIST. The internal implementation details are not fully public. In this S3 style design, strong consistency comes from writing durable data first and then atomically publishing the new object version through a linearizable metadata store. GET, HEAD, DELETE, and LIST read the committed metadata state, so a successful write is immediately visible to subsequent reads. For concurrent writers, use an atomic current-version update or conditional ETag based write to prevent lost updates.

Multipart Upload Recovery

A 5 TB file upload fails at 90%? Multipart upload keeps the completed parts. 1) Initiate and get upload_id. 2) Upload parts independently and in parallel. 3) Retry only failed parts. 4) Complete with the ordered part ETag list. The service publishes the completed object by committing metadata and part references rather than copying the full payload again. With at most 10,000 parts, a 5 TB design requires an average part size of about 500 MB or larger. Incomplete uploads are cleaned up by a lifecycle rule after 7 days.

Handling Hot Objects

Use CDN caching with CloudFront, request coalescing on cache misses so 1000 concurrent GETs can share one backend fetch, dynamic read replica expansion for hot objects, and signed URLs that target the CDN or an authorized regional data endpoint.

Garbage Collection

Lazy GC: The scanner reads the metadata store for live objects, and reads data nodes for stored blocks. Blocks not referenced are marked as garbage. After a 48-hour grace period, they are reclaimed, similar to JVM garbage collection and LSM tree compaction.

Additional Considerations

Storage Tiers

TierUse CaseDurabilityAvailabilityAccess / Retrieval
StandardFrequently accessed11 nines99.99%< 100ms
Infrequent AccessMonthly access11 nines99.9%< 100ms
Glacier Flexible RetrievalRare archive access11 nines99.99% after restoreMinutes to 12 hours after restore
Deep ArchiveLong term archive11 nines99.99% after restore9 to 48 hours after restore

For this interview design, the < 100ms values for the online tiers are latency targets, not universal AWS guarantees. Glacier retrieval depends on the selected restore tier, and Deep Archive requires a restore before normal object access.

Lifecycle automation transitions data predictably (such as after 30 days to IA, after 90 days to Glacier, and after 365 days to Deep Archive), applied daily by the Lifecycle Manager.

Security, Encryption, and Tenant Isolation

Object storage is multi tenant, so authorization and data isolation must apply to both metadata and the data plane. Authenticate through IAM or an equivalent identity layer, authorize bucket and object operations before placement, and carry the authorization context into pre signed URLs.

  • Encryption in transit: Require TLS for client, metadata, and data node traffic.
  • Encryption at rest: Encrypt object data and metadata with managed keys or customer managed KMS keys, with key rotation and audit logging.
  • Tenant isolation: Scope metadata partitions, quotas, lifecycle policies, and signed URLs by tenant and bucket.
  • Data exposure: Do not expose internal data node addresses directly when a gateway or CDN can issue a signed regional endpoint. The endpoint should resolve to current healthy placement rather than permanently binding the client to one storage node.
  • Audit: Record object writes, deletes, policy changes, permission failures, and administrative actions in an append only audit stream.

Event Notifications

In this design, object events are published to Kafka after the metadata commit. Consumers must be idempotent because event delivery is asynchronous and may contain duplicates. For Amazon S3, native event notification destinations include SQS, SNS, Lambda, and EventBridge, and S3 event notifications are designed for at least once delivery. Example use cases include thumbnail generation when an image is uploaded, ETL pipeline triggers when logs are uploaded, and search index updates when an object is deleted.

Interview Walkthrough

  • 25-minute cut

    Skip the deeper storage internals unless the interview targets staff level.

    • Split metadata store from blob data nodes (5 min)
    • Write path: replicate data before metadata commit (6 min)
    • Multipart upload for objects > 5 GB (5 min)
    • Erasure coding for cold and archive tiers (5 min)
    • Background scrubbing compares replica checksums (4 min)
  • Split metadata store (bucket, key, version, replica locations) from blob data (append only volumes on data nodes), establishing the defining architectural boundary.
  • Write path order matters: stream data to 3 replicas and verify checksums first, then atomically publish the current metadata state. The linearizable metadata commit provides the namespace visibility guarantee in this design.
  • Multipart upload for objects > 5 GB: upload parts independently, then complete assembly by updating metadata only with no data copy.
  • An illustrative RS(10,4) erasure coding profile yields 1.4x storage overhead versus 3x for hot replication and survives loss of any 4 fragments.
  • Background scrubbing covers the stored corpus continuously with each region of data targeted at least weekly, while any bit-rot detected on read triggers re-replication from a healthy copy.
  • Lifecycle policies automate tier transitions (Standard to IA to Glacier) to optimize cost without manual intervention.
  • Strong consistency applies to object reads, writes, deletes, and LIST. A newly committed object is immediately visible to subsequent reads and listings in this design.
  • Common pitfall: publishing the metadata pointer before durable data exists. The design prevents this by committing metadata only after the required replicas and checksums are verified.

Engineering Trade-offs

Erasure Coding vs Replication for Durability

Blob storage trades durability cost against read latency. Replication is simpler for hot data, while erasure coding reduces storage overhead for colder data.

Pure replication (3 copies): 3x storage overhead. Simple, fast recovery. Used for hot tier.

Erasure Coding ⭐ (Reed-Solomon): An illustrative 6+3 scheme stores 6 data fragments plus 3 parity fragments, so it has 1.5x storage overhead and survives loss of any 3 fragments. An illustrative RS(10,4) scheme has 1.4x overhead and survives loss of any 4 fragments. These are design choices rather than claims about undisclosed S3 internals. The trade-off is lower storage cost with higher reconstruction I/O because recovery must read enough surviving fragments. Hot tier: 3 copy replication for simpler fast recovery. Cold tier: erasure coding for cost efficiency.

Consistency: Eventual vs Strong Read After Write

Amazon S3 introduced strong read after write consistency for object operations in December 2020, and LIST is also strongly consistent. The internal metadata implementation is not fully public. In this S3 style design, the metadata service atomically publishes the current object version only after the required data replicas are durable. GET, HEAD, DELETE, and LIST then read the committed metadata state. Concurrent writers use conditional ETag or version checks so one write cannot silently overwrite another.

Multipart Upload: How Large File Uploads Work

A single PUT is limited to 5 GB in Amazon S3, while current S3 objects can be up to 50 TB, so a 5 TB interview design must use multipart upload. A network interruption at 99% should not force a full restart. Multipart upload sends parts in parallel, for example 100 parts at once for much higher throughput, and failed parts are retried independently. Completion publishes the assembled object through metadata and part references rather than copying the full payload again. For a 5 TB object, at most 10,000 parts means the average part size must be about 500 MB or larger. A lifecycle rule aborts incomplete uploads after 7 days to prevent orphaned storage.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...