Interview Setup
Interview Prompt
Design an S3-style blob storage system storing 100+ trillion objects at exabyte scale, serving 100M+ requests/sec with 11 nines durability and strong read after write consistency.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Consistency model: strong read after write or eventual? | The design chooses strong read after write so a successful object write is immediately visible to GET, HEAD, and LIST in the authoritative namespace. Eventual consistency would require clients or services to tolerate stale reads and possibly use retries or read repair. |
| Replication (3x full copies) or erasure coding for durability? | Replication = 3x storage cost and simpler fast recovery. A Reed-Solomon design such as 6+3 uses 1.5x storage overhead and slower reconstruction. The exact coding profile is a system design choice, not an assumption about S3 internals. |
| Max object size and multipart threshold? | This design uses a 5 TB maximum object size and a 100 MB threshold as interview assumptions. A single S3 PUT supports up to 5 GB, while multipart upload handles larger objects. |
| Pre-signed URLs for direct client upload: in scope? | Signed data-plane URLs offload large payload transfer from the API control plane. The storage nodes still validate the signature and authorization context. |
Scope
In scope
- Object metadata store
- Erasure coding vs replication
- Multipart upload
- Garbage collection
- Pre-signed URLs
- Consistency model (strong read after write)
Out of scope (state explicitly)
- Client desktop/mobile app implementation
- End-user file preview rendering for every format
- Building raw block storage hardware
Functional Requirements
Start by confirming object size range and consistency model with your interviewer. Ask about multipart uploads, versioning, lifecycle tiers, and pre signed URL requirements.
- Upload objects (blobs) from 1 byte to a 5 TB design limit. Single PUT operations handle up to 5 GB, whereas larger objects use multipart uploads.
- Download objects via GET with support for range requests (partial downloads)
- Delete objects with immediate namespace visibility change and lazy physical storage reclamation
- List objects in a bucket with prefix filtering and pagination
- Multipart upload: upload large objects in parts, assemble on completion
- Versioning: maintain multiple versions of the same object key
- Object metadata: custom key-value headers, content-type, content-encoding
- Access control: per-bucket and per-object ACLs, IAM policies
- Pre-signed URLs: generate time-limited URLs for upload/download without credentials
- Lifecycle policies: auto-transition to cheaper storage tiers, auto-delete after N days
Non-Functional Requirements
Call out durability and commit ordering early. Strong read after write consistency means the current object version becomes visible only after the required durable data has been acknowledged and the metadata transaction commits. For architectural fundamentals of object stores compared to POSIX filesystems, see Storage Types: Block vs File vs Object.
- Durability target: 99.999999999% (11 nines), achieved through replication, erasure coding, integrity checks, and continuous repair
- Availability target: 99.99% for reads, 99.9% for writes in this design
- Scalability: Exabytes of storage, millions of objects per bucket
- Throughput: 100K+ requests/sec per bucket, multi-Gbps per object
- Consistency: Strong read after write consistency
- Low Latency: First-byte in < 100ms for most objects
- Cost Efficiency: Tiered storage (hot, warm, cold, archive)
Capacity Estimations
Run this math before you size data nodes. Total exabytes and object count determine metadata store pressure. Replication factor then determines the raw storage multiplier.
| Metric | Calculation | Value |
|---|---|---|
| Total stored objects | Given (assumption documented in value) | 100+ trillion (S3-scale) |
| Total storage | Given (assumption documented in value) | Exabytes |
| Requests / sec (global) | From Requests / day ÷ 86400 (+ peak factor in value) | 100M+ |
| Avg object size | Given (typical workload assumption) | 100 KB (highly variable) |
| Metadata per object | Given | ~1 KB |
| Metadata storage | 100T x 1KB | 100 PB logical metadata before index/replication overhead |
| Logical payload at avg size | 100T × 100 KB | ~10 EB |
| 3x raw capacity if all payload were hot | 10 EB × 3 | ~30 EB before filesystem / parity overhead |
| Write throughput | Given (assumption documented in value) | 10M objects/sec |
| Full payload write bandwidth | 10M × 100 KB | ~1 TB/sec peak-equivalent |
| Replicated hot tier write bandwidth | 10M × 100 KB × 3 | ~3 TB/sec including 3 copies |
| Full payload request bandwidth upper bound | 100M × 100 KB | ~10 TB/sec if every request transfers a full object |
| Replication factor | Given (assumption documented in value) | 3 (cross-AZ) + erasure coding for cold |
Architecture Diagram
Walk your interviewer through the write path first. The system has an Object Service, Placement Service, Metadata Store, and replicated Data Nodes. The write path streams data to a primary node, replicates it across the selected failure domains, verifies integrity, and then commits the object version in the metadata store. A linearizable metadata commit makes the new version visible to subsequent reads.
Component Deep Dives
In the room: write path order matters. Replicate data to three nodes and verify checksums BEFORE writing metadata for strong read after write.
Write Path (PUT Object)
Walk your interviewer through each stage on the diagram, starting with the PUT write path because the ordering between data durability and metadata visibility is the consistency guarantee your interviewer will probe. This follows all seven write steps: placement, replication, checksum verification, and metadata commit last.
- Client → API Gateway → Object Service (authenticates + authorizes)
- If object > 5GB → require multipart upload
- Object Service asks Placement Service: "Where to store 3 replicas?"
- Object Service streams data to primary data node
- Primary replicates to secondary nodes using the chosen replication protocol
- After all 3 replicas written + checksums verified: atomically commit the new object version, current-version pointer, and durable lifecycle event record
- Metadata commit is LAST step → the object becomes visible only after durable data exists
Read Path (GET Object)
The GET path is simpler but still metadata-first. Query the current version and data locations, then stream from the closest healthy replica.
- Client → API Gateway → Object Service
- Object Service queries Metadata Store with (bucket, key)
- Gets data locations:
[{dn-17, vol-3, offset-4096}, ...] - Selects closest/healthiest data node
- Streams data directly from data node to client
- Verifies checksum on the fly
- If checksum mismatch → try next replica + trigger repair
For small objects below 256 KB, cache metadata and payloads in Redis or memory when the hot object rate justifies it. Key cached reads by object version or invalidate them from the same metadata commit path so a successful overwrite or delete cannot leave stale data behind. For large objects, use range requests so clients can download different byte ranges in parallel.
Data Node: Block Format
Storage is organized as an append only volume file (log structured). The volume file contains fixed-size 64 MB blocks appended sequentially, similar to chunk abstractions in a Distributed File System (HDFS). Objects ≥ 64 MB span multiple blocks. Multiple small objects are packed into a single 64 MB block to avoid wasted space. An in-memory block index maps object_id → (volume_id, block_offset, length) for O(1) random access. Deletion marks the version or block references unreachable, and compaction later reclaims the physical space.
Event Bus Design (Kafka)
Topic: object-lifecycle-events Partitions: 128 (partition by hash(bucket_name, object_key)) Events: ObjectCreated, ObjectRemoved, ObjectRestoreInitiated, StorageClassTransition Retention: 30 days (billing dispute + compliance replay) Producers: Object Service after metadata commit through a durable outbox/event log Consumers: Lifecycle Manager, Garbage Collector, usage metering, audit pipeline Topic: replication-repair-events Partitions: 64 (partition by data_node_id) Events: ReplicaLost, ChecksumMismatch, RebuildStarted, RebuildCompleted Producers: Anti-entropy scrubber, Placement Service Consumers: Replication Repair workers (rebuild 3rd replica cross-AZ) Topic: multipart-upload-events Events: UploadInitiated, PartUploaded, UploadCompleted, UploadAborted Partition key: upload_id Consumers: stale multipart abort job (Lifecycle Manager) PUT path: validate IAM → write data nodes → commit metadata + durable event record → publish lifecycle event → 200 Async: GC scans tombstones, lifecycle rules transition IA → Glacier, repair fixes bit-rot Publication rule: If metadata commits but Kafka is temporarily unavailable, the durable event record is retried Object visibility does not depend on a one-shot Kafka publish succeeding
API Design
Object and Multipart APIs
Use conditional writes when concurrent clients may update the same key. A current version pointer or ETag lets the metadata store enforce optimistic concurrency.
PUT /{bucket}/{key} # Up to 5 GB in one request
GET /{bucket}/{key} # Full or ranged download
HEAD /{bucket}/{key} # Read current metadata
DELETE /{bucket}/{key} # Delete current version or add delete marker
GET /{bucket}/{key}?versionId={v} # Read a specific version
If-None-Match: * # Create only if current key is absent
If-Match: <etag> # Replace only if current version matches
POST /{bucket}/{key}?uploads # Initiate multipart upload
PUT /{bucket}/{key}?partNumber={n}&uploadId={id} # Upload part
GET /{bucket}/{key}?uploadId={id} # List uploaded parts
POST /{bucket}/{key}?uploadId={id} # Complete multipart upload
DELETE /{bucket}/{key}?uploadId={id} # Abort multipart upload
PUT /{bucket} # Create bucket
DELETE /{bucket} # Delete bucket when empty
GET /{bucket}?list-type=2 # List objects with prefix and continuation token
POST /api/presign
{ "method": "PUT", "bucket": "my-bucket", "key": "photo.jpg", "expires": 900 }
Returns a short lived signed regional data endpointFor retries, use an idempotency key for control plane operations where duplicate submission would be harmful. Multipart parts are naturally deduplicated by upload_id and part_number, while versioned PUTs may intentionally create a new version for each accepted write.
Pre-signed URL Contract
interface PresignRequest {
method: "GET" | "PUT";
bucket: string;
key: string;
expiresInSeconds: number;
contentType?: string;
checksumSha256?: string;
}
interface PresignResponse {
url: string;
expiresAt: string;
}The signed request carries only the permissions and constraints needed for the transfer. The data plane validates the signature, expiry, bucket, key, and any declared content constraints before accepting the request.
Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Metadata Store (FoundationDB / DynamoDB)
Primary key: (bucket_name, object_key, version_id).
{
"bucket": "my-photos",
"key": "2026/march/sunset.jpg",
"version_id": "v_abc123",
"size": 2457600,
"etag": "etag-abc123",
"content_type": "image/jpeg",
"storage_class": "STANDARD",
"checksum_sha256": "abc123...",
"data_chunks": [
{"object_start": 0, "length": 2457600, "node_id": "dn-17", "volume_id": "vol-3", "offset": 4096},
{"object_start": 0, "length": 2457600, "node_id": "dn-42", "volume_id": "vol-1", "offset": 8192},
{"object_start": 0, "length": 2457600, "node_id": "dn-63", "volume_id": "vol-7", "offset": 2048}
],
"custom_metadata": { "x-amz-meta-photographer": "Alice" },
"delete_marker": false,
"created_at": "2026-03-14T10:00:00Z",
"version_generation": 17
}Object and Listing Indexes
Point lookup index: hash(bucket_name, object_key) -> current_version_id used for GET/HEAD/DELETE and updated atomically with the metadata transaction Version index: (bucket_name, object_key, version_id) -> object metadata + data locations Current namespace index: (bucket_name, object_key) -> current_version_id ordered by object_key so LIST can range scan by prefix with a continuation cursor delete markers are represented in the current namespace state range partition this index by key ranges rather than hashing so prefix scans stay efficient Multipart state: (bucket_name, object_key, upload_id) -> uploaded part numbers, ETags, checksums, placement
Bucket Metadata
CREATE TABLE buckets (
bucket_name TEXT PRIMARY KEY,
owner_id UUID NOT NULL,
region TEXT NOT NULL,
versioning TEXT DEFAULT 'disabled',
storage_class TEXT DEFAULT 'STANDARD',
lifecycle_rules JSONB,
cors_config JSONB,
policy JSONB,
created_at TIMESTAMPTZ DEFAULT NOW()
);Fault Tolerance
How 11 Nines Durability Is Approached
The 11 nines figure is a design target. Replication, erasure coding, integrity verification, repair speed, and failure domain separation all contribute to the probability of losing an object.
- 3x replication across AZs: For illustration only, if one AZ has a 10-3 probability of a relevant failure over a stated period and the failures are independent, simultaneous loss of all 3 AZs is 10-9. This redundancy math is not a proof of 11 nines object durability.
- Checksums at every storage layer: Object-level SHA-256 and block-level checksums detect corruption. TCP and TLS protect transport integrity but do not substitute for persistent object checksums.
- Background scrubbing: Continuously scrub stored blocks so every object is checked at least weekly. Any checksum mismatch triggers repair from a healthy replica.
- Rebuild on failure: Placement Service detects failure through heartbeat, triggers re-replication, and throttles rebuild traffic at 100 MB/s per node.
- Erasure coding for cold data: Example design RS(10,4) can lose any 4 fragments and still recover. It has 1.4x storage overhead versus 3x full replication. This is an illustrative coding profile, not a claim about S3 internals.
- Cross-region replication (optional): Asynchronous replication to another region provides disaster recovery rather than immediate cross region consistency.
Strong Read-After-Write Consistency
Amazon S3 provides strong read after write consistency for object operations and LIST. The internal implementation details are not fully public. In this S3 style design, strong consistency comes from writing durable data first and then atomically publishing the new object version through a linearizable metadata store. GET, HEAD, DELETE, and LIST read the committed metadata state, so a successful write is immediately visible to subsequent reads. For concurrent writers, use an atomic current-version update or conditional ETag based write to prevent lost updates.
Multipart Upload Recovery
A 5 TB file upload fails at 90%? Multipart upload keeps the completed parts. 1) Initiate and get upload_id. 2) Upload parts independently and in parallel. 3) Retry only failed parts. 4) Complete with the ordered part ETag list. The service publishes the completed object by committing metadata and part references rather than copying the full payload again. With at most 10,000 parts, a 5 TB design requires an average part size of about 500 MB or larger. Incomplete uploads are cleaned up by a lifecycle rule after 7 days.
Handling Hot Objects
Use CDN caching with CloudFront, request coalescing on cache misses so 1000 concurrent GETs can share one backend fetch, dynamic read replica expansion for hot objects, and signed URLs that target the CDN or an authorized regional data endpoint.
Garbage Collection
Lazy GC: The scanner reads the metadata store for live objects, and reads data nodes for stored blocks. Blocks not referenced are marked as garbage. After a 48-hour grace period, they are reclaimed, similar to JVM garbage collection and LSM tree compaction.
Additional Considerations
Storage Tiers
| Tier | Use Case | Durability | Availability | Access / Retrieval |
|---|---|---|---|---|
| Standard | Frequently accessed | 11 nines | 99.99% | < 100ms |
| Infrequent Access | Monthly access | 11 nines | 99.9% | < 100ms |
| Glacier Flexible Retrieval | Rare archive access | 11 nines | 99.99% after restore | Minutes to 12 hours after restore |
| Deep Archive | Long term archive | 11 nines | 99.99% after restore | 9 to 48 hours after restore |
For this interview design, the < 100ms values for the online tiers are latency targets, not universal AWS guarantees. Glacier retrieval depends on the selected restore tier, and Deep Archive requires a restore before normal object access.
Lifecycle automation transitions data predictably (such as after 30 days to IA, after 90 days to Glacier, and after 365 days to Deep Archive), applied daily by the Lifecycle Manager.
Security, Encryption, and Tenant Isolation
Object storage is multi tenant, so authorization and data isolation must apply to both metadata and the data plane. Authenticate through IAM or an equivalent identity layer, authorize bucket and object operations before placement, and carry the authorization context into pre signed URLs.
- Encryption in transit: Require TLS for client, metadata, and data node traffic.
- Encryption at rest: Encrypt object data and metadata with managed keys or customer managed KMS keys, with key rotation and audit logging.
- Tenant isolation: Scope metadata partitions, quotas, lifecycle policies, and signed URLs by tenant and bucket.
- Data exposure: Do not expose internal data node addresses directly when a gateway or CDN can issue a signed regional endpoint. The endpoint should resolve to current healthy placement rather than permanently binding the client to one storage node.
- Audit: Record object writes, deletes, policy changes, permission failures, and administrative actions in an append only audit stream.
Event Notifications
In this design, object events are published to Kafka after the metadata commit. Consumers must be idempotent because event delivery is asynchronous and may contain duplicates. For Amazon S3, native event notification destinations include SQS, SNS, Lambda, and EventBridge, and S3 event notifications are designed for at least once delivery. Example use cases include thumbnail generation when an image is uploaded, ETL pipeline triggers when logs are uploaded, and search index updates when an object is deleted.
Interview Walkthrough
- 25-minute cut
Skip the deeper storage internals unless the interview targets staff level.
- Split metadata store from blob data nodes (5 min)
- Write path: replicate data before metadata commit (6 min)
- Multipart upload for objects > 5 GB (5 min)
- Erasure coding for cold and archive tiers (5 min)
- Background scrubbing compares replica checksums (4 min)
- Split metadata store (bucket, key, version, replica locations) from blob data (append only volumes on data nodes), establishing the defining architectural boundary.
- Write path order matters: stream data to 3 replicas and verify checksums first, then atomically publish the current metadata state. The linearizable metadata commit provides the namespace visibility guarantee in this design.
- Multipart upload for objects > 5 GB: upload parts independently, then complete assembly by updating metadata only with no data copy.
- An illustrative RS(10,4) erasure coding profile yields 1.4x storage overhead versus 3x for hot replication and survives loss of any 4 fragments.
- Background scrubbing covers the stored corpus continuously with each region of data targeted at least weekly, while any bit-rot detected on read triggers re-replication from a healthy copy.
- Lifecycle policies automate tier transitions (Standard to IA to Glacier) to optimize cost without manual intervention.
- Strong consistency applies to object reads, writes, deletes, and LIST. A newly committed object is immediately visible to subsequent reads and listings in this design.
- Common pitfall: publishing the metadata pointer before durable data exists. The design prevents this by committing metadata only after the required replicas and checksums are verified.
Engineering Trade-offs
Erasure Coding vs Replication for Durability
Blob storage trades durability cost against read latency. Replication is simpler for hot data, while erasure coding reduces storage overhead for colder data.
Pure replication (3 copies): 3x storage overhead. Simple, fast recovery. Used for hot tier.
Erasure Coding ⭐ (Reed-Solomon): An illustrative 6+3 scheme stores 6 data fragments plus 3 parity fragments, so it has 1.5x storage overhead and survives loss of any 3 fragments. An illustrative RS(10,4) scheme has 1.4x overhead and survives loss of any 4 fragments. These are design choices rather than claims about undisclosed S3 internals. The trade-off is lower storage cost with higher reconstruction I/O because recovery must read enough surviving fragments. Hot tier: 3 copy replication for simpler fast recovery. Cold tier: erasure coding for cost efficiency.
Consistency: Eventual vs Strong Read After Write
Amazon S3 introduced strong read after write consistency for object operations in December 2020, and LIST is also strongly consistent. The internal metadata implementation is not fully public. In this S3 style design, the metadata service atomically publishes the current object version only after the required data replicas are durable. GET, HEAD, DELETE, and LIST then read the committed metadata state. Concurrent writers use conditional ETag or version checks so one write cannot silently overwrite another.
Multipart Upload: How Large File Uploads Work
A single PUT is limited to 5 GB in Amazon S3, while current S3 objects can be up to 50 TB, so a 5 TB interview design must use multipart upload. A network interruption at 99% should not force a full restart. Multipart upload sends parts in parallel, for example 100 parts at once for much higher throughput, and failed parts are retried independently. Completion publishes the assembled object through metadata and part references rather than copying the full payload again. For a 5 TB object, at most 10,000 parts means the average part size must be about 500 MB or larger. A lifecycle rule aborts incomplete uploads after 7 days to prevent orphaned storage.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.