Interview Setup
Interview Prompt
Design Dropbox or Google Drive cloud file storage and multi device synchronization for 500M users. Users upload files up to 50 GB, sync across desktop, mobile, and web devices, share folders, and resolve conflicts when multiple devices edit while offline.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Is cross user deduplication acceptable, or must each user's bytes be isolated? | Hash based deduplication saves 30 to 50 percent of storage capacity but introduces the risk that an attacker can infer whether another user stores an identical file. Dropbox historically accepted this trade off for global efficiency. |
| How fast must changes propagate to other devices? | Synchronization under one second requires persistent WebSocket connections paired with Kafka event streaming, whereas multi minute propagation allows simpler client polling. That choice directly shapes the notification service architecture. |
| File sync or character level collaboration such as Google Docs? | Dropbox preserves both files by creating a conflicted copy upon concurrent edits, whereas Google Docs requires Operational Transformation or Conflict Free Replicated Data Types for character level collaboration, fundamentally altering the concurrency model. |
| What is the typical edit pattern, such as full file replacement or small in place modification? | Content defined chunking saves up to 25 times more bandwidth on a 100 MB file with a 1 KB edit, whereas fixed size chunking shifts boundaries and forces reuploading the entire file. |
Scope
In scope
- File chunking and deduplication
- Delta sync (rsync)
- Conflict resolution
- Content addressed chunk storage
- Metadata DB design
- Notification of changes across devices
Out of scope (state explicitly)
- Client desktop/mobile app implementation
- End user file preview rendering for every format
- Building raw block storage hardware
Functional Requirements
Start by asking your interviewer about file upload and download sizes, multi device synchronization expectations, and collaboration scope. Conflict resolution and content defined chunking are the main design pivots for scalable cloud storage.
- Upload Files: Support uploading files of any type and size up to 50 GB.
- Download Files: Allow authenticated users to download stored files quickly and reliably.
- Cross Device Synchronization: Automatically synchronize file modifications across desktop, mobile, and web clients.
- File Versioning: Maintain complete revision history and support point in time rollbacks to previous versions.
- Sharing and Permissions: Enable sharing files and folders with specific users or via public links with view or edit permissions.
- Offline Access: Permit users to read and edit files offline, queuing updates to synchronize automatically upon reconnection.
- Conflict Resolution: Detect and resolve simultaneous concurrent edits on the same file without silent data loss.
- Real Time Notifications: Push instant notifications to connected devices whenever shared folders or files are updated.
- Hierarchical Organization: Support nested directory structures with folder creation, renaming, and moving operations.
- Block Level Deduplication: Prevent duplicate chunks from being stored multiple times across files and user accounts.
Non-Functional Requirements
Interview discussions focus heavily on synchronization latency and metadata consistency across devices. Separating relational metadata from raw object storage early is critical because they operate under fundamentally different latency and durability requirements.
- Extreme Durability: Target 99.999999999% (11 nines) durability for all committed file chunks, keeping the probability of permanent data loss extremely low.
- High Availability: Maintain 99.99% availability for file downloads and directory metadata queries.
- Consistency Separation: Enforce strong consistency for directory metadata and access control lists, while allowing eventual consistency for multi device sync notification propagation.
- Low Latency Sync: Propagate file changes to all online connected devices within 5 seconds of upload completion.
- Bandwidth Efficiency: Minimize network data transfer by transmitting only modified byte chunks rather than full files.
- Global Scalability: Support 500 million registered users, 100 billion files, and 150 petabytes of replicated storage.
- End to End Security: Protect assets using server side AES-256 encryption at rest, TLS 1.3 in transit, and granular access control lists. Optional client side zero knowledge encryption can be enabled for workloads that do not require server side content deduplication.
Capacity Estimations
Work through storage and bandwidth calculations before choosing a chunking strategy. Quantifying average file sizes, storage allocations per user, and synchronization frequencies dictates your metadata throughput and reveals the magnitude of deduplication savings.
| Metric | Calculation | Value |
|---|---|---|
| Users | Given platform baseline assumption | 500M |
| Average files per user | Typical active workload assumption | 200 |
| Total files | 500M users x 200 files/user | 100B |
| Average file size | Typical cloud storage workload distribution | 500 KB |
| Total logical storage | 100B files x 500 KB average file size | 50 PB |
| With replication (3x) | 50 PB x 3 replicas across availability zones | 150 PB |
| DAU | Product usage assumption | 100M |
| File sync events / day | Workload assumption (10 sync events per active user per day) | 1B |
| Sync events / sec | 1B file sync events ÷ 86,400 seconds | ~12K |
| Metadata reads / sec | Assume ~4 metadata reads per sync event: 12K x 4 | ~50K |
| Average changed data per sync | Typical document or code delta modification (10% of file) | 50 KB |
Architecture Diagram
In an interview setting, separate directory metadata (file tree, versions, permissions) from blob storage (content addressed chunks) before detailing the synchronization protocol.
Walk your interviewer through the architecture by distinguishing between the file write path and the read path. Files are partitioned into content defined chunks, hashed for deduplication, and stored directly in immutable object storage. Directory metadata, including folder hierarchies, version histories, and sharing permissions, resides in PostgreSQL. The client sync agent compares local chunk hashes with server metadata, uploads only newly detected delta chunks through pre signed S3 URLs, and triggers asynchronous notifications across registered user devices.
Component Deep Dives
Chunking: The Key Innovation
The platform architecture divides responsibilities across three core layers: client side synchronization agents, content addressed block storage services, and relational metadata services.
Partitioning files into discrete chunks provides four essential architectural advantages. Delta synchronization allows the client to upload only altered chunks when a large file is modified. Deduplication ensures that identical chunks across different files or users are stored only once. Parallel transfers allow multiple HTTP connections to use available bandwidth efficiently. Resumability lets an interrupted upload restart from the last successfully written chunk rather than retransmitting the entire file.
Fixed size chunking has a boundary shift problem. If a byte is inserted at the beginning of a file, the byte offset of every later block changes and all later hashes may change. Content defined chunking instead uses a rolling hash such as a Rabin fingerprint across a sliding byte window, typically 48 bytes, to identify content dependent boundaries. Because the cut points depend on local byte sequences, an insertion or deletion usually changes only a small local region before the boundary pattern realigns. Each chunk is hashed with SHA-256 to create a content addressed identifier. During upload initialization, the client sends the ordered hash list to the server, which returns only the chunks that are missing from storage.
Simplified pseudocode: content defined chunking with a Rabin fingerprint
chunks = []
window_size = 48 # bytes
min_chunk = 256 * 1024 # 256 KB minimum
max_chunk = 8 * 1024 * 1024 # 8 MB maximum
target_chunk = 4 * 1024 * 1024 # 4 MB target
position = 0
chunk_start = 0
while position < len(data):
hash = rabin_hash(data[position:position+window_size])
chunk_size = position - chunk_start
if (hash % target_chunk == 0 and chunk_size >= min_chunk) or chunk_size >= max_chunk:
chunks.append(data[chunk_start:position])
chunk_start = position
position += 1
chunks.append(data[chunk_start:])
return chunksDesktop Sync Agent
Client sync agents continuously reconcile local file system state with remote cloud metadata while minimizing CPU and battery consumption.
- File System Watcher: Monitors local folders using native OS kernel events such as inotify on Linux, FSEvents on macOS, and ReadDirectoryChangesW on Windows to detect file additions, modifications, and deletions without continuous disk polling.
- Local Chunking and Diffing: When a local file is modified, the agent executes content defined chunking, computes SHA-256 hashes for all chunks, compares the result against its local SQLite cache of known hashes, and uploads only newly generated chunks through pre signed S3 URLs.
- Remote Change Reconciliation: Upon receiving a WebSocket notification from the server, the agent requests metadata updates, compares the remote chunk manifest against local blocks, downloads only missing chunks, and reconstructs the updated file atomically.
Chunk Storage Gateway
The Chunk Storage Gateway is a stateless control and verification tier. The normal data path sends chunk payloads directly between the client and S3 through pre signed URLs. The gateway authorizes uploads, coordinates storage references, and validates object metadata without proxying large payloads through application servers.
- Upload Pipeline: The Metadata Service checks the ordered canonical SHA-256 manifest and returns pre signed S3 URLs only for missing chunks. The client uploads each chunk directly to S3 with an object-level checksum for the uploaded payload and records the canonical SHA-256 of the uncompressed content. S3 validates the uploaded bytes, while a background verifier can decompress and recheck the canonical content hash when compression is enabled.
- Download Pipeline: The Metadata Service returns short lived signed URLs for the required chunks. The client fetches chunks directly from S3, decompresses them when required, verifies the canonical SHA-256 digest of the original chunk content, and reconstructs the file.
Metadata Service
The Metadata Service maintains the authoritative state of directory hierarchies, version manifests, and user access permissions with strict transactional consistency. It also maintains the durable change journal used for offline catch up after notification events have expired.
- Authoritative Relational Store: Uses PostgreSQL to store file and folder relationships, enforce uniqueness constraints, and support ACID transactions for metadata operations.
- Low Latency Cache Aside: Employs Redis to cache frequently accessed directory trees, user permissions, and file metadata to reduce read load on primary database instances.
- Atomic Version Commit: When an upload finishes, the service executes one ACID transaction that creates the new version entry, updates the active chunk hash array, changes reference counters for the chunks whose references changed, and writes an event to the transactional outbox for reliable Kafka dispatch.
- Folder Move Validation: Before changing a folder's parent_id, the service verifies that the target parent exists, is itself a folder, and is not the folder being moved or one of its descendants. The parent change and related permission checks commit atomically.
Notification Service
The notification service delivers real time synchronization alerts to registered user devices and collaborators when files are modified. Device fan out is asynchronous and bounded so a large shared resource does not block the metadata commit path.
- Persistent Edge Connections: Maintains bidirectional WebSocket connections or HTTP long polling channels across millions of active desktop and mobile clients.
- Kafka Driven Dispatch: Consumes file-changed events from Kafka, resolves the affected owner and shared-resource subscriber set, and asynchronously delivers compact delta notifications containing the updated version identifier.
- Offline Catch Up Protocol: When a disconnected device regains network connectivity, it issues a cursor based query to the Metadata Service for changes in each synchronization scope since its last recorded cursor.
Conflict Resolution
Concurrent edits across offline or disconnected devices are handled gracefully through optimistic concurrency control without silent data loss.
Dropbox resolves concurrency by detecting revision mismatches during metadata commits. When two clients edit the same file simultaneously, the first commit succeeds in updating the primary version. The second client commit is rejected for revision conflict, prompting the sync agent to upload the alternate file under an explicit conflicted copy name such as quarterly_financials (Alice's conflicted copy).pdf. This preserves both complete versions on disk, eliminates the risk of overwriting collaborator edits, and allows users to manually merge changes at their convenience.
Event Bus Design (Kafka)
An asynchronous event bus decouples the metadata commit path from downstream notification delivery, search indexing, and audit logging.
# Kafka Event Bus Topology for File State Propagation and Sync Notifications
topics:
file-changed:
partitions: 128
partition_key: "owner_id (preserves file tree ordering for each user)"
retention: "7 days"
producers:
- "Metadata Service through the transactional outbox and CDC"
consumers:
- "Notification Service"
- "Sync Workers"
- "Elasticsearch Indexer"
payload_schema:
file_id: "UUID"
owner_id: "UUID"
version: "integer (committed metadata version)"
change_type: "create | update | rename | move | delete | share"
is_deleted: "boolean"
manifest_version: "integer (committed chunk manifest version)"
modified_by: "UUID"
timestamp: "ISO-8601 string"
sync-events:
partitions: 64
partition_key: "user_id"
producers:
- "Notification Service"
consumers:
- "WebSocket gateways for each device"
cluster_configuration:
replication_factor: 3
min_insync_replicas: 2
producer_guarantees: "Idempotent publishing with enable.idempotence=true and a transactional outbox on metadata commits"
dead_letter_queue: "file-changed-dlq"
processing_paths:
synchronous_path: "Client verifies uploaded chunks, the Metadata Service commits the metadata transaction and outbox event in PostgreSQL, then returns HTTP 201 Created without waiting for Kafka delivery"
asynchronous_path: "CDC publishes file-changed events to Kafka. Notification workers deliver delta notifications, Sync Workers process downstream state, and the Elasticsearch Indexer updates derived search documents."API Design
Sync Service Contracts
The synchronization protocol exposes RESTful endpoints for chunk upload initialization, direct to storage binary transfers, upload completion, file downloads, and cursor based delta synchronization.
// Core cloud file storage and delta sync API contracts
interface InitiateUploadRequest {
fileName: string;
fileSizeBytes: number;
parentFolderId: string;
chunkHashes: string[]; // Ordered list of SHA-256 chunk hashes
}
interface InitiateUploadResponse {
fileId: string;
uploadId: string;
chunksNeeded: string[]; // Chunks missing from storage
uploadUrls: Record<string, string>; // Presigned S3 PUT URLs keyed by chunk hash
expiresInSeconds: number;
}
interface CompleteUploadRequest {
uploadId: string;
baseVersion: number;
chunkHashesOrdered: string[];
}
interface CompleteUploadResponse {
fileId: string;
version: number;
status: "committed";
}
interface ChunkDownloadDescriptor {
chunkHash: string; // SHA-256 of canonical uncompressed content
downloadUrl: string;
byteSize: number;
compression: "none" | "zstd" | "lz4";
}
interface FileDownloadResponse {
fileId: string;
fileName: string;
fileSizeBytes: number;
requestedVersion: number;
chunks: ChunkDownloadDescriptor[];
nextCursor?: string;
hasMoreChunks: boolean;
}
interface ShareResourceRequest {
resourceId: string;
resourceType: "file" | "folder";
userEmail: string;
permission: "view" | "edit";
}
interface CreateShareLinkRequest {
resourceId: string;
resourceType: "file" | "folder";
permission: "view" | "edit";
expiresAt?: string;
password?: string;
}
interface CreateShareLinkResponse {
shareId: string;
shareUrl: string;
expiresAt?: string;
}
interface CreateFolderRequest {
name: string;
parentFolderId?: string;
}
interface RenameResourceRequest {
resourceId: string;
newName: string;
}
interface MoveResourceRequest {
resourceId: string;
newParentFolderId: string;
}
interface RestoreVersionRequest {
version: number;
}
interface RestoreVersionResponse {
fileId: string;
sourceVersion: number;
newVersion: number;
status: "restored";
}
interface FileDeltaChange {
fileId: string;
name: string;
parentFolderId: string;
version: number;
changeType: "create" | "update" | "rename" | "move" | "delete" | "share";
isDeleted: boolean;
updatedAt: string;
}
interface SyncChangesResponse {
cursor: string;
hasMore: boolean;
changes: FileDeltaChange[];
}
// Client synchronization and metadata service interface
interface CloudStorageSyncService {
initiateUpload(request: InitiateUploadRequest): Promise<InitiateUploadResponse>;
uploadChunk(presignedUrl: string, chunkData: Uint8Array): Promise<void>;
completeUpload(fileId: string, request: CompleteUploadRequest): Promise<CompleteUploadResponse>;
downloadFile(fileId: string, version?: number, cursor?: string): Promise<FileDownloadResponse>;
shareResource(request: ShareResourceRequest): Promise<{ shareId: string }>;
createShareLink(request: CreateShareLinkRequest): Promise<CreateShareLinkResponse>;
createFolder(request: CreateFolderRequest): Promise<{ folderId: string }>;
renameResource(request: RenameResourceRequest): Promise<void>;
moveResource(request: MoveResourceRequest): Promise<void>;
restoreVersion(fileId: string, request: RestoreVersionRequest): Promise<RestoreVersionResponse>;
getChangesSince(cursor: string, limit?: number): Promise<SyncChangesResponse>;
}Initiate Chunked Upload
The client initiates an upload by sending the file name, total byte size, destination folder, and an ordered list of SHA-256 chunk hashes. The server inspects its database and returns presigned S3 upload URLs strictly for chunks that are not already present in storage.
POST /api/v1/files/upload/init HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
Content-Type: application/json
{
"file_name": "quarterly_financials.pdf",
"file_size": 10485760,
"parent_folder_id": "8f3a9e1b-4c2d-4e5f-9a1b-3c4d5e6f7a8b",
"chunk_hashes": [
"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c",
"b8c7d6e5f4a3b2c1d0e9f8a7b6c5d4e3f2a1b0c9d8e7f6a5b4c3d2e1f0a9b8c7"
]
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"file_id": "d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"upload_id": "upl_9f8e7d6c-5b4a-3c2d-1e0f-9a8b7c6d5e4f",
"chunks_needed": [
"4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c"
],
"upload_urls": {
"4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c": "https://storage.driveplatform.com/chunks/4a/4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c?signature=xyz789&expires=3600"
},
"expires_in_seconds": 3600
}Upload Chunk Binary Data
The client streams binary chunk payloads directly to object storage using time limited presigned S3 URLs, bypassing application servers to maximize throughput.
PUT /chunks/4a/4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c?signature=xyz789&expires=3600 HTTP/1.1
Host: storage.driveplatform.com
Content-Type: application/octet-stream
Content-Length: 4194304
[Raw binary chunk payload: 4,194,304 bytes]
HTTP/1.1 200 OK
x-amz-checksum-sha256: <base64-sha256-of-uploaded-payload-bytes>
x-amz-meta-canonical-sha256: <sha256-of-uncompressed-chunk-content>Complete Upload
Once all missing chunks are confirmed in object storage, the client submits the ordered hash manifest to finalize the version commit.
POST /api/v1/files/upload/complete HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
Content-Type: application/json
{
"upload_id": "upl_9f8e7d6c-5b4a-3c2d-1e0f-9a8b7c6d5e4f",
"base_version": 0,
"chunk_hashes_ordered": [
"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c",
"b8c7d6e5f4a3b2c1d0e9f8a7b6c5d4e3f2a1b0c9d8e7f6a5b4c3d2e1f0a9b8c7"
]
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"file_id": "d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"version": 1,
"status": "committed"
}Download File Manifest
To retrieve a file or a specific historical version, the client requests its download manifest. The response contains an ordered chunk list and temporary download URLs. Large manifests are paginated so the client can reconstruct the file incrementally.
GET /api/v1/files/d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d/download?version=1 HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"file_id": "d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"file_name": "quarterly_financials.pdf",
"file_size": 10485760,
"requested_version": 1,
"has_more_chunks": false,
"chunks": [
{
"chunk_hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"download_url": "https://storage.driveplatform.com/chunks/e3/e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855?signature=abc123",
"byte_size": 4194304,
"compression": "none"
},
{
"chunk_hash": "4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c",
"download_url": "https://storage.driveplatform.com/chunks/4a/4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c?signature=def456",
"byte_size": 4194304,
"compression": "none"
},
{
"chunk_hash": "b8c7d6e5f4a3b2c1d0e9f8a7b6c5d4e3f2a1b0c9d8e7f6a5b4c3d2e1f0a9b8c7",
"download_url": "https://storage.driveplatform.com/chunks/b8/b8c7d6e5f4a3b2c1d0e9f8a7b6c5d4e3f2a1b0c9d8e7f6a5b4c3d2e1f0a9b8c7?signature=ghi789",
"byte_size": 2097152,
"compression": "none"
}
]
}Share File or Folder
Users grant access permissions to collaborators by specifying their email address and permission level.
POST /api/v1/files/d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d/share HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
Content-Type: application/json
{
"user_email": "colleague@example.com",
"permission": "edit"
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"share_id": "shr_3c4d5e6f-7a8b-9c0d-1e2f-3a4b5c6d7e8f",
"permission": "edit",
"status": "granted"
}Create Share Link
Users can create a public or restricted link for a file or folder. The link carries a scoped permission, can expire independently of membership permissions, and can require a password.
POST /api/v1/files/d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d/share-link HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
Content-Type: application/json
{
"resource_id": "d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"resource_type": "file",
"permission": "view",
"expires_at": "2026-09-06T10:15:30Z"
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"share_id": "shr_8a7b6c5d-4e3f-2d1c-0b9a-8f7e6d5c4b3a",
"share_url": "https://driveplatform.com/s/shr_8a7b6c5d",
"expires_at": "2026-09-06T10:15:30Z"
}Restore Historical Version
A restore operation creates a new active version that references the selected historical manifest. It does not rewind the current version number, so existing sync cursors and audit history remain monotonic.
Get Changes Since Cursor
Clients retrieve incremental directory tree changes using cursor based pagination, allowing quick recovery after network interruptions or offline work sessions.
GET /api/v1/sync/changes?cursor=cur_948f2b1a&limit=100 HTTP/1.1
Host: api.driveplatform.com
Authorization: Bearer <token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"cursor": "cur_a8b7c6d5",
"has_more": false,
"changes": [
{
"file_id": "d1c2b3a4-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"name": "quarterly_financials.pdf",
"parent_folder_id": "8f3a9e1b-4c2d-4e5f-9a1b-3c4d5e6f7a8b",
"version": 2,
"change_type": "update",
"is_deleted": false,
"updated_at": "2026-09-05T10:15:30Z"
}
]
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 422 Unprocessable Entity: password does not meet security requirements or MFA verification is required 423 Locked: account temporarily locked after repeated failed login attempts 440 Login Timeout: WebSocket connection session expired, client reconnect is required 400 Bad Request (INVALID_CHUNK_HASH): provided SHA-256 hash does not match uploaded chunk payload 403 Forbidden (INSUFFICIENT_PERMISSIONS): authenticated user lacks required view or edit permissions on target resource 404 Not Found (FILE_NOT_FOUND): specified file or folder identifier does not exist or has been soft deleted 409 Conflict (VERSION_MISMATCH): client base version is stale. Server generates a conflicted copy to preserve data 413 Payload Too Large (QUOTA_EXCEEDED): upload would exceed user storage quota limit
Data Model
PostgreSQL: File Metadata
The data layer separates relational metadata for directory hierarchies and file versions from content addressed binary chunk storage in S3.
Stores directory tree nodes with unique constraints on parent folder and name to prevent duplicate siblings, using index structures optimized for folder listings.
CREATE TABLE files (
file_id UUID PRIMARY KEY,
owner_id UUID NOT NULL,
name VARCHAR(256) NOT NULL,
parent_id UUID REFERENCES files(file_id) ON DELETE RESTRICT, -- Parent folder UUID (NULL for root)
is_folder BOOLEAN DEFAULT FALSE,
size BIGINT,
mime_type VARCHAR(128),
current_version INT DEFAULT 1,
is_deleted BOOLEAN DEFAULT FALSE, -- Soft delete tombstone
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
updated_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
deleted_at TIMESTAMP WITH TIME ZONE NULL
);
-- PostgreSQL NULL values do not make a normal UNIQUE(parent_id, name, owner_id) constraint unique for root entries.
CREATE UNIQUE INDEX uq_files_child_name
ON files(owner_id, parent_id, name)
WHERE parent_id IS NOT NULL AND is_deleted = FALSE;
CREATE UNIQUE INDEX uq_files_root_name
ON files(owner_id, name)
WHERE parent_id IS NULL AND is_deleted = FALSE;
CREATE INDEX idx_files_owner
ON files(owner_id);
CREATE INDEX idx_files_parent
ON files(parent_id);PostgreSQL: File Versions
Maintains an immutable historical record of every file revision, linking each version to its ordered array of SHA-256 chunk hashes.
CREATE TABLE file_versions (
version_id UUID PRIMARY KEY,
file_id UUID NOT NULL REFERENCES files(file_id) ON DELETE RESTRICT,
version_number INT NOT NULL,
size BIGINT NOT NULL,
chunk_hashes TEXT[] NOT NULL, -- Ordered SHA-256 chunk digests
modified_by UUID NOT NULL,
legal_hold BOOLEAN NOT NULL DEFAULT FALSE, -- Prevent permanent purge during legal hold
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
UNIQUE (file_id, version_number)
);
-- For very large manifests, normalize chunk_hashes into a child table keyed by file_id and version_number.
-- That keeps metadata rows bounded while preserving ordered manifest reads.PostgreSQL: Chunks (Content Addressable)
Tracks global content addressed chunks with atomic reference counting to support cross user deduplication and asynchronous garbage collection. The canonical chunk hash covers the uncompressed content, while the stored representation may use the deployment's selected compression codec.
CREATE TABLE chunks (
chunk_hash VARCHAR(64) PRIMARY KEY, -- SHA-256 hexadecimal digest of uncompressed content
size INT NOT NULL, -- Canonical uncompressed byte size
storage_path TEXT NOT NULL, -- S3 object key
reference_count INT NOT NULL DEFAULT 0, -- References from committed file versions
compressed_size INT, -- Stored size when compression is enabled
compression VARCHAR(8) DEFAULT 'none',
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);
-- Reference count changes occur in the same PostgreSQL transaction as the file version commit.
-- Garbage collection rechecks metadata after a grace period before deleting an object whose count is zero.PostgreSQL: Sharing and Permissions
Manages detailed user access control lists on individual files and shared folders. Public and restricted share links use separate capability records so link expiry does not change membership permissions.
CREATE TABLE sharing (
share_id UUID PRIMARY KEY,
file_id UUID REFERENCES files(file_id) ON DELETE CASCADE,
folder_id UUID REFERENCES files(file_id) ON DELETE CASCADE,
shared_with UUID NOT NULL,
permission VARCHAR(16) NOT NULL CHECK (permission IN ('view', 'edit')),
shared_by UUID NOT NULL,
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
CHECK ((file_id IS NOT NULL) <> (folder_id IS NOT NULL))
);
CREATE UNIQUE INDEX uq_sharing_file_user
ON sharing(file_id, shared_with)
WHERE file_id IS NOT NULL;
CREATE UNIQUE INDEX uq_sharing_folder_user
ON sharing(folder_id, shared_with)
WHERE folder_id IS NOT NULL;
-- Public and restricted share links are separate capabilities. Their expiration can change without changing membership ACLs.
CREATE TABLE share_links (
share_id UUID PRIMARY KEY,
file_id UUID REFERENCES files(file_id) ON DELETE CASCADE,
folder_id UUID REFERENCES files(file_id) ON DELETE CASCADE,
token_hash CHAR(64) NOT NULL UNIQUE,
permission VARCHAR(16) NOT NULL CHECK (permission IN ('view', 'edit')),
created_by UUID NOT NULL,
password_hash TEXT,
expires_at TIMESTAMP WITH TIME ZONE,
created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
CHECK ((file_id IS NOT NULL) <> (folder_id IS NOT NULL))
);Deletion and Garbage Collection
A delete operation first marks the file or folder as deleted in PostgreSQL and records the change in the transactional outbox. A folder delete hides the subtree immediately at the metadata and authorization layer and emits a subtree deletion marker so clients can treat all descendants as deleted even while individual descendant tombstones are materialized asynchronously. The soft delete keeps historical versions available during the retention window. When a version becomes eligible for permanent purge, the purge worker removes the version metadata and decrements the reference counters for its chunks in a PostgreSQL transaction. A separate garbage collector deletes an S3 object only after the reference count reaches zero, the safety grace period has elapsed, and no legal hold still applies. Public share links are revoked as soon as the resource is soft deleted. A version marked with legal_hold remains protected from permanent purge.
Durable Sync Change Log
Cursor based offline catch up should not depend on Kafka retention alone. The Metadata Service writes a compact durable change journal alongside the metadata transaction or materializes it from the file change stream. Each record contains the resource owner, the affected file identifier, the version, the change type, the deletion state, the timestamp, and a monotonically increasing sequence within the synchronization scope. A synchronization scope represents either the user owned tree or a shared resource visible to that user. Online devices can use Kafka driven notifications for low latency delivery, while devices that reconnect after the Kafka retention window resume from the durable journal using the appropriate scope sequence as their synchronization cursor. This ensures shared folder changes are included even when the original Kafka notification has expired.
S3: Content Addressed Chunk Storage
# S3 Object Storage Layout for Content Addressed Chunks
bucket_name: "file-chunks-prod"
prefix_format: "chunks/{sha256_prefix_2chars}/{full_sha256_hash}"
example_key: "chunks/4a/4a5a9c9f6d7a8b9c0d1e2f3a4b5c6d7e8f9a0b1c2d3e4f5a6b7c8d9e0f1a2b3c"
storage_properties:
immutability: "Immutable objects that are never modified in place"
compression: "The deployment may use one canonical client side zstd or LZ4 representation. The object checksum validates the uploaded payload, while chunk_hash remains the SHA-256 digest of the canonical uncompressed content"
encryption: "Server side AES-256 with KMS envelope encryption"
lifecycle_rules:
active_tier: "S3 Standard for newly written chunks"
infrequent_access: "S3 Infrequent Access for older chunks after 30 days under the lifecycle policy"
archive_tier: "S3 Glacier Flexible Retrieval for chunks after 90 days under the lifecycle policy"Kafka Topics Summary
# Kafka Topic Architecture for File Synchronization
topics:
file-changed:
partition_key: "owner_id"
description: "Emitted when a file or directory changes. The event identifies the committed version and change type so consumers can fetch authoritative metadata or the file manifest when needed."
sync-events:
partition_key: "user_id"
description: "Carries delta notifications for each device through persistent WebSocket gateways"Fault Tolerance
Specific: Data Integrity and Corruption Prevention
The system is designed to preserve data integrity, support fault recovery, and maintain high availability across network disconnections, storage failures, and concurrent modification conflicts.
| Concern | Solution |
|---|---|
| Chunk loss in S3 | The storage layer targets 99.999999999% (11 nines) object durability, backed by asynchronous cross region replication for disaster recovery. |
| Metadata DB loss | PostgreSQL primary with synchronous replication to an in region standby and daily snapshot backups to S3. |
| Partial upload failure | Resumable chunk uploads allow the client to retry only failed or missing chunks without reuploading the entire file. |
| Sync conflict | The system creates an explicit conflicted copy that preserves both revisions without silently overwriting data. |
| Client crash during sync | After restart, the client compares its local hash state with server metadata and requests only the missing delta chunks. |
| Dedup reference counting | Reference counts change inside metadata transactions, and background garbage collection removes a chunk only after the reference count reaches zero and a safety grace period has elapsed. |
- End to End Cryptographic Checksums: The client calculates a SHA-256 digest for each uncompressed chunk. The storage verification path decompresses the object when needed and checks the original chunk content against the recorded digest, and the downloading client verifies the digest again before reassembling the file on disk.
- Robust Encryption at Rest and Transit: Chunks are transmitted over TLS 1.3 and encrypted server side using AES-256 with KMS key rotation, with optional client side zero knowledge encryption where users retain private decryption keys.
- Immutable Write Once Storage: Content addressed chunks in S3 are write once and never updated in place. Modifying file content generates new chunks under distinct hashes, so an existing chunk object is never overwritten in place.
Additional Considerations
Access Control Lists and Sharing Model
Access control inheritance, storage lifecycle tiering, quota attribution, and real time collaborative editing represent advanced architectural considerations for enterprise deployments.
Each file and directory record stores ownership metadata and permission references. Shareable links use cryptographically signed capability tokens with limited lifetime and explicit read or edit permissions, with optional passwords and expiration times. The Metadata Service evaluates permissions on every metadata request so unauthorized clients never receive valid chunk hashes or download URLs. Team folders support downward permission inheritance, while an explicit child ACL can override inherited access.
Client Synchronization Protocol
The delta synchronization endpoint accepts a local file manifest that maps paths to ordered chunk hashes. The server compares that manifest with PostgreSQL metadata and returns only the chunks that are missing along with remote metadata changes that the client must apply. Cursor based pagination streams large change sets without loading an entire directory tree into memory. When concurrent edits are detected, the server keeps both versions and instructs the client to persist the alternate version as a conflicted copy. A restore operation creates a new active version that references the selected historical manifest instead of rewinding the current version number.
Bandwidth Optimization
- Delta Synchronization: By transferring only modified chunks, content defined chunking ensures that small edits within large documents consume negligible network bandwidth.
- Payload Compression: Applying zstd or LZ4 compression to chunks prior to transmission reduces network egress volume and object storage consumption by an average of 20 to 40 percent.
- Global Deduplication: Identical chunks uploaded across thousands of users are stored once globally, reducing total platform storage footprint by 30 to 50 percent.
Storage Lifecycle Tiering
- Hot Storage (S3 Standard): Retains frequently accessed files and newly modified chunks from the preceding 30 days in the hot S3 Standard tier.
- Warm Storage (S3 Infrequent Access): Automatically transitions chunks that have been idle for 30 to 90 days to a lower cost storage tier according to the lifecycle policy.
- Cold Archive (S3 Glacier): Moves chunks that have remained untouched for more than 90 days into S3 Glacier Flexible Retrieval to reduce ongoing storage costs.
Quota Management and Accounting
- Tier Allocations: Supports a 15 GB free tier alongside paid enterprise tiers ranging from 100 GB to unlimited storage.
- Usage Attribution: A user's quota is measured using the logical bytes referenced by the user's files, independent of physical savings from cross user deduplication.
- Shared Folders: Shared file storage is billed against the file owner rather than invited viewers or editors, preventing quota double charging.
Enterprise Security and Governance
- Administrative Audit Logging: Records immutable audit trails capturing every file access, modification, deletion, and sharing event.
- Data Loss Prevention (DLP): Scans uploaded files asynchronously for sensitive data such as credit card numbers, personal health information, or API credentials.
- Legal Hold: Allows administrators to freeze file versions and prevent permanent deletion during active regulatory or legal investigations.
Real Time Collaborative Editing Extension
While file synchronization systems like Dropbox operate on complete file chunks, real time multi user document editors require character level synchronization:
- Operational Transformation (OT): Employs a centralized server to serialize and transform concurrent editing operations, as explored in Google Docs Collaborative Editing.
- Conflict Free Replicated Data Types (CRDT): Uses unique operation or element identifiers and deterministic merge metadata so clients can reconcile concurrent edits without requiring a single centralized ordering point.
Interview Walkthrough Strategy
- 25 Minute Interview Strategy
Prioritize content defined chunking and metadata separation before diving into staff level storage math or cross user deduplication privacy trade offs.
- Functional and Non Functional Requirements (3 min)
- Metadata Database vs Blob Storage Separation (4 min)
- Content Defined Chunking and Deduplication Mechanics (7 min)
- Delta Synchronization Protocol and Conflict Resolution (6 min)
- Resumable Uploads and Storage Lifecycle Tiering (5 min)
- Separate directory metadata in PostgreSQL from raw binary chunks in S3 early in the interview using System Design Interview Patterns, explaining why relational metadata and content addressed blobs have completely different consistency profiles.
- Contrast fixed size chunking with Content Defined Chunking (CDC) using Rabin fingerprints, demonstrating how a 1 byte insertion at position zero can invalidate every fixed block while CDC limits the changed region until the content dependent boundaries realign.
- Explain content addressable storage where chunk SHA-256 hashes act as S3 keys, enabling global deduplication across files and user accounts to save 30 to 50 percent of total storage capacity.
- Apply Consistent Hashing to distribute chunk hashing workloads and partition block storage gateways across servers evenly without cascading rebalances during scale out.
- Detail the delta sync protocol where the client compares its local hash list against server metadata, downloading only missing chunks and uploading only newly generated deltas.
- Walk through conflict resolution for offline concurrent edits, explaining why general purpose file storage platforms adopt the conflicted copy pattern rather than complex collaborative text algorithms.
- Quantify network bandwidth using Back-of-the-Envelope Estimation, calculating that a 1 KB modification to a 100 MB file reduces network transfer from 100 MB down to approximately 4 MB with CDC.
- Highlight the common architectural anti pattern of proxying multi gigabyte file uploads through application servers, demonstrating that clients must stream binary chunks directly to S3 via pre signed URLs.
Related Problems
File chunking, content addressable storage, and collaborative editing connect directly to these related system design architectures:
- Google Docs Collaborative Editing explores real time character level synchronization using Operational Transformation and CRDTs.
- Blob Storage (S3) examines distributed content addressed object storage, LSM tree chunk indexing, and multi region durability.
- P2P File Transfer (BitTorrent) details decentralized chunk verification, torrent piece selection algorithms, and Merkle tree hash integrity.
- Storage Types: Block, File, and Object Storage provides foundational comparisons across storage abstractions and file system semantics.
- Merkle Tree explores hierarchical cryptographic data verification for rapid diff detection across distributed replicas.
Engineering Trade-offs
Fixed Size Chunking vs Content Defined Chunking (CDC)
Interview discussions frequently challenge your choices on chunking algorithms, metadata database selection, and conflict handling. Walk through each trade off and state what you would pick for a cloud storage platform operating at Dropbox scale.
Fixed size chunking divides files into static byte blocks such as 4 MB segments. When a user inserts a single byte at the beginning of a document, later boundaries shift and the resulting hashes can change across the rest of the file, which can force a full reupload. Content Defined Chunking instead uses a sliding window rolling hash such as a Rabin fingerprint to place boundaries according to the data content rather than byte offset. An insertion changes a local region until the boundary pattern realigns. For a 100 MB file with a 1 KB modification, the example assumes CDC transfers approximately 4 MB compared with 100 MB for fixed size chunking, giving a 25 fold reduction in bandwidth.
Why PostgreSQL for Metadata (Not MongoDB or Cassandra)
File metadata workloads require strong relational capabilities including directory hierarchy traversal, B tree indexes on parent folder identifiers, and strict uniqueness constraints on directory names. PostgreSQL provides ACID transaction guarantees to ensure that creating a file version and updating its chunk list happen atomically. While NoSQL databases like MongoDB or Cassandra offer horizontal scalability, Cassandra lacks relational joins, transactions, and unique constraints, making directory tree consistency fragile. PostgreSQL with horizontal sharding by owner_id and a Redis cache aside layer provides both the relational integrity and the throughput required at exabyte scale.
Conflict Resolution: Operational Transformation vs CRDT vs Conflicted Copies
Real time collaboration engines such as Google Docs use Operational Transformation with centralized operation ordering. Some collaborative editors use Conflict Free Replicated Data Types to merge concurrent edits without a single centralized coordinator. These techniques require document aware data models and can carry additional memory and algorithmic complexity. General purpose cloud storage platforms synchronize arbitrary binary files, including compiled binaries, videos, and spreadsheets, where automatic merging is generally not possible. Dropbox style storage uses optimistic concurrency control with the conflicted copy pattern, preserving both versions when a conflict occurs.
Deduplication: Within User vs Cross User
Within user deduplication stores duplicate files or chunks only once per account. It limits storage savings compared with global deduplication. Cross user deduplication compares chunk hashes across all users and can save 30 to 50 percent of total platform storage. It also introduces an existence disclosure risk because an attacker may infer whether another user stores a specific file by observing whether an upload initialization skips the chunk. Production systems can mitigate this through blinded hash verification, rate limits, or by accepting the privacy trade off in exchange for the storage savings.
Upload Strategy: Whole File vs Chunked Direct Upload vs Direct Streaming
Uploading files as monolithic single objects creates severe reliability issues because a network failure near the end of a 10 GB transfer can force a restart from byte zero. Streaming large payloads through application servers also turns backend instances into high bandwidth bottlenecks. The preferred architecture uses direct chunk uploads with pre signed S3 URLs. The client uploads multiple 4 MB chunks in parallel directly to object storage, retries failed chunks independently, skips chunks that already exist through deduplication, and keeps application servers focused on metadata operations.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.