System Design Problem

Design a Video Streaming Platform (YouTube / Netflix)

Commonly Asked By:NetflixGoogleAmazonDisney

Interview Setup

Interview Prompt

Design a video streaming platform like YouTube or Netflix. Users upload videos, the platform transcodes them into multiple quality levels, and viewers stream with adaptive bitrate playback. Support upload, playback, and view count analytics.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Should the platform support VOD, live streaming, or both?VOD follows an ingest, transcode, and CDN edge caching workflow, whereas live streaming introduces RTMP or SRT ingest, low-latency segment publishing, and no full-file pre-transcoding.
What is the expected average video duration and daily upload volume?A 10-minute average across 500,000 uploads per day dictates an average ingest workload of 58 video-hours per minute, while sizing transcoding worker farms for a peak platform capacity headroom of 500 video-hours per minute.
Do we need digital rights management or are signed URLs sufficient?Premium subscription catalogs require Widevine or FairPlay license servers with CENC encryption, whereas user-generated platforms often rely on time-bounded pre-signed CDN cookies or URLs.
View counts: should updates be real-time or eventually consistent?Synchronous updates against a relational table cause database lock exhaustion on viral videos, making stream processing in Apache Flink and Redis serving accelerators the industry standard for low-latency display counts.

Scope

In scope

  • Video upload with resumable S3 multipart ingest and transactional outbox
  • Adaptive bitrate playback (HLS and DASH with CMAF packaging)
  • CDN edge delivery with edge signed URLs and origin shielding
  • High-throughput view count aggregation pipeline (Kafka, Flink, Redis, Cassandra)
  • Capacity estimation for storage growth and peak egress bandwidth
  • Recommendation integration (candidate feed API and watch telemetry ingestion)

Out of scope (state explicitly)

Functional Requirements

Start by clarifying video-on-demand vs live streaming, upload and transcoding scope, and whether recommendations operate in-band. Video-on-demand adaptive bitrate streaming and global CDN delivery are core to the platform architecture.

  • Upload videos: Users upload videos with title, description, tags, and thumbnails via direct-to-storage resumable multipart upload.
  • Stream videos: Adaptive bitrate streaming (HLS and DASH with CMAF) dynamically adjusts playback quality based on network throughput.
  • Search videos: Full-text search across titles, descriptions, creator tags, and channels powered by inverted indexes.
  • Recommendations: Personalized candidate video feeds and watch telemetry ingestion (ranking model internals and offline training are out of scope).
  • Channels and Subscriptions: Subscribe to creator channels and receive new content notifications.
  • Engagement and History: Likes, comments, sharing, and watch history with seamless cross-device playback resumption.
  • Live Streaming Architecture: Architecture extension supporting low-latency ingest, real-time chunk transcoding, and LL-HLS delivery.
  • Monetization and Content Protection: Premium subscription access, dynamic ad insertion markers, and Common Encryption (CENC) digital rights management.

Non-Functional Requirements

Interviewers focus heavily on video startup latency and rebuffering ratios. Because egress bandwidth constitutes the primary ongoing operational expense, edge caching and origin shielding strategies must be established before designing individual microservices.

  • Platform Control-Plane Availability: 99.99% uptime for upload initiation, user authentication, and video metadata CRUD services (at most 4.3 minutes of cumulative downtime per month).
  • Playback Streaming Success: 99.9% availability for user-perceived playback delivery (at most 43.8 minutes of cumulative streaming degradation per month).
  • Low Startup Latency (TTFF): Time-to-first-frame under 1.5 seconds at p95 and under 2.0 seconds at p99 from nearest edge PoP.
  • Smooth Playback: Rebuffer ratio under 0.5% of total watch duration under standard network conditions.
  • Scalability: Support 200M DAU, 25M peak concurrent viewers, and 500K daily uploads (58 video-hours per minute baseline workload, with 500 video-hours per minute platform peak ingest capacity headroom).
  • Data Durability: 99.999999999% (11 9's) durability for uploaded master source videos and transcoded media chunks across multi-region object storage.
  • Global Edge Delivery: Worldwide low-latency manifest resolution and segment delivery via distributed multi-PoP Content Delivery Networks.
  • Cost Efficiency: Automated S3 lifecycle tiering from Standard to Glacier, combined with greater than 95% edge cache hit ratios to minimize origin egress charges.

Capacity Estimations

Calculate network bandwidth and storage volume before sizing CDN edge distributions and origin storage. Video bitrate and peak concurrent viewers dictate edge egress bandwidth, while daily upload volume and multi-rendition ladders govern storage growth.

MetricCalculationValue
Daily Active Users (DAU)Given (product assumption)200M users
Videos watched / day200M DAU x ~6 videos / user1.2B views/day (~13,888 views/sec avg, ~50,000 views/sec peak)
Avg video durationGiven (typical workload assumption)10 min
Average streaming bitrateBlended global average across resolutions and network profiles2.5 Mbps
Peak 1080p bitrateBaseline full-HD video profile (highest standard rendition)5 Mbps
Peak concurrent viewersPrimary peak concurrency assumption (evening peak)25M viewers
Peak CDN edge egress bandwidth25M concurrent viewers x 5 Mbps peak bitrate125 Tbps
Peak origin egress bandwidthLogical origin-bound bandwidth before additional origin-shield coalescing: 125 Tbps x 5% edge miss6.25 Tbps
Daily CDN edge egress volume200M DAU x 1 hr/day x 2.5 Mbps blended bitrate225 PB/day (~6.75 EB/month)
Daily origin egress volume225 PB/day x 5% cache miss (logical upper bound before shield coalescing)11.25 PB/day (~337 PB/month)
Videos uploaded / dayGiven baseline workload (~5.8 uploads/sec avg, ~25 uploads/sec peak)500K uploads/day
Avg original video size10 min video at ~6.6 Mbps master ingest bitrate500 MB
Raw master storage / day500K uploads/day x 500 MB250 TB/day
Transcoded storage / day250 TB x 3 assumed encoded/packaged-media expansion factor covering rendition ladders, codecs, audio, and packaging overhead750 TB/day
Total new media storage growth250 TB raw masters + 750 TB packaged renditions1 PB/day (1,000 TB/day)
Existing catalog storageCumulative multi-year video archive~1 EB (exabyte)
Transcoding baseline workload500K uploads/day x 10 min = 5M min/day ≈ 83,333 hours/day58 video-hours/min
Platform ingest capacity headroomPeak ingest capacity target (~8.6x average baseline workload)500 video-hours/min
Playback beacon ingestion rate25M peak viewers emitting heartbeats every 10 seconds2.5M events/sec

Architecture Diagram

In the room: clarify video-on-demand vs live streaming upfront, because live video introduces sub-5-second glass-to-glass latency constraints and real-time segment transcoding.

Walk your interviewer through the architecture by contrasting write and read paths, because video upload and video playback have fundamentally different latency and throughput budgets. Uploads land directly in raw object storage via S3 multipart URLs, record processing state through an atomic MySQL Transactional Outbox, and stream into Kafka via Debezium CDC to trigger distributed transcoding. Worker fleets transcode chunks in parallel across the canonical 5-rendition ladder, packaging files into CMAF-compliant HLS and DASH segments stored in S3 origin buckets. Viewers stream media globally through edge CDNs backed by origin shielding, while lightweight playback heartbeats flow into Kafka and Apache Flink to drive real-time Redis counters and durable ClickHouse analytics without stressing database tables.

Loading...

Component Deep Dives

Each core component addresses distinct challenges across the media lifecycle, starting with the upload and transcoding pipeline, moving to CDN edge distribution, and concluding with asynchronous telemetry and discovery services.

Upload Flow (Transactional Outbox Pattern)

Keep the upload data path completely decoupled from application API servers so video payloads never traverse API gateways. Direct-to-storage transfers prevent application servers from becoming high-bandwidth bottlenecks.

  1. The client requests a pre-signed multipart upload URL from the Upload Service, specifying file size, MIME type, and expected chunk count.
  2. The Upload Service validates creator permissions and generates short-lived pre-signed S3 upload URLs with scoped object keys.
  3. The client uploads raw video chunks directly to S3 in 5 MB to 10 MB parts using S3 multipart upload, allowing retries of failed parts without restarting the upload.
  4. Upon completing all chunk uploads, the client invokes the complete-upload endpoint, supplying part numbers and S3 entity tags (ETags).
  5. The Upload Service verifies S3 chunk completeness and executes an atomic MySQL transaction: updating the video record status to processing and inserting an event into the outbox_events table.
  6. Debezium Change Data Capture reads the MySQL binary log and streams the video-uploaded event to Kafka with at-least-once durability, completely eliminating dual-write failure windows between MySQL and Kafka.
  7. The distributed workflow orchestrator consumes the event from Kafka to schedule the transcoding Directed Acyclic Graph (DAG).

Video Processing Pipeline: Codecs, GOP Chunking, and Packaging

Transcoding represents the most computationally intensive write-path stage, requiring keyframe-aligned segment parallelization and adaptive bitrate packaging.

GOP Alignment and Parallel Chunking: Workers probe the master container and split the video strictly along Group of Pictures (GOP) and Instantaneous Decoder Refresh (IDR) keyframe boundaries into 4-second chunks. Splitting at arbitrary byte offsets is avoided because it corrupts temporal compression references. Distributed GPU workers pull chunks from S3 and transcode them in parallel across the canonical 5 baseline video renditions: 240p (400 Kbps), 360p (800 Kbps), 480p (1.5 Mbps), 720p (3 Mbps), and 1080p (5 Mbps), while extracting and encoding an independent 128 Kbps AAC and Opus audio track. A conditional 4K rendition (15 to 20 Mbps) is generated only when the uploaded master file supplies native 4K resolution, strictly prohibiting artificial upscaling. Workloads produce dual codec ladders: H.264 for universal device playback, alongside HEVC or AV1 for modern hardware to achieve 30% to 50% bandwidth savings.

Adaptive Bitrate Packaging with CMAF: Common Media Application Format (CMAF) standardizes fragmented MP4 (fMP4) containers across streaming protocols. Packagers encapsulate chunked media so a single set of media segments is referenced simultaneously by HLS master playlists (master.m3u8) and DASH manifest descriptors (manifest.mpd), cutting origin storage requirements in half compared to storing separate MPEG-TS and fMP4 assets. Client players monitor buffer occupancy and network throughput: downgrading quality when the buffer falls below 5 seconds, and upgrading when buffer occupancy exceeds 15 seconds with verified bandwidth headroom. Switching occurs strictly at keyframe boundaries to guarantee artifact-free transitions.

Digital Rights Management (DRM) and CENC: Protected catalog content is secured using ISO Common Encryption (CENC AES-128). Video chunks are encrypted using CTR mode (cenc) for DASH, Google Widevine, and Microsoft PlayReady, or CBC mode (cbcs) for HLS and Apple FairPlay. Protection System Specific Header (PSSH) metadata is written directly into media container headers. Authenticated client players exchange session entitlement tokens with a dedicated DRM license server to acquire decryption keys before decoding frames.

CDN and Edge Delivery Architecture

Edge caching is the foundational layer that makes low-latency global video streaming economically viable by absorbing over 95% of traffic before it reaches origin storage.

  • Tiered Caching: Popular videos in the top 20% are preferentially retained across edge PoPs based on access frequency, LRU eviction policies, and cache TTL headers, serving over 80% of streaming bytes directly from edge memory and NVMe drives.
  • Warm Content Handling: Less frequently viewed content is fetched from origin upon first request and cached at edge PoPs with standard time-to-live intervals.
  • Long-Tail Content: Rarely accessed archive assets stream from origin storage through pass-through streaming to prevent cache thrashing and preserve edge capacity for active titles.
  • Global Edge Footprint: Over 200 edge locations (Points of Presence) terminate TLS connections close to viewers worldwide, minimizing round-trip latency.
  • Origin Shielding and Request Coalescing: An intermediate centralized caching layer aggregates regional edge cache misses. Origin shield nodes coalesce concurrent requests for identical missing segments into one or a small number of upstream origin fetches, protecting origin storage from thundering herds.
  • Cache Ingestion Modes: Highly anticipated releases are proactively pre-warmed to regional edge PoPs, while catalog titles use lazy pull-on-demand fetching.

Recommendation Service Integration

Recommendation integration supplies personalized candidate feeds and ingests watch telemetry asynchronously, while machine learning ranking model internals and offline training pipelines remain out of scope.

  • Collaborative Filtering: Analyzes cross-user consumption patterns to identify videos frequently co-watched across similar user cohorts.
  • Content-Based Filtering: Recommends media sharing similar metadata attributes including genre, creator, tags, language, and duration.
  • Deep Learning Embeddings: Ingests audio and video frame embeddings to discover latent visual and acoustic similarities across catalog items.
  • Low-Latency Serving: Pre-computes candidate recommendations using offline Spark pipelines, hydrates results into Redis clusters, and serves recommendations to clients in under 50 milliseconds.

View Count Aggregation Pipeline

View count aggregation buffers high-velocity telemetry writes through Kafka and Apache Flink so viral video playback surges never overwhelm database counters.

For this design, a qualified/countable view is defined as continuous playback exceeding 30 seconds or 25% of total video duration (whichever is shorter), filtering out accidental clicks and immediate bounces. The client video player transmits lightweight telemetry beacons every 10 seconds to a stateless ingestion API that forwards events directly to the view-events Kafka topic. The topic uses composite partition keys (video ID combined with a shard bucket index from 0 to 99) to prevent viral videos from overloading individual Kafka brokers. Downstream Apache Flink applications process partition streams across tumbling windows, deduplicating records by user ID and video ID across 24-hour sliding windows. Flink flushes 1-minute aggregated increments to Redis via atomic INCRBY commands against views:{video_id}, providing an approximate serving counter for frontend display. Authoritative view events are written durably to Cassandra and ClickHouse. An hourly reconciliation batch job recomputes deduplicated totals from ClickHouse, synchronizing MySQL display counters and resetting Redis keys to prevent counter drift.

DAG Scheduler and Workflow Orchestration

The workflow orchestrator coordinates complex task dependencies across distributed worker pools using durable state machines.

The video processing pipeline involves complex dependencies naturally structured as a Directed Acyclic Graph (DAG). Production architectures employ durable workflow engines like Temporal or cadence-based distributed orchestrators with built-in leader election, state checkpointing, and automatic worker failover rather than single-point-of-failure leaders. Independent tasks such as audio extraction, thumbnail generation, and multi-resolution chunk transcoding run concurrently across GPU workers, while the packager waits for all constituent video and audio tracks to complete. Every task is idempotent based on composite keys of video ID, rendition name, and chunk index. If a 1080p transcode task experiences a transient failure while 240p through 720p succeed, the orchestrator retries only the 1080p task up to 3 times via a dead-letter queue. If product policy permits graceful degradation, the platform publishes a baseline manifest containing the available renditions so viewers can begin streaming immediately, updating the master manifest once the 1080p task finishes.

Loading...

Search Service (Elasticsearch)

Video search leverages inverted index architectures established in Search Engine, indexing rich textual metadata rather than raw video bytes.

  • Indexed Documents: Inverted indices cover video titles, full descriptions, creator tags, channel identifiers, and timestamped speech-to-text transcripts.
  • Query Capabilities: Supports full-text search evaluated via BM25 relevance scoring, n-gram prefix autocomplete suggestions, and edit-distance fuzzy matching for spelling mistakes.
  • Asynchronous Synchronization: Metadata mutations committed to MySQL are captured via Debezium Change Data Capture (CDC), streamed through Kafka, and indexed into Elasticsearch within 2 seconds of publication.
  • Tombstone Processing: When a video is soft-deleted in MySQL, CDC propagates a deletion event that purges the document from Elasticsearch, while playback authorization checks verify current database status.

Event Bus Design (Kafka)

The distributed event bus decouples producers from asynchronous consumer fleets and buffers high-volume traffic bursts.

YAML
# Kafka Event Bus Architecture for Video Ingestion and Telemetry
topics:
  video-uploaded:
    partitions: 64 # Partitioned by video_id: preserves sequential job dispatch per video
    retention_days: 7
    producers:
      - "Upload Service via MySQL Transactional Outbox + Debezium CDC (eliminates dual-write failure window)"
    consumers:
      - "DAG Scheduler / transcode worker fleet (initiates distributed transcode pipeline)"

  video-ready:
    partitions: 64 # Partitioned by video_id
    retention_days: 7
    producers:
      - "HLS/DASH packaging task (emitted only after all baseline renditions and manifests are durably committed in S3 and MySQL)"
    consumers:
      - "Video Metadata cache warmers"
      - "Elasticsearch indexer"
      - "CDN pre-warm worker"

  view-events:
    partitions: 512 # Partitioned by composite key (video_id:shard_bucket) to prevent viral video broker hotspots
    retention_hours: 24 # High-volume telemetry stream (2.5M events/sec peak)
    producers:
      - "Stateless Playback Beacon Ingestion API (client heartbeats sent every 10 seconds)"
    consumers:
      - "Apache Flink streaming job (deduplicates views, flushes approximate counts to Redis INCRBY, and writes durable records to Cassandra/ClickHouse)"
    deduplication: "Qualified view after 30 seconds watch time or 25% duration; unique view per (user_id, video_id) per 24-hour sliding window"

  search-index-updates:
    partitions: 32 # Partitioned by video_id
    producers:
      - "Debezium CDC connector on MySQL videos table (captures inserts, metadata updates, and deletion tombstones)"
    consumers:
      - "Elasticsearch synchronizer (asynchronously indexes or deletes documents within 2 seconds)"

operational_guarantees:
  replication_factor: 3
  min_insync_replicas: 2
  synchronous_path: "Upload negotiation -> S3 multipart direct upload -> atomic MySQL transaction (videos row + outbox_events row) -> HTTP 200/201"
  asynchronous_path: "MySQL CDC -> Kafka video-uploaded -> transcode pipeline -> S3 packaging -> video-ready -> search and CDN; playback beacons -> view-events -> Flink -> Redis and ClickHouse"

API Design

Establish clear API contracts for video chunk upload negotiation, adaptive manifest fetching, and asynchronous watch-progress telemetry across the core viewer lifecycle. Note that availableResolutions contains the 5 baseline renditions; 4K appears only when a native-4K source has successfully produced the optional 4K rendition.

Video Streaming Client API Contracts

TYPESCRIPT
// Core video streaming platform client contracts and service signatures
interface InitiateUploadRequest {
  title: string;
  fileSizeBytes: number;
  format: "mp4" | "mov" | "mkv";
}

interface InitiateUploadResponse {
  videoId: string;
  uploadId: string;
  partUploadUrls: { partNumber: number; uploadUrl: string }[];
  expiresInSeconds: number;
}

interface CompletePartItem {
  partNumber: number;
  etag: string;
}

interface CompleteUploadRequest {
  parts: CompletePartItem[];
}

interface CompleteUploadResponse {
  videoId: string;
  status: "processing";
}

interface UpdateMetadataRequest {
  title: string;
  description: string;
  tags: string[];
  category: string;
  visibility: "public" | "unlisted" | "private";
}

interface ManifestResponse {
  videoId: string;
  masterManifestUrl: string;
  thumbnailUrl: string;
  durationSeconds: number;
  // availableResolutions contains the 5 baseline renditions; 4K appears only when a native-4K source has successfully produced the optional 4K rendition.
  availableResolutions: ("240p" | "360p" | "480p" | "720p" | "1080p" | "4k")[];
  audioTracks: { language: string; codec: "aac" | "opus"; bitrateKbps: number }[];
}

interface PlaybackBeaconRequest {
  userId: string;
  videoId: string;
  sessionToken: string;
  watchDurationSeconds: number;
  playbackPositionSeconds: number;
  renderedQuality: "240p" | "360p" | "480p" | "720p" | "1080p" | "4k";
  deviceType: "mobile" | "desktop" | "tv";
}

interface DeleteVideoResponse {
  videoId: string;
  status: "removed";
  tombstone: true;
}

// Client streaming and upload service contracts
interface VideoStreamingService {
  initiateUpload(request: InitiateUploadRequest): Promise<InitiateUploadResponse>;
  completeUpload(videoId: string, request: CompleteUploadRequest): Promise<CompleteUploadResponse>;
  updateMetadata(videoId: string, metadata: UpdateMetadataRequest): Promise<{ status: "processing" | "ready" }>;
  deleteVideo(videoId: string): Promise<DeleteVideoResponse>;
  getStreamManifest(videoId: string): Promise<ManifestResponse>;
  sendPlaybackBeacon(beacon: PlaybackBeaconRequest): Promise<void>;
}

Upload Video Endpoints

HTTP
POST /api/v1/videos/upload-url HTTP/1.1
Host: api.streamplatform.com
Content-Type: application/json

{
  "title": "System Design in 10 Minutes",
  "file_size_bytes": 524288000,
  "format": "mp4"
}

HTTP/1.1 200 OK
Content-Type: application/json

{
  "video_id": "video-uuid",
  "upload_id": "s3-multipart-upload-xyz",
  "part_upload_urls": [
    { "part_number": 1, "upload_url": "https://s3.amazonaws.com/video-originals/video-uuid/part-1?signature=..." },
    { "part_number": 2, "upload_url": "https://s3.amazonaws.com/video-originals/video-uuid/part-2?signature=..." }
  ],
  "expires_in_seconds": 3600
}

POST /api/v1/videos/video-uuid/complete-upload HTTP/1.1
Host: api.streamplatform.com
Content-Type: application/json

{
  "parts": [
    { "part_number": 1, "etag": "\"d41d8cd98f00b204e9800998ecf8427e\"" },
    { "part_number": 2, "etag": "\"0cc175b9c0f1b6a831c399e269772661\"" }
  ]
}

HTTP/1.1 200 OK
Content-Type: application/json

{
  "video_id": "video-uuid",
  "status": "processing"
}

POST /api/v1/videos/video-uuid/metadata HTTP/1.1
Host: api.streamplatform.com
Content-Type: application/json

{
  "title": "System Design in 10 Minutes",
  "description": "Comprehensive high-level architecture walkthrough",
  "tags": ["system design", "tutorial"],
  "category": "education",
  "visibility": "public"
}

HTTP/1.1 201 Created
Content-Type: application/json

{
  "video_id": "video-uuid",
  "status": "processing"
}

DELETE /api/v1/videos/video-uuid HTTP/1.1
Host: api.streamplatform.com
Authorization: Bearer creator-session-token

HTTP/1.1 200 OK
Content-Type: application/json

{
  "video_id": "video-uuid",
  "status": "removed",
  "tombstone": true
}

Stream Video Manifest Endpoint

HTTP
GET /api/v1/videos/video-uuid/manifest HTTP/1.1
Host: api.streamplatform.com
Authorization: Bearer session-token

HTTP/1.1 200 OK
Content-Type: application/json

{
  "manifest_url": "https://cdn.example.com/videos/video-uuid/master.m3u8?token=hmac-signature",
  "thumbnail_url": "https://cdn.example.com/videos/video-uuid/thumb.jpg",
  "duration_seconds": 600
}

Search and Recommendation Endpoints

HTTP
GET /api/v1/search?q=system+design&type=video&sort=relevance HTTP/1.1
Host: api.streamplatform.com

HTTP/1.1 200 OK
Content-Type: application/json

{
  "query": "system design",
  "total_results": 14200,
  "results": [
    {
      "video_id": "video-uuid-1",
      "title": "System Design in 10 Minutes",
      "channel_id": "channel-101",
      "duration_sec": 600,
      "thumbnail_url": "https://cdn.example.com/videos/video-uuid-1/thumb.jpg"
    }
  ]
}

GET /api/v1/recommendations?limit=20 HTTP/1.1
Host: api.streamplatform.com
Authorization: Bearer session-token

HTTP/1.1 200 OK
Content-Type: application/json

{
  "recommendations": [
    {
      "video_id": "video-uuid-2",
      "title": "Kafka Internals Deep Dive",
      "score": 0.94
    }
  ]
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint
403 Forbidden (PLAYBACK_UNAUTHORIZED): expired CDN signature, invalid HMAC token, or geo-restricted content
409 Conflict (INVALID_UPLOAD_STATE): multipart upload was already completed, aborted, or expired
410 Gone (VIDEO_DELETED): video catalog asset was permanently removed by creator or copyright tombstone
423 Locked (VIDEO_PROCESSING): transcoding pipeline has not yet completed baseline playback manifests

Data Model

The data model separates relational transactional metadata and outbox records in MySQL from unstructured media blobs in S3, while streaming high-velocity view events into Cassandra and ClickHouse for durability and analytics.

MySQL: Video Metadata and Transactional Outbox (Sharded by video_id)

SQL
-- Authoritative video catalog metadata
CREATE TABLE videos (
    video_id        BIGINT PRIMARY KEY,
    channel_id      BIGINT NOT NULL,
    title           VARCHAR(100) NOT NULL,
    description     TEXT,
    tags            JSON,
    category        VARCHAR(50),
    duration_sec    INT,
    status          ENUM('uploading', 'processing', 'ready', 'failed', 'removed') NOT NULL DEFAULT 'uploading',
    visibility      ENUM('public', 'unlisted', 'private') NOT NULL DEFAULT 'public',
    view_count      BIGINT DEFAULT 0, -- Periodically materialized display count
    like_count      INT DEFAULT 0,
    dislike_count   INT DEFAULT 0,
    manifest_url    TEXT,
    thumbnail_url   TEXT,
    created_at      TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    updated_at      TIMESTAMP DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
    deleted_at      TIMESTAMP NULL DEFAULT NULL,
    INDEX idx_channel_created (channel_id, created_at DESC)
);

-- Transactional outbox table ensuring dual-write safety
CREATE TABLE outbox_events (
    event_id        BIGINT AUTO_INCREMENT PRIMARY KEY,
    aggregate_type  VARCHAR(50) NOT NULL,
    aggregate_id    BIGINT NOT NULL,
    event_type      VARCHAR(50) NOT NULL,
    payload         JSON NOT NULL,
    created_at      TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    INDEX idx_created (created_at)
);

S3: Video Storage Structure

YAML
# S3 Bucket Hierarchy for Raw Masters and Packaged Assets
bucket_video_originals:
  structure: "/{video_id}/original.mp4"
  lifecycle: "Normally retained as raw source in Glacier Flexible Retrieval after 30 days for future codec re-transcoding; creator deletion, copyright takedown, or legal-erasure tombstones override normal retention and trigger permanent purge after the 30-day safety window"

bucket_video_transcoded:
  structure:
    - "/{video_id}/master.m3u8" # Master HLS manifest referencing variant playlists
    - "/{video_id}/manifest.mpd" # Master DASH manifest referencing adaptation sets
    - "/{video_id}/audio/playlist.m3u8" # Independent 128 Kbps AAC / Opus audio track
    - "/{video_id}/audio/segment_001.mp4"
    - "/{video_id}/240p/playlist.m3u8" # 240p (400 Kbps) video rendition
    - "/{video_id}/240p/segment_001.mp4"
    - "/{video_id}/360p/playlist.m3u8" # 360p (800 Kbps) video rendition
    - "/{video_id}/360p/segment_001.mp4"
    - "/{video_id}/480p/playlist.m3u8" # 480p (1.5 Mbps) video rendition
    - "/{video_id}/480p/segment_001.mp4"
    - "/{video_id}/720p/playlist.m3u8" # 720p (3 Mbps) video rendition
    - "/{video_id}/720p/segment_001.mp4"
    - "/{video_id}/1080p/playlist.m3u8" # 1080p (5 Mbps) video rendition
    - "/{video_id}/1080p/segment_001.mp4"
    - "/{video_id}/4k/playlist.m3u8" # Conditional 4K (15-20 Mbps, only if uploaded master is native 4K)
    - "/{video_id}/4k/segment_001.mp4"

Cassandra: Durable Qualified View Events

To prevent viral video playback events from saturating individual partitions, the composite partition key includes a shard_bucket derived from hash(user_id) % 100. Cassandra stores deduplicated, qualified view records rather than raw 10-second playback heartbeats.

SQL
CREATE TABLE view_events (
    video_id        BIGINT,
    view_date       DATE,
    shard_bucket    INT,
    view_hour       INT,
    user_id         UUID,
    session_id      UUID,
    watch_duration  INT,
    quality         VARCHAR,
    device          VARCHAR,
    country         VARCHAR,
    PRIMARY KEY ((video_id, view_date, shard_bucket), view_hour, user_id, session_id)
);

Redis: View Counters and Hot Video Cache

Redis serves as an ephemeral low-latency read accelerator, while the authoritative source of truth resides in Cassandra and ClickHouse. If Redis loses state, serving counters are rehydrated from authoritative Cassandra and ClickHouse aggregates, with retained Kafka events replayed when catching up from a recent window.

REDIS
# Increment approximate real-time view counter (flushed from Flink 1-minute tumbling windows)
INCRBY views:{video_id} 42

# Cache video metadata for hot playback paths
HSET video:meta:{video_id} title "System Design in 10 Minutes" channel_id "1002" status "ready" manifest_url "https://cdn.example.com/videos/xyz/master.m3u8" thumbnail_url "https://cdn.example.com/videos/xyz/thumb.jpg"
EXPIRE video:meta:{video_id} 3600

# Rebuilding Redis counters on cache loss:
# 1. Query authoritative aggregated view counts from ClickHouse or Cassandra
# 2. SET views:{video_id} <authoritative_count> (retained Kafka events may also be replayed while available within the 24-hour window)

Kafka Topics Summary

YAML
topics:
  video-uploaded: "Emitted via MySQL Transactional Outbox + CDC to trigger distributed worker nodes for transcoding"
  video-ready: "Emitted once all baseline renditions and manifests are committed, notifying cache warmers and search"
  view-events: "High-throughput stream (partitioned by video_id:shard_bucket) ingesting 10s playback heartbeats into Flink"
  search-index-updates: "Captures MySQL CDC stream via Debezium to synchronize or tombstone Elasticsearch documents"

Fault Tolerance

The system is designed to maintain high playback availability and reliable uploads across transcoding worker crashes, CDN edge anomalies, and viral view surges.

ConcernSolution
Upload failureResumable S3 multipart upload where the client retries from the last verified chunk with idempotent part numbers
Transcoding failure
  • Task-level retries on GOP-aligned chunks with dead-letter queue routing
  • optional degraded publish of ready renditions while failed tiers retry
CDN edge failureAnycast routing and DNS health checks redirect viewer traffic to adjacent healthy edge PoPs or alternate CDN providers
Origin failureS3 cross-region replication with multi-region origin shield caches absorbing traffic during regional failover
Video corruptionEnd-to-end cryptographic checksum verification at each ingest and transcode stage, re-transcoding from preserved raw source on mismatch
Popularity surgePredictive CDN pre-warming of manifests and initial segments, request coalescing at origin shield, and sharded counter buckets

Specific: Handling a Viral Video Surge

  1. Cold vs Hot Asset Behavior: On first access, an edge cache miss routes through the origin shield to S3 origin storage. Once popular, edge PoPs retain immutable segments, serving over 95% of reads directly from edge memory and NVMe caches without contacting origin.
  2. Origin Shield Request Coalescing: For residual edge misses, the intermediate origin shield consolidates concurrent requests for the identical segment into one or a small number of upstream origin fetches, preventing origin saturation during traffic spikes.
  3. Manifest vs Segment Hotness: Manifests have shorter TTLs (1 to 2 hours once finalized, or 5 to 30 seconds during live streaming) to allow dynamic variant updates, while media segments are immutable and cached with long TTLs (days to months).
  4. Predictive Pre-warming: Trending videos and promoted content have their master manifests and opening segments pre-warmed to regional CDN edge PoPs prior to broadcast.
  5. Sharded Counter Updates: View count telemetry is sharded across random bucket keys in Kafka and Cassandra to prevent partition hotspots, aggregated in Flink, and flushed to Redis counters.

Additional Considerations

Digital rights management, live streaming architectures, deletion lifecycles, and recommendation ranking represent advanced operational considerations.

Video Segment Prefetching

  • Proactive Buffering: The client video player prefetches the upcoming 2 to 3 segments while actively rendering the current chunk, preventing playback starvation.
  • Seek Cancellation: If the user scrubs forward or backward to an unbuffered timestamp, the player immediately cancels outstanding prefetch requests and begins buffering from the new seek offset.
  • Adaptive Buffer Sizing: When observed network bandwidth is high, the player expands the prefetch window, whereas on constrained links it reduces lookahead buffering to preserve memory.

Thumbnail Generation

  • Automated Extraction: Worker nodes extract candidate frames at 10% timeline intervals and score them using computer vision models to select the most visually engaging frame based on entropy and facial detection.
  • Hover Previews: The pipeline stitches together a lightweight 6-second animated storyboard preview rendered when viewers hover over video thumbnails in navigation feeds.

Copyright and Content ID

  • Audio Fingerprinting: Generates acoustic fingerprints from the extracted audio stream to match waveforms against indexed catalogs of copyrighted musical recordings.
  • Video Fingerprinting: Computes perceptual hashes across video keyframes to detect near-duplicate re-uploads and edited versions.
  • Policy Enforcement: Automated policy rules enforce copyright decisions by blocking public uploads, muting unlicensed audio tracks, monetizing through ad injection, or routing disputes to rights holders.

Multi-Tier Authorization and DRM Security

  • Upload Security: Short-lived S3 pre-signed multipart upload URLs scoped to specific object keys with size and MIME type constraints, valid for 1 hour.
  • Edge Playback Authorization: Time-bounded CDN signed URLs or signed cookies containing an HMAC signature evaluated over wildcard paths like /videos/{id}/*, validated at the edge PoP without contacting origin.
  • Digital Rights Management (DRM): Common Encryption (CENC AES-128) protects licensed content. Video players acquire decryption keys from dedicated Widevine, FairPlay, or PlayReady license servers upon presenting authenticated session tokens.

Cost Optimization and Storage Tiering

  • Storage Tiering: High-traffic videos remain in Standard S3 storage, while assets with fewer than 10 monthly views transition automatically to Glacier Flexible Retrieval.
  • Encoding Efficiency: Transcoding engines avoid upscaling, encoding up to 4K only when the uploaded master file supplies native 4K resolution.
  • Multi-CDN Arbitrage: Directing traffic across multiple CDN vendors reduces egress pricing tiers while safeguarding uptime against vendor-specific edge outages.
  • Source Preservation: Raw masters are normally retained in archival storage so future codec upgrades can be re-transcoded without requesting new uploads from creators. However, creator deletion, copyright takedown, or legal-erasure tombstones override this retention policy, triggering a permanent purge after the 30-day safety window.

Live Streaming Architecture

  • Live Ingest: Broadcasters stream raw video feeds over RTMP or SRT directly into live transcoding gateway servers.
  • Real-Time Transcoding: Live encoding nodes process incoming chunks within sub-second deadlines to keep pace with real-time video feeds.
  • Low-Latency Packaging: Packaging engines generate 2-second segments using Low-Latency HLS (LL-HLS) or chunked transfer DASH to minimize transmission lag.
  • Glass-to-Glass Latency: Architectures maintain end-to-end latency below 5 seconds for standard streams, and under 3 seconds when using LL-HLS partial segments.
  • Live DVR Capabilities: The ingest pipeline continuously persists live video segments into object storage to enable rewind and instant replay functionalities.

Delete and Tombstone Semantics

Video deletion follows a strict lifecycle starting from authoritative catalog state and propagating through derived caches, search indexes, and media storage.

  1. The creator issues a deletion request via the API, which updates the video record in MySQL with status set to removed and records an outbox event.
  2. Debezium CDC streams the deletion tombstone to Kafka.
  3. The metadata service purges cached metadata from Redis (video:meta:{video_id}) and removes the video identifier from recommendation candidate sets.
  4. The search synchronizer receives the event and deletes the corresponding document from the Elasticsearch index.
  5. The CDN invalidation worker submits a cache purge request for the master manifest URL (master.m3u8), while playback authorization endpoints immediately reject new token requests.
  6. An asynchronous cleanup worker purges original master files and packaged media segments from S3 following a 30-day retention safety window.

Related Problems

Video ingestion and chunked transformation pipelines connect directly to Video Transcoding Pipeline. Global edge caching strategies, origin shielding, and pre-signed token validation are explored in detail in Content Delivery Network (CDN). Personalized video suggestions and candidate generation models link to Video Recommendation Engine and Recommendation System. Full-text metadata indexing and token retrieval patterns mirror Search Engine, while automated moderation of media uploads is covered in Content Moderation System. For foundational distributed systems principles, review CDN and Edge Delivery, Network Protocols (HTTP, gRPC, WebSocket, DNS), Scaling 0 to 1M Users, and System Design Interview Patterns.

Interview Walkthrough

  • 25-minute cut

    Skip staff-level live streaming and DRM details unless interviewing for senior or staff roles.

    • Functional and non-functional requirements with VOD vs live scope (3 min)
    • Upload architecture, multipart S3 ingest, and transactional outbox pipeline (7 min)
    • HLS and DASH adaptive bitrate delivery mechanics with CMAF packaging (6 min)
    • CDN placement, origin shielding, and edge signed cookie security (6 min)
    • Asynchronous view count aggregation architecture (3 min)
  • Clarify video-on-demand vs live streaming early, because live video introduces real-time transcoding, short HLS segments, and sub-5-second glass-to-glass latency constraints.
  • Walk through the upload path sequentially, demonstrating how direct-to-S3 ingest triggers asynchronous workers that transcode source media into multiple resolutions, package files as HLS or DASH segments, and persist assets into partitioned object storage.
  • Position a distributed CDN in front of video delivery, explaining how client-side adaptive bitrate algorithms dynamically switch quality levels based on measured network throughput.
  • When scoping for VOD only, establish that low-latency live streaming is out of scope, which allows 6-to-10-second segments where glass-to-glass latency is unconstrained.
  • If live streaming is in scope, target under 5 seconds of latency using 2-second segments and LL-HLS partial segments, discussing how segment duration trades off against CDN edge cache hit ratios.
  • Enforce strict separation between relational metadata in MySQL and raw media segments in S3, ensuring video payloads never traverse backend API application servers.
  • Quantify network bandwidth using Back-of-the-Envelope Estimation, demonstrating an illustrative sub-scenario where 1 million concurrent viewers streaming at 5 Mbps generate 5 Tbps of peak egress, scaling up to 125 Tbps for the primary 25 million peak concurrent assumption.
  • Highlight the common pitfall of streaming video blobs directly from origin storage without an edge CDN, which causes viral traffic surges to exhaust origin bandwidth and take down the platform.

Engineering Trade-offs

Interview discussions frequently evaluate HLS vs DASH protocol selection, multi-codec transcoding investments, and where state aggregation occurs. Walk through each trade-off and state what you would pick for a YouTube-scale platform.

HLS vs DASH vs WebRTC: Choosing the Streaming Protocol

HLS provides universal device compatibility across Apple devices, desktop browsers, native iOS, Android, and smart TVs. DASH provides an open, vendor-neutral standard with broader codec flexibility. In contrast, WebRTC delivers sub-500-millisecond latency for bidirectional communication over UDP and SRTP, but lacks standard HTTP CDN edge caching, standardized CENC DRM, and incurs high media relay compute costs, making it economically unviable for mass-scale video-on-demand streaming to millions of concurrent viewers. Major platforms standardize on HLS and DASH because both leverage standard HTTP CDN edge caching infrastructure, while Common Media Application Format (CMAF) unifies both protocols under single-chunk fragmented MP4 containers.

Codec Selection: H.264 vs H.265 vs AV1

H.264 delivers universal hardware decoding across virtually every consumer device. H.265/HEVC offers 50% higher compression efficiency at equivalent quality but carries complex licensing royalties. AV1 achieves 30% greater bandwidth savings than H.265 as a royalty-free standard, but demands 50 to 100 times higher compute time during encoding. Production platforms address this trade-off through multi-codec encoding ladders, generating H.264 renditions for universal playback while serving AV1 or HEVC variants exclusively to compatible clients to optimize egress costs.

Pre-Signed Upload URL: Why Upload Directly to S3?

A naive approach routing video from the client through the API server into S3 turns application gateways into high-bandwidth bottlenecks and doubles ingress network costs. Direct-to-storage uploads using pre-signed S3 URLs allow the API tier to handle lightweight metadata exclusively, offloading high-volume payload transfers directly to object storage with native resumability through S3 multipart uploads.

View Count: Why Not Just INCREMENT a Database Counter?

At 25 million peak concurrent viewers, issuing SQL UPDATE statements to increment a single row causes severe row lock contention and database exhaustion. Production video architectures buffer client playback heartbeats in memory before flushing batched events to Kafka. A stream processing engine like Apache Flink aggregates view records across tumbling windows, incrementing atomic Redis counters for immediate approximate display while streaming validated counts to ClickHouse for hourly reconciliation.

Storage Tiering: The 80/20 Rule for Video

Video viewership follows a power-law distribution where 20% of catalog assets generate over 80% of total streaming volume. Hot storage holds this top 20% along with all newly uploaded content from the preceding 7 days in SSD-backed S3 Standard. Warm storage maintains assets receiving more than 10 views monthly in S3 Standard-Infrequent Access, while cold storage archives videos with fewer than 10 monthly views that exceed 1 year of age into S3 Glacier. Automated lifecycle rules monitor access frequency and transition segments between storage classes, reducing multi-petabyte storage expenses by millions of dollars annually.

Why MySQL for Video Metadata

Primary video metadata workloads require relational query capabilities such as channel listings sorted chronologically, tag filters, and administrative queries requiring joins. Video metadata payloads are compact at under 1 KB per video, which means 1 billion catalog records occupy only 1 TB of storage and easily fit into a sharded MySQL deployment. While NoSQL datastores like Cassandra excel at high-volume immutable time-series telemetry, they lack relational joining capabilities and introduce consistency complexities when managing relational catalog entities.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...