System Design Problem

Design a Live Streaming Platform like Twitch

Commonly Asked By:AmazonGoogleMetaByteDance

Interview Setup

Interview Prompt

Design a live video streaming platform supporting 100K concurrent streamers and 50M concurrent viewers with low latency (2-5s), interactive chat, and automatic VOD recording.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What's the end-to-end latency target?LL-HLS (2-5s) vs standard HLS (15-30s) vs WebRTC (<1s) fundamentally alters ingest and CDN architecture.
Peak concurrent viewers and streamers?50M viewers x 4 Mbps = 200 Tbps egress: CDN is the entire architecture.
Is chat required at the same scale?5M msg/sec peak with viral channels (50K msg/sec) requires dedicated fan-out pipeline separate from video.
Are streams archived for VOD?800 TB/day archive storage with dual-write during transcode vs post-stream processing.

Scope

In scope

  • RTMP/SRT video ingest
  • Real-time transcoding ladder
  • LL-HLS packaging & CDN distribution
  • Chat infrastructure with fan-out
  • VOD archival & clip creation

Out of scope (state explicitly)

Functional Requirements

Start by confirming scope with your interviewer: live video ingest and playback are the core, but chat scale, VOD archival, and DVR rewind each change the architecture. Ask whether LL-HLS (<2s latency) is required or standard HLS (~5s) is acceptable, because that single answer shapes segment duration and CDN caching.

  • Live video ingest: Ingest RTMP/SRT streams from OBS Studio or mobile apps
  • Adaptive bitrate transcoding: Transcode live stream to 1080p, 720p, 480p, 360p + audio-only
  • Low-latency playback: Deliver via LL-HLS/HLS with 2-5 second glass-to-glass latency
  • Chat system: Real-time chat per channel (up to 50K msgs/sec for top channels)
  • VOD recording: Automatically save completed streams for on-demand playback
  • Clips generation: Allow viewers to clip the last 30-60 seconds of a live stream
  • Stream discovery: Browse by category, search, live viewer count ranking
  • DVR capability: Allow viewers to pause and rewind up to 2 hours of live content

Non-Functional Requirements

Your interviewer will probe whether you understand that egress bandwidth, rather than ingest, is the primary scaling challenge. Call out the 200 Tbps viewer-side number early, as it justifies CDN-first delivery and why chat must run on a completely separate real-time path from video segments.

  • Ultra-Low Latency: 2-5 seconds glass-to-glass (streamer camera to viewer screen)
  • High Availability: 99.99% for video playback; stream MUST NOT drop during failover
  • Massive Scalability: 100K concurrent streams, 50M concurrent viewers
  • Cost Efficient: Video bandwidth is the primary cost center; optimize CDN usage
  • Smooth Playback: Zero buffering; ABR smoothly adapts to viewer's network
  • Chat Consistency: Chat messages in chronological order; rate-limited per user

Capacity Estimations

Run this math aloud before drawing boxes, because the 200 Tbps egress line item is what makes interviewers lean in. Compare it to 600 Gbps ingest: origin can handle ingest regionally, but no single datacenter serves 50M concurrent viewers. That gap is why every byte must terminate at a CDN edge POP.

MetricCalculationValue
Concurrent streamersGiven100,000
Concurrent viewersGiven50,000,000
Avg viewer stream bitrateGiven4 Mbps
Total egress bandwidth50M viewers x 4 Mbps200 Tbps
Total ingest bandwidth100K streamers x 6 Mbps600 Gbps
Chat messages / sec (peak)Given5,000,000/sec
VOD storage / day100K streams x 4 hrs x 2 GB/hr~800 TB/day
Target latencyGiven2-5 seconds (LL-HLS)

Architecture Diagram

In the room: ask HLS vs WebRTC early, because segment latency (2s to 10s) vs sub-second interaction changes your entire ingest path.

Walk your interviewer through two parallel pipelines: video flows from RTMP ingest to transcode to HLS segments to CDN, while chat flows WebSocket to Kafka to fan-out per channel. Never let chat backpressure block segment generation, because they share a viewer experience but not infrastructure.

Loading...

Component Deep Dives

Next we walk through each box on the diagram. Work left to right on the video path first, where ingest and transcode accumulate latency, then cover chat fan-out, which is a separate scaling problem disguised as a feature.

Ingest: Receiving the Live Stream

Every streamer connects with a stream key. Treat ingest as your authentication and quality gate before bytes hit expensive transcode GPUs.

Streamer uses OBS Studio to broadcast via RTMP (Real-Time Messaging Protocol) to the ingest server. The stream key is unique per channel and acts as authentication. Typical settings: 1080p 60fps, 6 Mbps bitrate, x264 encoder, keyframe interval 2s.

The RTMP Ingest Server receives the RTMP stream, decodes it, and validates the stream key (Redis lookup), bitrate limits, and correct keyframe interval. If valid, it forwards raw frames to the transcoding cluster.

RTMP vs SRT vs WebRTC for Ingest

ProtocolProsCons
RTMPUniversal, mature, every streaming software supports itTCP-based (higher latency), no built-in encryption, technically deprecated
SRT ⭐UDP-based (lower latency), built-in AES encryption, forward error correctionLess universal than RTMP (growing adoption)
WebRTCNo software needed (browser), ultra-low latency (< 500ms)Complex at scale (STUN/TURN), quality/bitrate constraints

Twitch uses RTMP for ingest (compatibility) + internal SRT transport. Future: migrating to SRT or QUIC-based protocols.

Real-Time Transcoding: The Hardest Part

Unlike VOD transcoding, live transcoding must be real-time. Each second of video must be transcoded in < 1 second (otherwise latency accumulates).

Pipeline per stream

  1. Receive raw frames from ingest (1080p 60fps = 60 frames/sec)
  2. Transcode to multiple qualities IN PARALLEL:
QualityResolutionFPSBitrateEncoder
Source1080p606 Mbpspassthrough
High720p603 Mbpsx264/NVENC
Medium480p301.5 Mbpsx264/NVENC
Low360p300.8 Mbpsx264
Audio onlyN/AN/A128 KbpsAAC

Real-time constraint: 1 second of 1080p 60fps must encode 60 frames across 4 qualities. x264 (CPU) takes ~5ms per frame → 60 x 4 = 1.2s (too slow sequentially). Solution: Parallel encoding: each quality on its own CPU core/GPU stream. GPU (NVENC): ~1ms per frame → 240ms total. At 100K concurrent streams: GPU approach needs ~17K GPUs vs 50K servers for CPU. GPU is 3x more cost-effective for live encoding.

HLS Segment Generation

Every 2 seconds, the encoded stream is cut into a segment (segment_000001.ts). The live playlist is updated with a sliding window of the last 3-5 segments. Clients poll the playlist every 1-2 seconds, discover new segments, and download them.

Low-Latency HLS (LL-HLS): Getting Below 3 Seconds

Standard HLS latency breakdown: Encoder buffer (2s) + Segment duration (6s) + Player buffer (18s) + CDN propagation (1s) = ~25-30 seconds.

How to reduce:

  1. Shorter segments (2s instead of 6): Reduces segment wait from 6s → 2s, but more HTTP requests and less CDN cache efficiency.
  2. LL-HLS with Partial Segments ⭐: Push 200ms "partial" segments instead of waiting for full 2-second segments. Client downloads partials as produced. Latency: encoder buffer (2s) + partial (0.2s) + CDN (0.5s) + player (0.5s) = ~3 seconds.
  3. HTTP/2 Server Push: Server pushes new partials without client polling.
  4. Preload hints: Client pre-connects and waits for the next partial before it's ready.
MethodLatency
Standard HLS25-30 seconds
Short segments6-10 seconds
LL-HLS ⭐2-5 seconds
WebRTC< 1 second (but doesn't scale)

Chat System: 50K Messages/Second Per Channel

Architecture: Client → WebSocket → Chat Gateway → Kafka → Chat Processor (Flink) → Fan-out.

Chat Gateway (WebSocket servers): Each server handles 100K WebSocket connections. 50M viewers → 500 WS servers. Viewer sends message → WS server validates → publishes to Kafka topic chat-messages keyed by channel_id.

Chat Processor (Flink) handles: rate limiting (max 1 msg per 1.5s per user), spam filter, banned words, subscriber checks, and emote parsing.

Fan-out: The Scaling Challenge: 200K viewers x 50K msgs/sec = 10B pushes/sec (impossible). Solutions:

  • Message batching: Batch messages per 100ms window → 1 batch per 100ms.
  • Message sampling: Show only 20 msgs/sec to each viewer (randomly sampled); highlighted messages always shown. Result: 200K x 20 = 4M pushes/sec (manageable).
  • Redis Pub/Sub: Each WS server subscribes to Redis Pub/Sub for channels its clients watch. 1 Redis publish → 500 server deliveries → each delivers to ~400 local clients.
  • Sharding hot channels: Top 10 channels get dedicated chat infrastructure.

VOD: Automatic Stream Archive

During live stream: HLS segments are copied to S3 "vod-archive" bucket. After stream ends: generate complete HLS manifest, DASH manifest, thumbnail, optionally split into chapters, run content moderation, and make available for on-demand playback.

Storage: 800 TB/day. Retention: 60 days (Twitch standard). Cost optimization: first 7 days on S3 Standard, 7-60 days on S3 Infrequent Access (50% cheaper).

Clips: Viewer clicks "Clip" → capture last 60 seconds from HLS segments → concatenate into MP4 → transcode to clip format → store permanently.

Event Bus Design (Kafka)

Topic: live-stream-events
  Partitions: 64 (partition by stream_id)
  Key: stream_id
  Retention: 3 days

Producer: Ingest server → publishes stream-started, stream-stopped, quality-dropped
Consumers:
  1. chat-fanout: Flink 100ms tumbling window batch → push to WS subscribers
  2. moderation: auto-flag banned words, spam velocity → moderation-events
  3. vod-archive: persist chat alongside VOD recording in S3

Topic: chat-messages (partition by channel_id, 256 partitions)
Topic: clip-requests: async highlight clip generation
Topic: stream-metrics (ingest bitrate, dropped frames, viewer count)

API Design

Start Stream

HTTP
POST /api/v1/streams/start
{
  "channel_id": "ch-uuid",
  "title": "Friday Night Gaming",
  "category": "Fortnite",
  "tags": ["English", "Competitive"],
  "language": "en"
}
Response: 200 OK
{
  "stream_id": "stream-uuid",
  "ingest_url": "rtmp://ingest-us-east.example.com/live",
  "stream_key": "live_sk_abc123def456",
  "recommended_settings": {
    "resolution": "1920x1080",
    "fps": 60,
    "bitrate": "6000 kbps",
    "keyframe_interval": 2,
    "encoder": "x264",
    "rate_control": "CBR"
  }
}

Get Live Stream (Viewer)

HTTP
GET /api/v1/streams/{channel_id}/live
Response: 200 OK
{
  "stream_id": "stream-uuid",
  "channel": {"name": "Ninja", "avatar": "..."},
  "title": "Friday Night Gaming",
  "viewer_count": 145230,
  "started_at": "2025-03-14T20:00:00Z",
  "manifest_url": "https://cdn.example.com/live/stream-uuid/master.m3u8",
  "chat_websocket": "wss://chat.example.com/ws/ch-uuid"
}

Send Chat Message

JSON
// WebSocket message
{
  "type": "chat_message",
  "channel_id": "ch-uuid",
  "content": "PogChamp that was insane!"
}

// Server broadcast
{
  "type": "chat_message",
  "id": "msg-uuid",
  "user": {"name": "viewer123", "color": "#FF0000"},
  "content": "PogChamp that was insane!",
  "timestamp": "2025-03-14T20:15:23Z"
}

Create Clip

HTTP
POST /api/v1/clips
{
  "stream_id": "stream-uuid",
  "title": "Insane play!",
  "duration_seconds": 30,
  "offset_seconds": -30
}
Response: 201 Created
{
  "clip_id": "clip-uuid",
  "url": "https://cdn.example.com/clips/clip-uuid.mp4",
  "status": "processing"
}

Browse Streams

HTTP
GET /api/v1/streams?category=gaming&sort=viewers_desc&limit=20

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
440 Login Timeout: WebSocket connection session expired, client reconnect is required
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

MySQL: Stream & Channel Metadata

SQL
CREATE TABLE channels (
    channel_id      UUID PRIMARY KEY,
    user_id         UUID NOT NULL UNIQUE,
    name            VARCHAR(50) UNIQUE NOT NULL,
    display_name    VARCHAR(50),
    description     TEXT,
    avatar_url      TEXT,
    banner_url      TEXT,
    stream_key      VARCHAR(64) NOT NULL,
    follower_count  INT DEFAULT 0,
    subscriber_count INT DEFAULT 0,
    partner         BOOLEAN DEFAULT FALSE,
    created_at      TIMESTAMP
);

CREATE TABLE streams (
    stream_id       UUID PRIMARY KEY,
    channel_id      UUID NOT NULL,
    title           VARCHAR(255),
    category_id     INT,
    language        CHAR(2),
    status          ENUM('live', 'ended') DEFAULT 'live',
    viewer_count    INT DEFAULT 0,
    peak_viewers    INT DEFAULT 0,
    started_at      TIMESTAMP,
    ended_at        TIMESTAMP,
    vod_url         TEXT,
    duration_seconds INT,
    INDEX idx_channel (channel_id, started_at DESC),
    INDEX idx_category_live (category_id, status, viewer_count DESC),
    INDEX idx_live_viewers (status, viewer_count DESC)
);

Redis: Live State

stream:{channel_id}    → Hash { stream_id, title, category, started_at, ingest_server, manifest_url, status }
viewers:{stream_id}    → INT (INCR/DECR on connect/disconnect)
stream_key:{key}       → channel_id
chat_rate:{user_id}:{channel_id}  → INT (INCR, check < 1 per 1.5 sec)
chat_banned:{channel_id}  → SET of user_ids
live_streams:{category}  → Sorted Set { channel_id: viewer_count }

Kafka Topics

Topic: stream-events         (stream started, ended, title changed)
Topic: chat-messages          (partitioned by channel_id for ordering)
Topic: viewer-events          (join, leave: for viewer count)
Topic: clip-requests          (async clip generation)
Topic: moderation-events      (ban, timeout, message delete)

S3: Video Storage

Bucket: live-segments (short retention, 48 hours)
  /{stream_id}/720p/segment_000001.ts
  /{stream_id}/1080p/segment_000001.ts
  /{stream_id}/master.m3u8

Bucket: vod-archive (60-day retention)
  /{stream_id}/vod/master.m3u8
  /{stream_id}/vod/720p/segment_000001.ts

Bucket: clips (permanent)
  /{clip_id}/clip.mp4

Fault Tolerance

ConcernSolution
Ingest server crash
  • Streamer's software auto-reconnects to backup ingest server
  • 2-5 sec gap
Transcoder crash
  • Hot standby transcoder per stream
  • failover in < 1 second
CDN edge failure
  • CDN auto-routes to next nearest PoP
  • viewer sees brief rebuffer
Origin storage failure
  • S3 11 nines durability
  • origin shield with multiple backends
Chat server crash
  • WS reconnect → new server
  • message history replayed from Kafka
Viewer count driftPeriodic reconciliation: count actual WS connections per stream
Stream key leaked
  • Streamer can regenerate stream key instantly
  • invalidate old key

Handling a Stream That Goes Viral (1K → 500K Viewers in Minutes)

CDN cache: Before viral, segments cached at 2-3 PoPs. After viral, viewers from 100+ PoPs. Solution: origin shield (absorbs 90% edge misses), CDN pre-push to all PoPs when viewer count > 50K, and multi-origin replication.

Transcoding: Migrate to dedicated transcoder (seamless, < 1s gap). Add more quality options for ABR.

Chat: Migrate channel to dedicated chat cluster, enable message sampling, add slow mode (1 msg per 5s per user).

Stream Latency Optimization Budget

End-to-end: ~2.3 seconds total. Each step optimized: encoder buffer (500ms), ingest PoP (200ms), GPU transcode (200ms), LL-HLS partial segment (200ms), CDN edge (500ms), player buffer (500ms), decode+render (100ms). Trade-off: Lower latency = smaller buffer = more rebuffering. Let viewer choose.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Bandwidth dominates cost: size CDN egress before compute (5 min)
    • Live path: RTMP ingest to GPU transcode to ABR HLS ladder (6 min)
    • Transcoder hot standby per stream for GPU failover (5 min)
    • Viewer count via periodic WebSocket sweeps, not per-segment (5 min)
    • Separate chat onto Kafka-backed cluster with rate limits (4 min)
  • Lead with the cost reality: bandwidth dominates, since 50M viewers at 4 Mbps is hundreds of Tbps egress, making CDN architecture the centerpiece.
  • Walk the live path: RTMP ingest to GPU transcode (ABR ladder) to LL-HLS segment packaging to CDN edge with origin shield.
  • Explain transcoder failover: hot standby per stream so a GPU crash causes <1 s gap, not a dead broadcast.
  • Cover viewer count via periodic WebSocket server sweeps rather than a central Redis counter that breaks on connection drops.
  • Separate chat onto its own Kafka-backed cluster with slow mode and message sampling when a stream goes viral.
  • Mention VOD capture: same ingest feeds a parallel recording pipeline for replay/clips after the stream ends.
  • Common pitfall: trying WebRTC for mass delivery, because it has no CDN support and cannot scale to millions of concurrent viewers.
  • Related systems: for VOD file chunking and multi-resolution encoding pipelines, see Video Transcoding Pipeline, and for discovery and home feed ranking, see Video Recommendation Engine.

Engineering Trade-offs

Live delivery trades latency against scale; these are the protocol and CDN choices your interviewer will ask you to justify.

HLS vs DASH vs WebRTC for Live Delivery

ProtocolProsCons
HLSUniversal (iOS, Android, all browsers), works with CDN, LL-HLS brings latency to 2-5sHigher latency than WebRTC
DASHOpen standard (no Apple dependency), LL-DASH similar to LL-HLSiOS Safari requires CMAF
WebRTCUltra-low latency (< 500ms)P2P doesn't scale to millions, no CDN support, no DRM

Industry moving to CMAF (unified format for HLS and DASH). Twitch uses a custom CMAF + chunked transfer encoding protocol achieving ~2-3 seconds end-to-end.

Viewer Count: Why It's Harder Than You Think

Approach 1: Central counter (Redis INCR/DECR): 50M viewers connecting/disconnecting causes millions of Redis ops/sec. WebSocket crashes inflate count (no DECR). Not recommended.

Approach 2: Periodic sweep ⭐ (Twitch's approach): Each WS server reports per-stream count every 10 seconds. Aggregator sums reports. Self-correcting (if server crashes, it stops reporting). 10-second staleness is acceptable ("145K viewers" doesn't need to be exact).

Approach 3: HyperLogLog: For unique viewer count (not concurrent). 12 KB per stream regardless of viewer count, ±0.81% error.

Cost Analysis: Bandwidth Is Everything

50M concurrent viewers x 4 Mbps avg = 200 Tbps egress. 200 Tbps ÷ 8 = 25 TB/s x 3600 = 90M GB/hour. At $0.03/GB CDN cost: ~$2.7M per hour. This is why Twitch operates at a loss on bandwidth alone.

Cost optimization: ABR pushes lower qualities (default to 720p), multi-CDN with volume discounts, own CDN infrastructure (Twitch/Amazon saves 50-70%), P2P CDN (experimental), and codec efficiency (HEVC/AV1 saves 30-50% bandwidth).

Transcoding: Dedicated Per-Stream vs Shared Pool

Dedicated: Isolated, predictable, but expensive and under-utilized.

Shared pool: Cost-efficient (6-8 streams per GPU), but noisy neighbor problem.

Hybrid ⭐ (Twitch): Tier 1 partners (> 1000 avg viewers) get dedicated. Tier 2 affiliates get shared pool (4 streams/instance). Tier 3 (< 10 viewers): source quality only (no transcoding).

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...