System Design Problem

Design a Video Transcoding Pipeline

Commonly Asked By:NetflixGoogleTwitchAWS

Interview Setup

Interview Prompt

Design a video transcoding pipeline that accepts uploads in any format and produces multi-resolution adaptive streaming packages (HLS/DASH) with thumbnails, subtitles, and optional DRM.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What's the upload volume and turnaround SLA?10K videos/hour with 30-min ready time drives segment parallelism vs one serial FFmpeg job.
How many codec/resolution variants per video?12 variants (6 resolutions x 2 codecs) with 3x expansion = 720 TB transcoded output per day.
Must failed steps retry independently?One bad 1080p segment should not restart transcode for all 12 variants.
GPU or CPU for which codecs?H.265 and AV1 run on GPU (~0.5 GPU-h per video) whereas H.264 runs on CPU (~2 CPU-h per video), requiring distinct worker pools and cost management.

Scope

In scope

  • Segment-based parallel transcoding
  • DAG job orchestration
  • HLS/DASH packaging
  • GPU/CPU worker pools
  • Priority queues
  • Per-step retry and DLQ

Out of scope (state explicitly)

Functional Requirements

Clarify input formats, output renditions (ABR ladder), and SLA. You'll want upload, probe, transcode to multiple bitrates, and package for HLS/DASH, and confirm live vs VOD scope.

In the room: spot instances save 60% to 80% on compute costs but require idempotent segment jobs. Highlight that tradeoff when discussing operational expense.

  • Ingest videos: Accept uploaded video files in any format (MP4, AVI, MOV, MKV, WebM)
  • Multi-resolution transcoding: Convert to multiple resolutions (240p, 360p, 480p, 720p, 1080p, 4K)
  • Multi-codec support: Encode in H.264, H.265/HEVC, VP9, AV1
  • Adaptive bitrate packaging: Package into HLS (.m3u8 + .ts) and DASH (.mpd + .m4s)
  • Audio processing: Extract, normalize, and transcode audio (AAC, Opus) at multiple bitrates
  • Subtitle extraction: Auto-generate subtitles via speech-to-text; support uploaded subtitle files
  • Thumbnail generation: Extract keyframes, generate sprite sheets for seek preview
  • DRM encryption: Encrypt segments with Widevine/FairPlay/PlayReady
  • Watermarking: Forensic or visible watermarking for content protection
  • Progress tracking: Real-time progress reporting for upload → transcode → ready
  • Priority queues: Premium content (paid creators) gets transcoded faster

Non-Functional Requirements

Transcode jobs complete within minutes for VOD; pipeline must survive worker preemption. GPU cost dominates.

  • Throughput: Process 10,000+ videos/hour
  • Latency: Standard video ready within 30 minutes; short videos (< 5 min) within 5 minutes
  • Durability: Original uploaded file NEVER lost; transcoded outputs can be regenerated
  • Fault Tolerance: Any step failure → retry that step, not the entire pipeline
  • Cost Efficient: GPU for H.265/AV1; CPU for H.264; spot instances for non-urgent work
  • Scalability: Auto-scale based on queue depth; handle viral upload spikes
  • Quality: Output quality comparable to or better than input
  • Idempotent: Re-running a failed job produces the same output

Capacity Estimations

Upload volume, average video length, and rendition count drive worker pool and storage sizing.

MetricCalculationValue
Videos uploaded / hourGiven (assumption documented in value)10,000
Avg original video durationGiven (typical workload assumption)10 minutes
Avg original file sizeGiven (typical workload assumption)1 GB
Upload storage / dayDerived from upstream throughput x size240 TB
Transcoded variants per video6 resolutions x 2 codecs12
Expansion factor (transcoded/original)Given~3x
Transcoded storage / dayDerived from upstream throughput x size720 TB
CPU-hours per video (H.264)Given~2 CPU-hours
GPU-hours per video (H.265/AV1)Given~0.5 GPU-hours
Total compute / day240K videos x (2 CPU-h + 0.5 GPU-h)~480K CPU-hours + ~120K GPU-hours

Architecture Diagram

We queue transcode jobs on upload, fan out parallel segment encodes on workers, then package manifests and push to CDN origin.

Loading...

Component Deep Dives

Walk through the upload, probe, parallel transcode, packaging, and publishing pipeline, addressing spot instance preemption and segment retries.

Segment-Based Parallel Transcoding: The Core Optimization

Parallelism comes from segmenting before encoding, allowing the orchestrator to fan out hundreds of small FFmpeg tasks instead of a handful of long ones.

Why split into segments? Without splitting: 1 video x 6 resolutions = 6 tasks at 20 minutes each = 20 min wall-clock. With splitting (10-second segments): 60 segments x 6 resolutions = 360 tasks at ~3 seconds each = ~30 seconds wall-clock with 40 workers.

GOP-aligned splitting: Must split at GOP boundaries (I-frame positions). FFmpeg: ffmpeg -i input.mp4 -c copy -f segment -segment_time 10 -reset_timestamps 1 segment_%03d.mp4. The -c copy flag ensures no re-encoding during split (fast, lossless).

Reassembly: FFmpeg concat demuxer. For HLS, segments ARE the final output: no reassembly needed! Just generate the .m3u8 playlist.

FFmpeg Command Breakdown

H.264 (CPU)

BASH
ffmpeg -i segment_005.mp4 -vf "scale=1280:720" -c:v libx264 -preset medium -crf 23 -profile:v high -level 4.0 -maxrate 3M -bufsize 6M -g 48 -sc_threshold 0 -an segment_005_720p.ts

H.265 (GPU NVENC)

BASH
ffmpeg -i segment_005.mp4 -vf "scale=1280:720" -c:v hevc_nvenc -preset p5 -rc:v vbr -cq 28 -maxrate 1.5M -bufsize 3M -g 48 -tag:v hvc1 segment_005_720p_h265.ts

Pipeline Orchestrator: Temporal/Step Functions

Kafka consumers alone can't express complex DAG dependencies. Temporal ⭐ provides DAG definition, per-step retries with exponential backoff, workflow state persistence, visibility dashboard, timeouts, and versioning. Each step has a retry policy with configurable initial interval, maximum interval, maximum attempts, and non-retryable errors.

@workflow
def transcode_pipeline(video_id, s3_key):
    probe_result = await probe_video(s3_key)
    segments = await split_video(s3_key, probe_result)
    resolutions = determine_resolutions(probe_result)
    transcode_futures = []
    for res in resolutions:
        for segment in segments:
            future = transcode_segment.async(segment, res, 'h264')
            transcode_futures.append(future)
    transcoded = await all(transcode_futures)
    manifest = await package_hls_dash(transcoded, video_id)
    await all(
        generate_thumbnails(s3_key, probe_result),
        extract_audio(s3_key),
        content_moderation(s3_key),
        generate_subtitles(s3_key)
    )
    await encrypt_drm(manifest, video_id)
    await update_video_status(video_id, 'ready')

Event Bus Design (Kafka)

Topic: video-uploaded
  Partitions: 128 (scale worker pools independently)
  Partition key: video_id (preserves per-video job ordering)
  Retention: 7 days (replay failed transcodes)
  Replication factor: 3, min.insync.replicas: 2

Producer: S3 event notification → API enriches with { video_id, s3_key, codec_hint, priority }
Consumer groups:
  1. orchestrator: start Temporal DAG (probe → GOP split → parallel transcode → package)
  2. status-updater: UPDATE videos SET status=processing|ready|failed in MySQL
  3. cdn-purger: invalidate CDN manifest on transcode-complete

Topic: transcode-progress (internal): per-segment completion for DAG resume
Topic: transcode-failed-dlq: poison jobs after 3 retries; alert ops

Sync path: presigned upload ACK < 200ms; processing is fully async
Async path: workers pull segment jobs from CPU/GPU pools; idempotent by segment_id
Lag alert: consumer lag > 300s → scale GPU pool

API Design

Show upload initiation, job status, and playback URL endpoints.

Initiate Upload

HTTP
POST /api/v1/videos/upload
{
  "filename": "vacation.mp4",
  "content_type": "video/mp4",
  "file_size_bytes": 1073741824,
  "title": "Summer Vacation 2025"
}
Response: 200 OK
{
  "video_id": "vid-uuid",
  "upload_url": "https://s3.amazonaws.com/originals/vid-uuid/upload?X-Amz-...",
  "upload_id": "multipart-upload-id",
  "max_chunk_size": 104857600
}

Check Transcoding Status

HTTP
GET /api/v1/videos/{video_id}/status
Response: 200 OK
{
  "video_id": "vid-uuid",
  "status": "transcoding",
  "pipeline_progress": {
    "probe": "completed",
    "split": "completed",
    "transcode": { "completed": 45, "total": 60, "percent": 75 },
    "package": "pending",
    "thumbnails": "completed",
    "moderation": "pending",
    "drm": "pending"
  },
  "estimated_completion": "2025-03-14T11:15:00Z",
  "started_at": "2025-03-14T10:45:00Z"
}

Retry Failed Pipeline

HTTP
POST /api/v1/videos/{video_id}/retry
{ "from_step": "transcode" }
Response: 200 OK
{ "status": "retrying", "retry_from": "transcode" }

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

MySQL: Video and Job Metadata

SQL
CREATE TABLE videos (
    video_id        VARCHAR(36) PRIMARY KEY,
    creator_id      VARCHAR(36) NOT NULL,
    title           VARCHAR(255),
    original_s3_key TEXT NOT NULL,
    original_format VARCHAR(10),
    duration_sec    INT,
    original_width  INT,
    original_height INT,
    original_codec  VARCHAR(20),
    original_bitrate_kbps INT,
    file_size_bytes BIGINT,
    status          ENUM('uploaded','probing','transcoding','packaging',
                         'moderating','ready','failed','removed') DEFAULT 'uploaded',
    preset          VARCHAR(20) DEFAULT 'standard',
    priority        ENUM('low','normal','high','urgent') DEFAULT 'normal',
    manifest_url    TEXT,
    thumbnail_url   TEXT,
    error_message   TEXT,
    created_at      TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    completed_at    TIMESTAMP,
    INDEX idx_status (status),
    INDEX idx_creator (creator_id, created_at DESC)
);

CREATE TABLE transcoding_tasks (
    task_id         VARCHAR(36) PRIMARY KEY,
    video_id        VARCHAR(36) NOT NULL,
    task_type       ENUM('probe','split','transcode','package','thumbnail',
                         'audio','moderation','drm','subtitle'),
    resolution      VARCHAR(10),
    codec           VARCHAR(10),
    segment_index   INT,
    status          ENUM('pending','running','completed','failed','retrying'),
    worker_id       VARCHAR(36),
    s3_input_key    TEXT,
    s3_output_key   TEXT,
    started_at      TIMESTAMP,
    completed_at    TIMESTAMP,
    duration_ms     INT,
    error_message   TEXT,
    retry_count     INT DEFAULT 0,
    INDEX idx_video (video_id, task_type),
    INDEX idx_status (status),
    INDEX idx_worker (worker_id)
);

S3: Storage Structure

Bucket: video-originals (NEVER deleted, cross-region replicated)
  /{video_id}/original.mp4

Bucket: video-segments (temporary, auto-delete after 7 days)
  /{video_id}/segments/segment_000.mp4

Bucket: video-transcoded (long-term, lifecycle policies)
  /{video_id}/h264/720p/segment_000.ts
  /{video_id}/h264/720p/playlist.m3u8
  /{video_id}/h265/720p/segment_000.ts
  /{video_id}/manifest.m3u8
  /{video_id}/manifest.mpd
  /{video_id}/thumbnails/thumb_001.jpg
  /{video_id}/subtitles/en.vtt
  /{video_id}/audio/aac_128k.m4a

Redis: Pipeline State + Queue Management

pipeline:{video_id}   → Hash { status, current_step, transcode_completed, transcode_total, started_at, estimated_completion }
task_queue:transcode:cpu  → Sorted Set { task_id: priority_score }
task_queue:transcode:gpu  → Sorted Set { task_id: priority_score }
worker:{worker_id}  → Hash { status, current_task, last_heartbeat }

Fault Tolerance

ConcernSolution
Original file lost
  • S3 cross-region replication
  • 11 nines durability
  • versioning enabled
Transcode worker crashTemporal detects heartbeat timeout → re-schedule on another worker
Segment transcode failure
  • Retry 3 times with backoff
  • after 3 failures → DLQ, alert ops
S3 upload failure
  • Retry with exponential backoff
  • S3 multipart for large segments
Worker pool exhaustion
  • Auto-scaling based on queue depth
  • alert if depth > 1000 for > 10 min
Corrupt input videoProbe step detects invalid file → fail fast, notify creator
Spot instance preemption
  • Task checkpointing
  • preempted task re-queued automatically

Handle Spot Instance Preemption

GPU instances are expensive. Spot instances save 60-80%. Strategy: use spot for transcoding workers (segment-based, each takes 3-10s, usually finishes before preemption). If preempted: worker marks task as "interrupted", Temporal re-schedules on another worker. Critical path (probe, package, publish) uses on-demand instances.

Quality Verification After Transcoding

Automated checks: duration match (±0.5s), frame count match, VMAF score (> 80 for 720p, > 85 for 1080p), audio-video sync (< 50ms drift), black frame/freeze detection, and bitrate compliance (±20% of target). These checks add ~5 seconds per video but catch 0.1% of errors.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Start with upload to S3 and Kafka job queue, returning job ID immediately without blocking on transcode (5 min)
    • Pipeline stages: probe metadata, GOP segment split, parallel FFmpeg workers, and manifest assembly (6 min)
    • Explain why 10-second segments cut wall-clock from ~20 min to ~30 sec via parallelism (5 min)
    • Priority queues in Redis sorted sets: premium creators and live replays jump the line (5 min)
    • Staff only: spot-instance idempotency, Temporal-style workflow compensation, and CDN propagation (4 min)
  • Frame upload as async: accept the file to S3, return a job ID immediately, because transcoding takes minutes and must never block the API.
  • Walk through the pipeline stages: probe, segment split, parallel transcode (CPU/GPU queues), package HLS/DASH, and publish to CDN origin.
  • Explain why Temporal (or similar) orchestrates the workflow, because heartbeat timeouts re-schedule failed segments on another worker automatically.
  • Cover priority queues in Redis sorted sets: premium creators and trending videos jump ahead of long-tail backlog.
  • Mention a tiered codec strategy: encode H.264 for all uploads immediately, adding H.265 or AV1 only when view counts justify the higher GPU cost.
  • Discuss spot instances for segment workers with checkpointing, keeping probe and package steps on on-demand instances.
  • Common pitfall: monolithic FFmpeg on a single worker for a 2-hour 4K video, where a single crash loses all progress instead of retrying individual segments.
  • Related systems: for continuous real-time broadcast instead of VOD chunking, see Live Streaming Platform, and for discovery pipelines, see Video Recommendation Engine.

Engineering Trade-offs

Compare segment-based against whole-file transcoding, spot against on-demand worker pools, and codec tradeoffs between H.264 and AV1.

CRF vs CBR vs VBR: Bitrate Control Strategies

StrategyDescriptionBest For
CBR (Constant Bitrate)Every second uses same bitrate. Predictable but wastes bits on simple scenes.Live streaming (consistent bandwidth)
VBR (Variable Bitrate)Bitrate varies by scene complexity. Better perceptual quality.VOD (pre-recorded)
CRF ⭐ (Constant Rate Factor)Target constant QUALITY, let bitrate vary. Best quality for given size.VOD transcoding (YouTube, Netflix). Use CRF + maxrate for best of both worlds.

Netflix's per-title encoding: Encode test segment at multiple CRFs, measure VMAF, and pick optimal CRF per video. For instance, animated content can use CRF 28 whereas an action movie requires CRF 20, yielding 20% to 40% bitrate savings over fixed CRF.

Per-Title vs Per-Shot Encoding

Per-Title: For each video, run convex hull analysis across CRF values and resolutions. Select optimal CRF per resolution maximizing VMAF/bitrate ratio. Up to 40% bandwidth savings.

Per-Shot (state of the art): Split video into shots (scene changes). Each shot gets its own encoding parameters. Dialogue (static): CRF 28, car chase (motion): CRF 20, credits: CRF 30. Additional 10-20% savings over per-title.

Cost Optimization: When to Encode Which Codec

Tier 1 (all videos): H.264 at 480p, 720p, 1080p. Cost: ~$0.02/video.

Tier 2 (> 100 views in first hour): Add H.265. Cost: ~$0.08/video (GPU).

Tier 3 (> 10K views): Add AV1. Cost: ~$0.50/video but saves 30% more bandwidth than H.265. At 10K views, ROI: 330x. Uploads trigger immediate H.264 encoding. An hourly job inspects view counts to schedule H.265 encoding, while a daily job schedules AV1 encoding for highly popular videos.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...