Interview Setup
Interview Prompt
Design a video transcoding pipeline that accepts uploads in any format and produces multi-resolution adaptive streaming packages (HLS/DASH) with thumbnails, subtitles, and optional DRM.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What's the upload volume and turnaround SLA? | 10K videos/hour with 30-min ready time drives segment parallelism vs one serial FFmpeg job. |
| How many codec/resolution variants per video? | 12 variants (6 resolutions x 2 codecs) with 3x expansion = 720 TB transcoded output per day. |
| Must failed steps retry independently? | One bad 1080p segment should not restart transcode for all 12 variants. |
| GPU or CPU for which codecs? | H.265 and AV1 run on GPU (~0.5 GPU-h per video) whereas H.264 runs on CPU (~2 CPU-h per video), requiring distinct worker pools and cost management. |
Scope
In scope
- Segment-based parallel transcoding
- DAG job orchestration
- HLS/DASH packaging
- GPU/CPU worker pools
- Priority queues
- Per-step retry and DLQ
Out of scope (state explicitly)
- Live streaming ingest (see Live Streaming Platform)
- Video recommendations (see Video Recommendation Engine)
- DRM license server
Functional Requirements
Clarify input formats, output renditions (ABR ladder), and SLA. You'll want upload, probe, transcode to multiple bitrates, and package for HLS/DASH, and confirm live vs VOD scope.
In the room: spot instances save 60% to 80% on compute costs but require idempotent segment jobs. Highlight that tradeoff when discussing operational expense.
- Ingest videos: Accept uploaded video files in any format (MP4, AVI, MOV, MKV, WebM)
- Multi-resolution transcoding: Convert to multiple resolutions (240p, 360p, 480p, 720p, 1080p, 4K)
- Multi-codec support: Encode in H.264, H.265/HEVC, VP9, AV1
- Adaptive bitrate packaging: Package into HLS (.m3u8 + .ts) and DASH (.mpd + .m4s)
- Audio processing: Extract, normalize, and transcode audio (AAC, Opus) at multiple bitrates
- Subtitle extraction: Auto-generate subtitles via speech-to-text; support uploaded subtitle files
- Thumbnail generation: Extract keyframes, generate sprite sheets for seek preview
- DRM encryption: Encrypt segments with Widevine/FairPlay/PlayReady
- Watermarking: Forensic or visible watermarking for content protection
- Progress tracking: Real-time progress reporting for upload → transcode → ready
- Priority queues: Premium content (paid creators) gets transcoded faster
Non-Functional Requirements
Transcode jobs complete within minutes for VOD; pipeline must survive worker preemption. GPU cost dominates.
- Throughput: Process 10,000+ videos/hour
- Latency: Standard video ready within 30 minutes; short videos (< 5 min) within 5 minutes
- Durability: Original uploaded file NEVER lost; transcoded outputs can be regenerated
- Fault Tolerance: Any step failure → retry that step, not the entire pipeline
- Cost Efficient: GPU for H.265/AV1; CPU for H.264; spot instances for non-urgent work
- Scalability: Auto-scale based on queue depth; handle viral upload spikes
- Quality: Output quality comparable to or better than input
- Idempotent: Re-running a failed job produces the same output
Capacity Estimations
Upload volume, average video length, and rendition count drive worker pool and storage sizing.
| Metric | Calculation | Value |
|---|---|---|
| Videos uploaded / hour | Given (assumption documented in value) | 10,000 |
| Avg original video duration | Given (typical workload assumption) | 10 minutes |
| Avg original file size | Given (typical workload assumption) | 1 GB |
| Upload storage / day | Derived from upstream throughput x size | 240 TB |
| Transcoded variants per video | 6 resolutions x 2 codecs | 12 |
| Expansion factor (transcoded/original) | Given | ~3x |
| Transcoded storage / day | Derived from upstream throughput x size | 720 TB |
| CPU-hours per video (H.264) | Given | ~2 CPU-hours |
| GPU-hours per video (H.265/AV1) | Given | ~0.5 GPU-hours |
| Total compute / day | 240K videos x (2 CPU-h + 0.5 GPU-h) | ~480K CPU-hours + ~120K GPU-hours |
Architecture Diagram
We queue transcode jobs on upload, fan out parallel segment encodes on workers, then package manifests and push to CDN origin.
Component Deep Dives
Walk through the upload, probe, parallel transcode, packaging, and publishing pipeline, addressing spot instance preemption and segment retries.
Segment-Based Parallel Transcoding: The Core Optimization
Parallelism comes from segmenting before encoding, allowing the orchestrator to fan out hundreds of small FFmpeg tasks instead of a handful of long ones.
Why split into segments? Without splitting: 1 video x 6 resolutions = 6 tasks at 20 minutes each = 20 min wall-clock. With splitting (10-second segments): 60 segments x 6 resolutions = 360 tasks at ~3 seconds each = ~30 seconds wall-clock with 40 workers.
GOP-aligned splitting: Must split at GOP boundaries (I-frame positions). FFmpeg: ffmpeg -i input.mp4 -c copy -f segment -segment_time 10 -reset_timestamps 1 segment_%03d.mp4. The -c copy flag ensures no re-encoding during split (fast, lossless).
Reassembly: FFmpeg concat demuxer. For HLS, segments ARE the final output: no reassembly needed! Just generate the .m3u8 playlist.
FFmpeg Command Breakdown
H.264 (CPU)
ffmpeg -i segment_005.mp4 -vf "scale=1280:720" -c:v libx264 -preset medium -crf 23 -profile:v high -level 4.0 -maxrate 3M -bufsize 6M -g 48 -sc_threshold 0 -an segment_005_720p.tsH.265 (GPU NVENC)
ffmpeg -i segment_005.mp4 -vf "scale=1280:720" -c:v hevc_nvenc -preset p5 -rc:v vbr -cq 28 -maxrate 1.5M -bufsize 3M -g 48 -tag:v hvc1 segment_005_720p_h265.tsPipeline Orchestrator: Temporal/Step Functions
Kafka consumers alone can't express complex DAG dependencies. Temporal ⭐ provides DAG definition, per-step retries with exponential backoff, workflow state persistence, visibility dashboard, timeouts, and versioning. Each step has a retry policy with configurable initial interval, maximum interval, maximum attempts, and non-retryable errors.
@workflow
def transcode_pipeline(video_id, s3_key):
probe_result = await probe_video(s3_key)
segments = await split_video(s3_key, probe_result)
resolutions = determine_resolutions(probe_result)
transcode_futures = []
for res in resolutions:
for segment in segments:
future = transcode_segment.async(segment, res, 'h264')
transcode_futures.append(future)
transcoded = await all(transcode_futures)
manifest = await package_hls_dash(transcoded, video_id)
await all(
generate_thumbnails(s3_key, probe_result),
extract_audio(s3_key),
content_moderation(s3_key),
generate_subtitles(s3_key)
)
await encrypt_drm(manifest, video_id)
await update_video_status(video_id, 'ready')Event Bus Design (Kafka)
Topic: video-uploaded
Partitions: 128 (scale worker pools independently)
Partition key: video_id (preserves per-video job ordering)
Retention: 7 days (replay failed transcodes)
Replication factor: 3, min.insync.replicas: 2
Producer: S3 event notification → API enriches with { video_id, s3_key, codec_hint, priority }
Consumer groups:
1. orchestrator: start Temporal DAG (probe → GOP split → parallel transcode → package)
2. status-updater: UPDATE videos SET status=processing|ready|failed in MySQL
3. cdn-purger: invalidate CDN manifest on transcode-complete
Topic: transcode-progress (internal): per-segment completion for DAG resume
Topic: transcode-failed-dlq: poison jobs after 3 retries; alert ops
Sync path: presigned upload ACK < 200ms; processing is fully async
Async path: workers pull segment jobs from CPU/GPU pools; idempotent by segment_id
Lag alert: consumer lag > 300s → scale GPU poolAPI Design
Show upload initiation, job status, and playback URL endpoints.
Initiate Upload
POST /api/v1/videos/upload
{
"filename": "vacation.mp4",
"content_type": "video/mp4",
"file_size_bytes": 1073741824,
"title": "Summer Vacation 2025"
}
Response: 200 OK
{
"video_id": "vid-uuid",
"upload_url": "https://s3.amazonaws.com/originals/vid-uuid/upload?X-Amz-...",
"upload_id": "multipart-upload-id",
"max_chunk_size": 104857600
}Check Transcoding Status
GET /api/v1/videos/{video_id}/status
Response: 200 OK
{
"video_id": "vid-uuid",
"status": "transcoding",
"pipeline_progress": {
"probe": "completed",
"split": "completed",
"transcode": { "completed": 45, "total": 60, "percent": 75 },
"package": "pending",
"thumbnails": "completed",
"moderation": "pending",
"drm": "pending"
},
"estimated_completion": "2025-03-14T11:15:00Z",
"started_at": "2025-03-14T10:45:00Z"
}Retry Failed Pipeline
POST /api/v1/videos/{video_id}/retry
{ "from_step": "transcode" }
Response: 200 OK
{ "status": "retrying", "retry_from": "transcode" }Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
MySQL: Video and Job Metadata
CREATE TABLE videos (
video_id VARCHAR(36) PRIMARY KEY,
creator_id VARCHAR(36) NOT NULL,
title VARCHAR(255),
original_s3_key TEXT NOT NULL,
original_format VARCHAR(10),
duration_sec INT,
original_width INT,
original_height INT,
original_codec VARCHAR(20),
original_bitrate_kbps INT,
file_size_bytes BIGINT,
status ENUM('uploaded','probing','transcoding','packaging',
'moderating','ready','failed','removed') DEFAULT 'uploaded',
preset VARCHAR(20) DEFAULT 'standard',
priority ENUM('low','normal','high','urgent') DEFAULT 'normal',
manifest_url TEXT,
thumbnail_url TEXT,
error_message TEXT,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
completed_at TIMESTAMP,
INDEX idx_status (status),
INDEX idx_creator (creator_id, created_at DESC)
);
CREATE TABLE transcoding_tasks (
task_id VARCHAR(36) PRIMARY KEY,
video_id VARCHAR(36) NOT NULL,
task_type ENUM('probe','split','transcode','package','thumbnail',
'audio','moderation','drm','subtitle'),
resolution VARCHAR(10),
codec VARCHAR(10),
segment_index INT,
status ENUM('pending','running','completed','failed','retrying'),
worker_id VARCHAR(36),
s3_input_key TEXT,
s3_output_key TEXT,
started_at TIMESTAMP,
completed_at TIMESTAMP,
duration_ms INT,
error_message TEXT,
retry_count INT DEFAULT 0,
INDEX idx_video (video_id, task_type),
INDEX idx_status (status),
INDEX idx_worker (worker_id)
);S3: Storage Structure
Bucket: video-originals (NEVER deleted, cross-region replicated)
/{video_id}/original.mp4
Bucket: video-segments (temporary, auto-delete after 7 days)
/{video_id}/segments/segment_000.mp4
Bucket: video-transcoded (long-term, lifecycle policies)
/{video_id}/h264/720p/segment_000.ts
/{video_id}/h264/720p/playlist.m3u8
/{video_id}/h265/720p/segment_000.ts
/{video_id}/manifest.m3u8
/{video_id}/manifest.mpd
/{video_id}/thumbnails/thumb_001.jpg
/{video_id}/subtitles/en.vtt
/{video_id}/audio/aac_128k.m4aRedis: Pipeline State + Queue Management
pipeline:{video_id} → Hash { status, current_step, transcode_completed, transcode_total, started_at, estimated_completion }
task_queue:transcode:cpu → Sorted Set { task_id: priority_score }
task_queue:transcode:gpu → Sorted Set { task_id: priority_score }
worker:{worker_id} → Hash { status, current_task, last_heartbeat }Fault Tolerance
| Concern | Solution |
|---|---|
| Original file lost |
|
| Transcode worker crash | Temporal detects heartbeat timeout → re-schedule on another worker |
| Segment transcode failure |
|
| S3 upload failure |
|
| Worker pool exhaustion |
|
| Corrupt input video | Probe step detects invalid file → fail fast, notify creator |
| Spot instance preemption |
|
Handle Spot Instance Preemption
GPU instances are expensive. Spot instances save 60-80%. Strategy: use spot for transcoding workers (segment-based, each takes 3-10s, usually finishes before preemption). If preempted: worker marks task as "interrupted", Temporal re-schedules on another worker. Critical path (probe, package, publish) uses on-demand instances.
Quality Verification After Transcoding
Automated checks: duration match (±0.5s), frame count match, VMAF score (> 80 for 720p, > 85 for 1080p), audio-video sync (< 50ms drift), black frame/freeze detection, and bitrate compliance (±20% of target). These checks add ~5 seconds per video but catch 0.1% of errors.
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Start with upload to S3 and Kafka job queue, returning job ID immediately without blocking on transcode (5 min)
- Pipeline stages: probe metadata, GOP segment split, parallel FFmpeg workers, and manifest assembly (6 min)
- Explain why 10-second segments cut wall-clock from ~20 min to ~30 sec via parallelism (5 min)
- Priority queues in Redis sorted sets: premium creators and live replays jump the line (5 min)
- Staff only: spot-instance idempotency, Temporal-style workflow compensation, and CDN propagation (4 min)
- Frame upload as async: accept the file to S3, return a job ID immediately, because transcoding takes minutes and must never block the API.
- Walk through the pipeline stages: probe, segment split, parallel transcode (CPU/GPU queues), package HLS/DASH, and publish to CDN origin.
- Explain why Temporal (or similar) orchestrates the workflow, because heartbeat timeouts re-schedule failed segments on another worker automatically.
- Cover priority queues in Redis sorted sets: premium creators and trending videos jump ahead of long-tail backlog.
- Mention a tiered codec strategy: encode H.264 for all uploads immediately, adding H.265 or AV1 only when view counts justify the higher GPU cost.
- Discuss spot instances for segment workers with checkpointing, keeping probe and package steps on on-demand instances.
- Common pitfall: monolithic FFmpeg on a single worker for a 2-hour 4K video, where a single crash loses all progress instead of retrying individual segments.
- Related systems: for continuous real-time broadcast instead of VOD chunking, see Live Streaming Platform, and for discovery pipelines, see Video Recommendation Engine.
Engineering Trade-offs
Compare segment-based against whole-file transcoding, spot against on-demand worker pools, and codec tradeoffs between H.264 and AV1.
CRF vs CBR vs VBR: Bitrate Control Strategies
| Strategy | Description | Best For |
|---|---|---|
| CBR (Constant Bitrate) | Every second uses same bitrate. Predictable but wastes bits on simple scenes. | Live streaming (consistent bandwidth) |
| VBR (Variable Bitrate) | Bitrate varies by scene complexity. Better perceptual quality. | VOD (pre-recorded) |
| CRF ⭐ (Constant Rate Factor) | Target constant QUALITY, let bitrate vary. Best quality for given size. | VOD transcoding (YouTube, Netflix). Use CRF + maxrate for best of both worlds. |
Netflix's per-title encoding: Encode test segment at multiple CRFs, measure VMAF, and pick optimal CRF per video. For instance, animated content can use CRF 28 whereas an action movie requires CRF 20, yielding 20% to 40% bitrate savings over fixed CRF.
Per-Title vs Per-Shot Encoding
Per-Title: For each video, run convex hull analysis across CRF values and resolutions. Select optimal CRF per resolution maximizing VMAF/bitrate ratio. Up to 40% bandwidth savings.
Per-Shot (state of the art): Split video into shots (scene changes). Each shot gets its own encoding parameters. Dialogue (static): CRF 28, car chase (motion): CRF 20, credits: CRF 30. Additional 10-20% savings over per-title.
Cost Optimization: When to Encode Which Codec
Tier 1 (all videos): H.264 at 480p, 720p, 1080p. Cost: ~$0.02/video.
Tier 2 (> 100 views in first hour): Add H.265. Cost: ~$0.08/video (GPU).
Tier 3 (> 10K views): Add AV1. Cost: ~$0.50/video but saves 30% more bandwidth than H.265. At 10K views, ROI: 330x. Uploads trigger immediate H.264 encoding. An hourly job inspects view counts to schedule H.265 encoding, while a daily job schedules AV1 encoding for highly popular videos.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.