Interview Setup
Interview Prompt
Design a video recommendation engine for a YouTube-scale platform: home feed, up-next, similar videos, driven by user watch behavior and video embeddings.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Homepage load QPS vs catalog size? | Serving 200K recommendation requests per second across 500M videos requires two-stage retrieval rather than brute-force scoring. |
| Latency budget for ranking? | 50ms inference budget splits retrieve vs rank phases. |
| Cold start users and new videos? | Trending + content features until embeddings exist. |
| Real-time vs batch feature refresh? | Watch-events stream updates user embedding near-real-time. |
Scope
In scope
- Candidate retrieval (ANN)
- Ranking model
- User/video embeddings
- Pre-computed rec cache
- Cold start
- Feedback loop
Out of scope (state explicitly)
- GPU training cluster
- Content moderation
- Ad auction ranking
Functional Requirements
Start by asking your interviewer about the homepage latency budget and catalog size. Serving recommendations across massive catalogs requires a two-stage funnel comprising candidate retrieval followed by heavy ranking, rather than brute-force scoring. The engine works in close coordination with the Video Transcoding Pipeline for content ingestion and the Podcast Delivery Platform for cross-media syndication patterns.
- Personalized recommendations: "Videos you might like" feed tailored to each user's watch history, likes, and preferences
- "Up Next" recommendation: After watching a video, suggest what to watch next (autoplay)
- Homepage feed: Curated mix of trending, personalized, and fresh content
- Similar videos: "Because you watched X": find videos related to a specific video
- Category/topic recommendations: "Trending in Technology", "Popular in Music"
- New user cold start: Recommend popular/trending content for users with no history
- Explain recommendations: "Recommended because you watched System Design Interview"
- Feedback loop: Incorporate explicit (like/dislike) and implicit (watch time, skip) signals
- Diversity: Avoid filter bubble; include exploratory recommendations
- Real-time updates: Recommendation changes within minutes of new user behavior
Non-Functional Requirements
Watch time, rather than click-through rate alone, is the north-star metric interviewers expect you to prioritize. Introduce the exploration vs exploitation tradeoff early: pure personalization risks filter bubbles, whereas catalog diversity and new-creator fairness serve as product constraints that shape ranking features.
- Low Latency: Recommendations served in < 200 ms
- Scalability: 1B+ users, 500M+ videos, 5B+ watches/day
- Freshness: Newly uploaded videos appear in recommendations within 1 hour
- Quality: Increase average watch time per session (the north star metric)
- Diversity: No more than 30% of recommendations from same creator/category
- Fairness: New creators' content gets a fair chance (exploration vs exploitation)
- Availability: 99.99%: recommendations are the main UI; failure = blank homepage
Capacity Estimations
Scoring 500M videos per request is computationally prohibitive, which is why system capacity hinges on the candidate funnel. The calculations below demonstrate why approximate nearest neighbor (ANN) retrieval of roughly 500 candidates followed by a specialized ranking model is required instead of brute-force scoring.
| Metric | Calculation | Value |
|---|---|---|
| DAU | Given (product assumption) | 500M |
| Videos in catalog | Given (assumption documented in value) | 500M |
| Watches / day | Given (assumption documented in value) | 5B |
| Recommendation requests / sec | From Recommendation requests / day ÷ 86400 (+ peak factor in value) | 200K (homepage loads, up-next) |
| User feature vector size | 256 floats | 1 KB |
| Video feature vector size | 256 floats | 1 KB |
| User features total | 1B x 1 KB | 1 TB |
| Video features total | 500M x 1 KB | 500 GB |
| Model inference latency budget | Given (assumption documented in value) | < 50 ms |
| Pre-computed recs cache | 500M users x 100 recs x 8B | 400 GB |
Architecture Diagram
In the room: name the two-stage funnel before drawing. Explain that you cannot score 500M videos per homepage request, so you must use ANN retrieval to shortlist roughly 500 candidates before ranking.
In the interview, define the two-stage funnel before drawing: candidate generation using ANN retrieval (approximately 500 items in under 10ms), followed by ranking using a cross-network model (under 40ms). Supporting operations such as training, feature backfills, and A/B experiments remain asynchronous via Kafka.
A YouTube-scale recommendation engine cannot score 500M videos per homepage load. It must funnel the catalog through candidate retrieval on embeddings into a ranking model that scores only hundreds of candidates in under 50ms.
Active users receive pre-computed feeds refreshed every few minutes, while cold-start users and cache misses fall back to trending and content-based signals. Watch events stream through Kafka to keep user features fresh without blocking video playback or the primary read path.
Component Deep Dives
Two-Stage Architecture: Candidate Generation to Ranking
Two-Stage Retrieval and Ranking Funnel
The synchronous request path operates as a two-stage funnel: Stage 1 retrieves approximately 500 candidates from a FAISS ANN index in under 10ms, while Stage 2 runs a cross-network ranking model on that shortlist in under 40ms. All other processes, including feature updates, model training, and A/B experiments, remain off the hot path via Kafka.
Event Bus Design (Kafka)
Watch behavior is the lifeblood of personalization. The player SDK emits idempotent watch events, and Flink consumers update the Redis feature store in near real time while a separate training pipeline batches historical data for nightly model retrains.
Topic: watch-events
Partitions: 256
Partition key: user_id (preserves per-user watch sequence for feature updates)
Retention: 7 days (training data replay)
Replication factor: 3, min.insync.replicas: 2
Producer: video player SDK (idempotent producer)
Event: { user_id, video_id, watch_pct, duration_sec, timestamp, device }
Consumer groups:
1. realtime-features: Flink updates user_features:{user_id} and user_watch_history List (last 50)
2. training-pipeline: batch to S3, followed by Spark feature engineering and weekly model retrain
3. ctr-computer: join recommendation-served and recommendation-clicked for A/B metrics
Topic: video-published: triggers embedding generation and ANN index update
Topic: model-events: A/B framework automated rollback upon metric regression
Sync path: GET /recommendations reads Redis feature store and Faiss ANN in under 50ms
Async path: watch events never block playback, and near-real-time features update in under 5 minutes
DLQ: watch-events-dlq, with alerts when Flink consumer lag exceeds 300sEmbedding-Based Retrieval and ANN Search
Embedding-Based Retrieval and ANN Search
User and video embeddings reside in a shared 256-dimensional vector space. The two-tower model trains offline so that the dot product of user and video embeddings correlates directly with watch probability. This mathematical alignment enables ANN search at 500M-video scale. Without separate user and item towers, pre-computing embeddings and searching in sub-10ms would be impossible.
User and video embeddings in 256-dim vector space. Two-Tower Model training: User tower: user_id embedding + avg watched video embeddings + demographics Video tower: video_id embedding + title embedding + category + channel Objective: Maximize dot(user_emb, pos_video_emb) - dot(user_emb, neg_video_emb) ANN Index (FAISS IVF-PQ): 500M videos x 256 x 4 bytes = 512 GB raw → compressed to 32 GB with IVF-PQ IVF: cluster into 4096 groups → search nearest 50 groups PQ: compress each vector to 64 bytes (16x compression) Query time: < 5 ms per user embedding → serves 200K QPS across 10 replicas
Near-Real-Time Feature Updates
Near-Real-Time Feature Updates
Recommendations feel fresh when user features update within minutes of a watch event. The streaming path updates lightweight features in Redis, while the heavy two-tower retraining job runs nightly and rebuilds the FAISS index.
User watches a video at 3:00 PM. By 3:05 PM, recommendations reflect this.
Pipeline:
1. User finishes video → watch event to Kafka
2. Flink streaming job:
a. Update user's recent watch history (last 50)
b. Update user interest vector (exponential moving average)
c. Lightweight user embedding update (avg of last 50 watched)
d. Update video stats (views, avg_watch_pct)
3. Next recommendation request (< 5 min later) uses updated features
Full model retraining (daily):
Spark job: extract all watch events from last 30 days
Train two-tower model on GPU cluster (4-8 hours)
Generate new embeddings for ALL users and videos
Build new FAISS index → blue-green deployment swapPost-Ranking Filters: Diversity and Quality
Post-Ranking Filters: Diversity and Quality
The ranking model optimizes engagement, while post-ranking filters enforce essential product constraints such as diversity, freshness, quality floors, and business rules so the feed does not devolve into a monoculture from a single channel.
After ranking model outputs top 100 scored candidates: 1. Diversity filter: Max 3 videos per channel, 30% per category. MMR algorithm. 2. Freshness boost: < 24h old → 1.2x, < 7 days → 1.1x 3. Quality filter: Remove < 40% avg watch pct, high dislike ratio, flagged content 4. Explore vs Exploit: 90% exploitation + 10% exploration (ε-greedy) 5. Business rules: Inject ads, boost premium, suppress blocked creators 6. Dedup: Remove watched, too-similar (pHash), reuploads
Cold Start for New Users and New Videos
Cold Start for New Users and New Videos
Cold start is where many recommendation system designs stumble in interviews. Address this with a deliberate fallback ladder (trending content, followed by content-based heuristics, and finally collaborative filtering) that gradually hands off to full personalization as watch history accumulates.
New User: - Onboarding: select interests → initialize interest vector - Fallback: popular by country, language, time of day, referral source - Progressive: 80% popular → 40% → 10% as watch history accumulates New Video: - Content-based embedding from title/description/tags (BERT → 256-dim) - Channel prior: inherit channel's avg engagement metrics - Exploration allocation: 10% slots for < 24h old videos - Multi-armed bandit (Thompson Sampling): learn CTR after ~1000 impressions - Blend: blended_emb = α x content_emb + (1-α) x collab_emb (α decays from 1.0 to 0.2)
API Design
Recommendation Endpoints
Home feed, up-next, similar videos, and explicit feedback each hit the same retrieval and ranking pipeline with different context headers. The examples below show the response shapes interviewers expect, including the reason field for explainability and source for debugging which candidate generator produced each item.
Get Home Feed Recommendations
GET /api/v1/recommendations/home?limit=20&cursor={last}
Headers: X-User-Id: user-uuid, X-Context: { "device": "mobile", "time_zone": "America/New_York" }
Response: 200 OK
{
"recommendations": [
{
"video_id": "vid-uuid",
"title": "System Design: URL Shortener",
"channel": "TechChannel",
"thumbnail": "https://cdn.example.com/thumb/vid-uuid.webp",
"duration": 1200,
"view_count": 1200000,
"published_at": "2025-03-10",
"score": 0.95,
"reason": "Because you watched 'System Design Interview Guide'",
"source": "collaborative_filtering"
}
]
}Get Up Next
GET /api/v1/recommendations/up-next?current_video={video_id}&watch_time=600&total_duration=1200
Response: 200 OK
{
"up_next": {
"video_id": "vid-next",
"title": "System Design: API Rate Limiter",
"reason": "Next in series",
"autoplay_in_seconds": 5
},
"related": [...]
}Get Similar Videos
GET /api/v1/recommendations/similar/{video_id}?limit=10
Response: 200 OK
{
"similar": [
{"video_id": "...", "title": "...", "similarity_score": 0.89, "reason": "Similar topic"}
]
}Send Feedback
POST /api/v1/recommendations/feedback
{
"video_id": "vid-uuid",
"action": "not_interested",
"reason": "already_watched"
}
Response: 200 OKCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 504 Gateway Timeout: search index shard responded slowly, narrow query parameters or retry
Data Model
Redis: Feature Store and Cache
Redis serves double duty as a real-time feature store and a recommendation cache. User and video embeddings, watch history, pre-computed similar-video lists, and per-context recommendation caches all live here with TTLs tuned to freshness vs memory cost. Kafka topics feed the training loop and CTR measurement without touching the synchronous read path.
# User features (updated in near-real-time by Flink)
user_features:{user_id} → Hash {
embedding: binary(1024),
preferred_categories: "tech,science,education",
avg_watch_duration: 480,
activity_level: "high",
country: "US",
language: "en",
last_updated: 1710400000
}
# User's recent watch history
user_watch_history:{user_id} → List [video_id_1, ..., video_id_50]
# Video features
video_features:{video_id} → Hash {
embedding: binary(1024), duration: 1200, view_count: 1200000, like_ratio: 0.95, avg_watch_pct: 0.72
}
TTL: 86400
# Pre-computed similar videos (item-item)
similar:{video_id} → List [vid_1, ..., vid_100]
TTL: 86400
# Cached recommendations per user
recs:{user_id}:{context} → List [vid_1, ..., vid_100]
TTL: 300
# Trending videos
trending:global → Sorted Set
trending:{category} → Sorted Set
TTL: 900Kafka Topics
Topic: watch-events (user watch behavior → training data + real-time features) Topic: recommendation-served (what was shown → for CTR computation) Topic: recommendation-clicked (what was clicked → for CTR computation) Topic: video-published (new video → needs embedding generation) Topic: model-events (model deployed, A/B test started/ended)
Fault Tolerance
| Concern | Solution |
|---|---|
| Feature store (Redis) down |
|
| FAISS index unavailable |
|
| Ranking model error | Circuit breaker → serve candidates by score from candidate generation (skip ranking) |
| Cold user (no history) | Trending + popular + category-based recommendations |
| Cold video (just uploaded) | Use content-based features (title, description, category) for initial embedding |
| Embedding drift |
|
| A/B test regression | Auto-rollback if treatment decreases avg watch time by > 2% |
| Thundering herd (homepage) | Pre-compute recommendations for active users every 5 minutes |
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50 and arch75 depth unless interviewing for a staff-level role.
- Two-stage funnel: ANN retrieval of roughly 500 candidates, ranking the top 20 (5 min)
- Two-tower model for Stage 1 embeddings and FAISS vector indexing (6 min)
- Cross-network ranker on shortlisted candidates using Redis feature stores (5 min)
- Pre-computing recommendations every 5 minutes for active users (5 min)
- Cold start handling: trending content, content features, and diversity injection (4 min)
- Frame the architecture as a two-stage funnel: candidate generation (billions down to hundreds) followed by ranking (hundreds down to 20), never ranking the full catalog per request.
- Explain the two-tower model for Stage 1: offline user and video embeddings paired with ANN search (FAISS or ScaNN) for sub-10ms retrieval.
- Cover Stage 2 cross-network ranking on approximately 200 candidates using fresh features from a Flink-updated feature store in Redis.
- Discuss the hybrid pre-compute pattern: refreshing base recommendations every 4 hours and re-ranking on demand using real-time watch signals.
- Mention the feedback loop: weighting implicit signals (watch time, completion percentage, skips) over raw CTR to avoid clickbait, and logging all interactions for model retraining.
- Address filter bubbles through epsilon-greedy exploration and diversity constraints (capping at 3 videos per channel and 30% per category).
- Avoid the common pitfall of attempting to use a cross-network model for candidate generation, as it cannot be indexed with ANN and requires billions of forward passes per request.
Engineering Trade-offs
Two-Tower vs Cross-Network
Recommendation systems trade model accuracy against inference latency, making the choice between two-tower retrieval and cross-network ranking the foundational architectural fork.
Stage 1 and Stage 2 require fundamentally different model architectures. A two-tower model pre-computes separate user and video embeddings offline so FAISS can perform ANN searches across billions of videos, though it cannot model rich cross-features between user and item. A cross-network ranker handles that deep interaction modeling on approximately 500 candidates per request. YouTube's production stack employs both: two-tower models for candidate retrieval and Wide & Deep models for final ranking.
Two-Tower Model (candidate generation): ✓ Offline computation, ANN-searchable, scales to billions ✗ Limited interaction modeling (no cross-features) Cross-Network (ranking): ✓ Models complex feature interactions ✗ Requires per-pair forward pass (1000 inferences/request) ✗ Can't do ANN search YouTube's architecture: Two-Tower (Stage 1) + Wide & Deep (Stage 2)
Watch Time vs CTR
Optimizing click-through rate alone rewards sensational clickbait where users click but abandon playback almost immediately. Watch time and watch completion percentage capture true user satisfaction more effectively, though they introduce a bias toward long-form videos. Production recommendation engines balance these tradeoffs through a multi-objective loss function weighted by likes, shares, and survey feedback.
CTR: Maximize P(click) → clickbait wins, low satisfaction Watch Time: Maximize E[watch_time] → rewards engaging content ✗ Bias toward long videos (60 min x 50% = 30 > 5 min x 100% = 5) Best: E[watch_time x watch_percentage] + multi-objective YouTube 2019+: w1xE[watch_time] + w2xP(like) + w3xP(share) - w4xP(dislike)
Batch Pre-Computation vs Real-Time Inference
Executing full real-time inference at 200K requests per second quickly overwhelms GPU serving clusters. The hybrid pattern resolves this bottleneck by pre-computing a base candidate shortlist every few hours and re-ranking on demand using fresh Flink-updated features, achieving sub-15ms latencies alongside near-real-time personalization.
Hybrid approach ⭐ (YouTube/Netflix):
1. Pre-compute "base" recs every 4 hours (recs_base:{user_id} → top 200)
2. Real-time personalization on request:
- Read base from Redis
- Fetch fresh user features (Flink-updated)
- Re-rank 200 candidates with latest features (fast)
- Apply post-rank filters
3. If base stale → trigger full pipeline, show cached while computing
Result: < 15 ms latency with near-real-time personalizationFilter Bubble Prevention
Engagement-only ranking inevitably creates ideological and thematic filter bubbles. Epsilon-greedy exploration, Thompson Sampling on uncertain predictions, and deterministic diversity constraints (limiting channels to 3 videos and categories to 30%) deliberately sacrifice marginal watch time in exchange for catalog discovery and creator ecosystem fairness.
Strategies: 1. ε-greedy: 10% random high-quality content 2. Thompson Sampling: explore where model is uncertain 3. Contextual Bandits: learn when to explore 4. Interest expansion: suggest boundary topics 5. Diversity constraints: max 3 per channel, 30% per category, min 2 fresh
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.