Interview Setup
Interview Prompt
Design TikTok, a short video platform where users upload 15 second to 10 minute videos and scroll through an AI personalized For You Page. With 500M daily active users watching an average of 150 videos each day, the recommendation engine serves as the platform's central product.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Should the architecture prioritize the interest based For You Page or the social graph following feed as primary? | The For You Page recommendation engine processes an estimated ~43.4K feed page requests per second when each response carries 20 videos, while the corresponding playback volume is ~870K video plays per second. The following feed relies on straightforward social graph fan out. |
| How quickly must newly uploaded videos become available for playback? | Determines transcoding pipeline latency vs publishing before all bitrates finish encoding, because users expect the automated pass path to make short clips available in under 60 seconds, while borderline content can remain in processing during human review. |
| What engagement signals should drive video ranking among watch time, likes, shares, and completion rate? | Prioritizing watch completion time over likes prevents clickbait, while completion rate is critical for 15 second clips and directly influences feature store design. |
| Should content safety moderation operate as a blocking upload gate or execute asynchronously? | Blocking moderation adds latency to the upload path, whereas asynchronous moderation returns quickly while content remains undiscoverable until it clears the required verification state. |
Scope
In scope
- Interest graph vs social graph architecture
- Engagement loop optimization and feature store
- Cold-start exploration for new videos and users
- Short-video ingestion and distributed transcoding pipeline
- Content safety and automated moderation at upload
- Capacity estimation with complete mathematical derivations
Out of scope (state explicitly)
- Generic recommendation architecture deep dive (see Recommendation System)
- Full ads auction and monetization stack
- Content moderation at scale (see Content Moderation System)
- Direct messaging and user to user chat (see Real-Time Chat)
Functional Requirements
TikTok is a broad product, so narrow the scope quickly with your interviewer. Focus primarily on video upload, feed recommendation, and user engagement, while confirming upfront that live streaming and creator monetization remain out of scope for the core design.
- Upload short videos: Support 15 second to 10 minute videos with background music, visual effects, and filter overlays.
- For You Page (FYP): Deliver an AI driven, personalized, infinite scroll video feed tailored to viewer interests.
- Social features: Provide interactive capabilities including likes, comments, shares, follows, duets, and stitches.
- Video creation tools: Support in app recording, multiple track audio overlays, trimming, and effects rendering.
- Search and Discover: Enable discovery by hashtags, original sounds, creator handles, and trending topic categories.
- Notifications: Deliver timely alerts for likes, comments, mentions, new followers, and viral trends through an asynchronous pipeline with preference filtering, deduplication, rate limiting, and APNs or FCM delivery.
- Creator analytics: Surface metrics covering video views, engagement rates, audience demographics, and traffic sources.
- Live streaming: Real-time broadcast capabilities with virtual gifting (addressed as out of scope).
- Monetization: Creator rewards, brand marketplace integrations, and in app purchases (addressed as out of scope).
Non-Functional Requirements
The platform demands sub-500ms feed loading and smooth global video playback. Interviewers typically prioritize recommendation serving latency and CDN edge delivery efficiency.
- Low Latency Feed: The For You Page loads within 500 milliseconds with seamless continuous playback enabled by client prefetching.
- High Throughput: Sustain more than 75 billion video plays per day across 500 million daily active users.
- Global CDN Distribution: Cache video segments at worldwide edge Points of Presence to target under 100 milliseconds Time-To-First-Byte.
- Recommendation Relevance: Optimize the recommendation pipeline for watch completion and repeat engagement as the core product differentiator.
- Upload Processing Velocity: For the automated pass path, transcode, inspect, and make uploaded media available globally within 60 seconds for short clips. Borderline moderation cases may remain in processing while human review completes.
- High Availability: Achieve 99.99% uptime for feed retrieval by providing graceful fallbacks to cached recommendation candidate pools.
- Massive Scalability: Support over 1 billion monthly active users and 10 million daily video uploads.
Capacity Estimations
Daily active users, watch session duration, and feed request rates directly govern storage growth, CDN edge bandwidth provisioning, and recommendation infrastructure capacity.
| Metric | Calculation | Value |
|---|---|---|
| DAU | Given platform scale | 500M |
| Videos watched / user / day | Average session duration of 30 to 60 minutes | 150 |
| Total video plays / day | 500M users x 150 plays | 75B |
| Video uploads / day | Given platform creator volume | 10M |
| Avg video size (encoded) | Aggregated multiple bitrate ladder (360p to 1080p) | 15 MB |
| Upload storage / day | 10M uploads x 15 MB | 150 TB |
| CDN bandwidth / day | 75B plays x 3 MB average consumed | 225 PB |
| Video plays / sec | 75B plays ÷ 86,400 seconds | ~870K |
| Feed page requests / sec | 75B plays ÷ 20 videos per response ÷ 86,400 seconds | ~43.4K (assuming 20 videos consumed per feed response) |
| Raw video embedding footprint | 10M videos x 128 dimensions x 4 bytes | ~5.12 GB per replica before HNSW overhead |
Architecture Diagram
The architecture separates asynchronous video upload, transcoding, moderation, analytics, trending, and notification pipelines from real time recommendation serving and feed delivery. The feed delivery path is heavily read intensive, retrieving precomputed candidates, applying lightweight session aware ranking, and caching personalized results for predictable sub-500ms service latency.
Component Deep Dives
For You Page Recommendation Pipeline (Two Tower Retrieval and Neural Ranking)
Focus on the feed request path first because recommendation latency and relevance drive user retention, while video upload and transcoding execute asynchronously in the background. The recommendation pipeline processes items through a staged funnel. It performs lightweight candidate recall using two tower vector embeddings, deep scoring using neural rankers, and business logic reranking for category diversity and user safety.
two_tower_retrieval_and_ranking_funnel:
stage_1_candidate_retrieval:
model: "Two Tower Dual Encoder"
corpus_size: "10,000,000 active videos"
candidate_output: "10,000 video candidates"
latency_budget: "30 milliseconds"
user_tower:
inputs: "Demographics, long term interest vectors, historical completion rates, session language"
output: "128-dimensional user embedding vector computed once per session and refreshed on interactions"
video_tower:
inputs: "Multimodal video features (audio spectrogram, visual keyframes, title text, hashtags)"
output: "128-dimensional video embedding vector precomputed offline during video ingestion"
vector_indexing: "Hierarchical Navigable Small World (HNSW) graph index in vector database"
scoring_metric: "Cosine similarity / dot product: dot_product(user_vector, video_vector)"
stage_2_deep_ranking:
model: "Single Tower Neural Ranker (Mixture of Experts)"
candidate_input: "10,000 candidates from Stage 1"
lightweight_pre_rank: "Reduce 10,000 candidates to 1,000 candidates before expensive neural inference"
neural_inference_input: "1,000 candidates"
ranked_output: "500 ranked videos"
latency_budget: "50 milliseconds"
features: "Dense cross features combining [user_features, video_features, interaction_history, context_features]"
prediction_targets:
- "P(complete): Probability of watching full video duration"
- "P(rewatch): Probability of looping the video a second or third time"
- "P(share): Probability of sharing video externally or to friends"
- "P(comment): Probability of leaving a comment"
- "P(like): Probability of tapping the heart icon"
- "P(skip): Probability of quickly swiping away"
utility_score_formula: "Score = w1*P(complete) + w2*P(rewatch) + w3*P(share) + w4*P(comment) + w5*P(like) - w6*P(skip)"
stage_3_business_rules_and_diversity:
candidate_input: "Top 500 ranked videos"
final_output: "200 curated video IDs stored in Redis feed buffer"
latency_budget: "10 milliseconds"
filtering_rules:
- "Category diversity: Maximum 2 consecutive videos from the same creator or hashtag theme"
- "Freshness exploration: Epsilon-greedy 10% slot reservation for newly uploaded videos"
- "Author fatigue penalty: Reduce ranking weight for creators viewed within the last 2 hours"
- "Safety verification: Exclude videos undergoing post publish report investigation"Precomputed Candidate Buffering vs On-Demand Generation
TikTok balances recommendation freshness against serving latency by precomputing candidate pools for active users in background workers. The system caches the top 200 video identifiers in Redis to provide predictable sub-50ms serving latency for the feed buffer.
precomputed_feed_buffering:
storage_topology:
feed_buffer_key: "feed:{user_id}"
data_structure: "Redis List"
buffer_capacity: "200 precomputed video IDs"
ttl: "2 hours"
read_latency: "< 50 milliseconds target via LRANGE feed:{user_id} 0 19"
request_reranking: "After loading cached IDs, apply lightweight session aware reranking using recent completions, skips, and seen state before returning up to 20 videos"
cache_lifecycle_and_invalidation:
background_worker_cadence: "Evaluates active users every 30 minutes to replenish candidate buffers"
active_user_optimization: "Recompute only for users with an active session plus a low buffer, significant feature change, or scheduled refresh opportunity"
idle_user_fallback: "Users inactive for > 1 hour use their retained buffer while it exists. If it has expired, fall back to regional trending candidates while refreshing personalization"
client_prefetching_strategy:
initial_load: "Client requests 20 videos on startup, immediately playing video 1 while buffering 2 and 3"
pipeline_prefetch: "As viewer watches video N, mobile player prefetches segments for videos N+1, N+2, and N+3"
replenishment_trigger: "When the viewer reaches item 150 in the 200-item buffer, the serving layer schedules asynchronous feed replenishment"Engagement Signal Hierarchy and Feature Store
The recommendation engine evaluates implicit and explicit user behaviors along a calibrated intent hierarchy, where video watch completion rate and repeated rewatches carry substantially more weight than lightweight interactions such as likes.
engagement_signal_hierarchy:
signal_1_watch_completion_rate:
importance: "Primary ranking signal (approximately 10x more predictive of user delight than likes)"
formula: "completion_rate = min(watch_time, video_duration) / video_duration"
replay_signal: "watch_time_ratio = watch_time / video_duration. Values above 1.0 indicate replay consumption"
significance:
single_loop: "completion_rate = 1.0 indicates full narrative consumption"
multi_loop: "watch_time_ratio > 1.0 (2nd or 3rd replay loop) represents the strongest positive affinity"
signal_2_share:
importance: "High-intent explicit social recommendation"
significance: "Viewer advocates content to external peers, indicating exceptional quality"
signal_3_comment:
importance: "Strong positive engagement signal"
significance: "Viewer invests active time writing feedback, driving community interaction"
signal_4_like:
importance: "Moderate positive signal"
significance: "Low cognitive friction tap, with high volume but relatively noisy compared to watch time"
signal_5_follow_after_watching:
importance: "High-value creator loyalty signal"
significance: "Directly converts viewer interest into sustained following relationship"
signal_6_skip_or_swipe_away:
importance: "Strong negative implicit signal"
significance: "Swiping away in under 3 seconds penalizes video score and related feature weights"
signal_7_explicit_not_interested:
importance: "Strongest negative explicit signal"
significance: "Explicitly tapping 'Not Interested' heavily suppresses the creator, audio track, and topic"
training_data_and_feedback_loop:
session_events: "Viewer interactions generate real time training tuples: (user_features, video_features) -> {completed, liked, shared, skipped}"
model_retraining: "Models undergo daily batch retraining on distributed Spark/GPU clusters and hourly streaming fine-tuning so viral trends can influence models within hours"Video Ingestion, Transcoding, and Automated Moderation Pipeline
Every uploaded video passes through an automated processing pipeline that generates multiple bitrate HLS renditions, extracts audio fingerprints, and executes computer vision safety classifiers before content is admitted to public feeds.
content_moderation_pipeline:
stage_1_automated_classifiers:
execution_time: "< 5 seconds per uploaded video"
image_analysis: "Convolutional and Vision Transformer models scanning keyframes for nudity, violence, and gore"
text_analysis: "NLP transformer models evaluating captions, subtitles, and OCR video text for hate speech"
audio_analysis: "Acoustic fingerprint matching against copyrighted sound and music databases"
metadata_heuristics: "Pattern detection identifying spam, scam links, and coordinated bot behaviors"
output: "Normalized confidence score between 0.0 and 1.0 across each violation category"
stage_2_decision_routing:
auto_reject:
condition: "Confidence score > 0.95 in any severe policy category"
action: "Block video publication immediately, log audit record, and notify creator with policy citation"
human_review_queue:
condition: "Confidence score between 0.70 and 0.95"
action: "Hold video in processing state, excluding from public feeds pending human operator inspection"
auto_approve:
condition: "Confidence score < 0.70 across all policy categories"
action: "Mark the video moderation-eligible. Publication still waits for required renditions and the authoritative publish transaction"
stage_3_human_review_operations:
queue_volume: "10,000 to 50,000 borderline videos per day routed to human moderation"
operational_sla: "Complete review within 2 hours of upload"
staffing_hierarchy: "Trained moderation operations with escalation pathways to specialized legal and policy teams"
stage_4_post_publish_monitoring:
signal_source: "Live user reports and anomaly detection algorithms tracking viewer flag spikes"
deprioritize_threshold:
condition: "Accumulates more than 10 user reports"
action: "Automatically suppress recommendation weight and remove from exploratory FYP slots"
takedown_threshold:
condition: "Accumulates more than 50 user reports"
action: "Instantly remove from all discovery feeds and enqueue high-priority human review"
accuracy_benchmarks:
appeals_process: "Creators can appeal removals, triggering an independent secondary human review"
target_false_positive_rate: "< 1.0% (legitimate videos mistakenly removed)"
target_false_negative_rate: "< 0.1% (violating media mistakenly exposed to viewers)"Event Streaming and Telemetry Bus (Apache Kafka)
Apache Kafka decouples high throughput video upload events and real time viewer engagement telemetry from backend database persistence, allowing downstream stream processors and machine learning training pipelines to consume updates without blocking the video delivery path.
kafka_cluster_topology:
video_lifecycle_topic:
topic_name: "video-lifecycle"
partition_key: "video_id"
partition_count: 256
retention_period: "7 days"
replication_factor: 3
purpose: "Carries upload-ready, published, and removed lifecycle events while preserving per-video ordering"
engagement_events_topic:
topic_name: "engagement-events"
partition_key: "user_id"
partition_count: 256
retention_period: "7 days"
replication_factor: 3
purpose: "Streams viewer telemetry plus asynchronously emitted engagement events after authoritative mutation commits"
social_events_topic:
topic_name: "social-events"
partition_key: "actor_user_id"
partition_count: 128
retention_period: "7 days"
replication_factor: 3
purpose: "Streams follows, mentions, and other social mutations used by notifications and relationship features"
event_schemas:
video_lifecycle_event:
event_id: "CHAR(36) unique event identifier"
event_type: "ENUM(upload_ready, published, removed)"
video_id: "VARCHAR(36) UUID"
creator_id: "BIGINT creator identifier"
s3_key: "raw/uploads/2026/03/video-uuid.mp4"
duration_sec: "INTEGER (video length in seconds)"
timestamp: "TIMESTAMP (epoch milliseconds)"
engagement_event:
event_id: "UUID event transaction ID"
user_id: "UUID viewer identifier"
creator_id: "UUID creator identifier when applicable"
video_id: "UUID video identifier"
event_kind: "ENUM(telemetry, mutation_event)"
action: "ENUM(watch_start, heartbeat, like, share, comment, skip, not_interested)"
watch_ms: "INTEGER (playback duration in milliseconds)"
country: "STRING country or region code"
device_type: "STRING client device category"
timestamp: "TIMESTAMP (epoch milliseconds)"
social_event:
event_id: "UUID event transaction ID"
actor_user_id: "UUID acting user identifier"
target_user_id: "UUID affected user identifier"
video_id: "UUID video identifier when applicable"
action: "ENUM(follow, mention)"
timestamp: "TIMESTAMP (epoch milliseconds)"
consumer_groups:
transcoding_workers:
purpose: "Consumes upload_ready lifecycle events to execute GPU transcoding, thumbnail generation, and initial moderation processing"
rec_feature_updater:
purpose: "Apache Flink job consuming engagement-events to update Redis user and video feature vectors in real time and repartitioning by video_id where video-level aggregates are required"
vector_index_updater:
purpose: "Consumes published and removed lifecycle events, upserts eligible 128-dimensional video embeddings into the HNSW vector index, and removes or tombstones deleted videos"
search_index_updater:
purpose: "Indexes published video metadata, hashtags, creator handles, captions, and audio identifiers, then removes or tombstones deleted videos"
model_training_pipeline:
purpose: "Apache Spark ETL job dumping events to S3 Parquet for offline deep neural network retraining"
creator_analytics_pipeline:
purpose: "Aggregates engagement and social events for near real time creator dashboards and daily reconciliation in ClickHouse"
trending_pipeline:
purpose: "Apache Flink job computes five minute velocity and acceleration for hashtags, sounds, and topic categories, then writes regional top K trends to Redis for discovery and viral notification triggers"
notification_pipeline:
purpose: "Consumes committed likes, comments, and social mutations, resolves recipients from creator_id or target user IDs, applies notification preferences and deduplication, rate limits bursts, and dispatches through APNs or FCM with retry and dead letter handling"
operational_paths:
upload_and_publish_path: "Client requests a presigned URL and uploads to S3. The publish API verifies that the object exists and records video metadata plus an upload_ready lifecycle outbox event in one MySQL transaction. Debezium CDC publishes the committed outbox event to Kafka after commit"
publication_transition: "After required renditions and safety checks succeed, a transaction sets status = published and writes a published lifecycle outbox event. A removal transaction sets status = removed and writes a removed lifecycle event. Conditional state transitions prevent duplicate publish or remove events"
feed_telemetry_path: "Client emits engagement heartbeats that flow into recommendation pipelines without blocking playback. Authoritative likes, comments, and follows are persisted through their mutation APIs and publish mutation events only after commit"
notification_path: "Committed like, comment, follow, and mention events resolve the recipient, check notification preferences, deduplicate and rate limit the notification, then enqueue APNs or FCM delivery. Delivery failures retry with backoff before entering a dead letter queue"
trending_path: "Flink aggregates five minute regional velocity and acceleration for hashtags, sounds, and categories. Top K results are written to Redis, and threshold crossings can emit viral trend notification events"
dead_letter_queue: "Failed engagement events route to engagement-events-dlq, while failed lifecycle events route to video-lifecycle-dlq. Alert when the relevant consumer lag exceeds 5 minutes"
idempotency: "Consumers use event_id to deduplicate retried CDC or Kafka deliveries"Following Feed Secondary Path
The social graph following feed uses fan out on write for ordinary creators and fan out on read for high follower creators. Follower edges use separate followee and follower access paths so creator fan out and a user's following list remain local to their intended shards. Privacy and block rules are applied before candidates reach the client. The For You Page remains the primary recommendation path and does not depend on the following feed being available.
Search and Discover Secondary Path
Published video events populate a search index for hashtags, creator handles, captions, and audio identifiers. A separate streaming trend detector computes regional velocity and acceleration for hashtags, sounds, and topic categories. Removal events delete or tombstone the corresponding search documents and trend entries, while privacy and moderation status remain authoritative in the metadata store. Search or trend service failure should degrade discovery independently without blocking playback or the For You Page.
API Design
Client API Contracts
The client API contracts define endpoints for infinite scroll feed retrieval, presigned object storage upload negotiation, and video publishing workflows.
TypeScript domain signatures that define the request envelopes, video item payloads, and client service methods.
type VideoId = string;
type UserId = string;
type MusicId = string;
type FeedCursor = string;
type IdempotencyKey = string;
type VideoStatus = "processing" | "published" | "rejected" | "removed";
interface CreatorProfile {
id: UserId;
username: string;
displayName: string;
avatarUrl: string;
verified: boolean;
}
interface AudioTrack {
id: MusicId;
title: string;
artist: string;
audioUrl: string;
durationSec: number;
}
interface VideoEngagementStats {
plays: number;
likes: number;
comments: number;
shares: number;
}
interface VideoFeedItem {
videoId: VideoId;
creator: CreatorProfile;
description: string;
music?: AudioTrack;
stats: VideoEngagementStats;
signedHlsPlaylistUrl: string;
playbackExpiresAt: string;
thumbnailUrl: string;
durationSec: number;
createdAt: string;
}
interface FeedResponse {
videos: VideoFeedItem[];
cursor?: FeedCursor;
hasMore: boolean;
}
interface UploadUrlRequest {
fileSizeBytes: number;
format: "mp4" | "mov" | "webm";
}
interface UploadUrlResponse {
uploadUrl: string;
videoId: VideoId;
expiresInSec: number;
}
interface PublishVideoRequest {
description: string;
musicId?: MusicId;
privacy: "public" | "friends" | "private";
tags: string[];
}
interface VideoStatusResponse {
videoId: VideoId;
status: VideoStatus;
publishedAt?: string;
moderationReason?: string;
}
interface CreateCommentRequest {
videoId: VideoId;
text: string;
}
interface CreateCommentResponse {
commentId: string;
videoId: VideoId;
createdAt: string;
}
interface FollowCreatorRequest {
creatorId: UserId;
following: boolean;
}
interface EngagementBeaconRequest {
eventId: string;
videoId: VideoId;
action: "watch_start" | "heartbeat" | "skip" | "not_interested";
watchDurationMs: number;
playbackPositionSec: number;
}
interface LikeVideoRequest {
videoId: VideoId;
liked: boolean;
}
interface ShareVideoRequest {
videoId: VideoId;
destination: "external" | "copy_link";
}
interface MutationOptions {
idempotencyKey: IdempotencyKey;
}
// Client API interface for TikTok feed, upload, status, and interaction workflows
interface TikTokApiService {
getFeed(count: number, cursor?: FeedCursor): Promise<FeedResponse>;
requestUploadUrl(request: UploadUrlRequest, options: MutationOptions): Promise<UploadUrlResponse>;
publishVideo(videoId: VideoId, payload: PublishVideoRequest, options: MutationOptions): Promise<VideoStatusResponse>;
getVideoStatus(videoId: VideoId): Promise<VideoStatusResponse>;
createComment(request: CreateCommentRequest, options: MutationOptions): Promise<CreateCommentResponse>;
likeVideo(request: LikeVideoRequest, options: MutationOptions): Promise<{ liked: boolean }>;
shareVideo(request: ShareVideoRequest, options: MutationOptions): Promise<{ acknowledged: boolean }>;
followCreator(request: FollowCreatorRequest, options: MutationOptions): Promise<{ following: boolean }>;
sendEngagementBeacon(beacon: EngagementBeaconRequest): Promise<{ acknowledged: boolean }>;
}Get Feed (For You Page)
Fetches a batch of personalized video items with opaque cursor based pagination and short lived signed media URLs. The serving layer filters unavailable videos before returning the final batch.
GET /api/v1/feed?count=20&cursor=f-cursor-20 HTTP/1.1
Host: api.tiktok.com
Authorization: Bearer <user_session_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"videos": [
{
"video_id": "v-uuid-1",
"creator": {
"id": "u102",
"username": "alice_dance",
"displayName": "Alice",
"avatarUrl": "https://cdn.tiktok.com/avatars/u102.jpg",
"verified": true
},
"description": "Dance challenge #fyp",
"music": {
"id": "m501",
"title": "Original Sound - DJ Groove",
"artist": "DJ Groove",
"audioUrl": "https://cdn.tiktok.com/audio/m501.mp3",
"durationSec": 60
},
"stats": {
"plays": 1523000,
"likes": 89200,
"comments": 3400,
"shares": 12300
},
"signedHlsPlaylistUrl": "https://cdn.tiktok.com/v-uuid-1/playlist.m3u8?token=...",
"playbackExpiresAt": "2026-03-14T08:10:00Z",
"thumbnailUrl": "https://cdn.tiktok.com/v-uuid-1/thumb.jpg",
"durationSec": 45,
"createdAt": "2026-03-14T08:00:00Z"
}
],
"cursor": "f-cursor-20",
"hasMore": true
}Upload and Publish Video
Negotiates presigned object storage credentials and triggers asynchronous transcoding. Mutating requests use idempotency keys, and the publish path verifies the completed upload before entering the processing state.
POST /api/v1/videos/upload-url HTTP/1.1
Host: api.tiktok.com
Authorization: Bearer <user_session_token>
Idempotency-Key: upload-request-001
Content-Type: application/json
{
"fileSizeBytes": 15728640,
"format": "mp4"
}
HTTP/1.1 200 OK
Content-Type: application/json
{
"upload_url": "https://s3.us-east-1.amazonaws.com/raw-uploads/video-uuid-1?signature=...",
"video_id": "video-uuid-1",
"expiresInSec": 3600
}
POST /api/v1/videos/video-uuid-1/publish HTTP/1.1
Host: api.tiktok.com
Authorization: Bearer <user_session_token>
Idempotency-Key: publish-request-001
Content-Type: application/json
{
"description": "Dance challenge #fyp",
"music_id": "m501",
"privacy": "public",
"tags": ["dance", "fyp"]
}
HTTP/1.1 202 Accepted
Content-Type: application/json
{
"video_id": "video-uuid-1",
"status": "processing"
}
GET /api/v1/videos/video-uuid-1/status HTTP/1.1
Host: api.tiktok.com
Authorization: Bearer <user_session_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"video_id": "video-uuid-1",
"status": "published",
"published_at": "2026-03-14T08:00:50Z"
}
POST /api/v1/videos/video-uuid-1/comments HTTP/1.1
Host: api.tiktok.com
Authorization: Bearer <user_session_token>
Idempotency-Key: comment-request-001
Content-Type: application/json
{
"text": "Great video!"
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"comment_id": "c-uuid-1",
"video_id": "video-uuid-1",
"created_at": "2026-03-14T08:06:00Z"
}Common Error Responses
Standardized API error responses across the video platform.
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Mutation Safety and Abuse Controls
Upload, publish, comment, like, share, and follow mutations require authentication, payload validation, rate limits, and idempotency keys. Like and follow operations use explicit desired state rather than toggle semantics, so retries are safe. The server validates the actual media format and duration after upload. Authoritative mutation events are emitted only after the mutation commits successfully. Feed and playback responses use short lived authorization tokens so removed or private media cannot remain accessible indefinitely through stale client state.
Data Model
MySQL (Vitess): Video Catalog Metadata
The storage tier employs polyglot persistence tailored to distinct access patterns. Vitess-sharded MySQL manages relational video metadata, creator ownership, and the transactional outbox. Redis stores short lived feed buffers and active-user real time features. The HNSW vector index stores video embeddings for ANN retrieval. Cassandra stores high-velocity interaction telemetry.
Relational storage manages author ownership, privacy levels, transcoding state, and materialized engagement counters. Video rows are sharded horizontally across Vitess clusters by creator ID.
CREATE TABLE videos (
video_id VARCHAR(36) PRIMARY KEY,
creator_id BIGINT NOT NULL,
description TEXT,
music_id VARCHAR(36),
duration_sec SMALLINT,
status ENUM('processing','published','rejected','removed') DEFAULT 'processing',
moderation_reason VARCHAR(512),
privacy ENUM('public','private','friends') DEFAULT 'public',
s3_key VARCHAR(512),
thumbnail_key VARCHAR(512),
view_count BIGINT DEFAULT 0,
like_count BIGINT DEFAULT 0,
comment_count BIGINT DEFAULT 0,
share_count BIGINT DEFAULT 0,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
published_at TIMESTAMP NULL,
INDEX idx_creator (creator_id, created_at DESC),
INDEX idx_creator_status_created (creator_id, status, created_at DESC)
-- Global status-based discovery is served by the search index rather than a cross shard scan
);MySQL Transactional Outbox
Video metadata changes and the corresponding Kafka publication event are committed together, preventing the dual-write failure where the database succeeds but the upload event is lost. Debezium CDC publishes the committed outbox record asynchronously.
CREATE TABLE video_event_outbox (
event_id CHAR(36) PRIMARY KEY,
video_id VARCHAR(36) NOT NULL,
creator_id BIGINT NOT NULL,
event_type VARCHAR(64) NOT NULL,
payload JSON NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
INDEX idx_video_created (video_id, created_at)
);
-- Vitess sharding uses creator_id so the outbox row is co-located with the video metadata transactionRedis: Feed Buffers and Active User Feature Store
In-memory data structures store short lived personalized feed lists, bounded recent-seen state, and active-user features for low latency recommendation serving.
# Pre-computed personalized video feed for recently active users
Key: feed:{user_id}
Type: List
Values: video_id (precomputed, 200 items)
TTL: 2 hours
Refresh: Every 30 minutes for users active within the trailing 30 minutes
Trigger: Active session plus low buffer, significant interaction change, or scheduled refresh
# Bounded set of recently seen video IDs for per-user deduplication
Key: seen:{user_id}
Type: Sorted Set
Members: video_id
Score: last_seen_epoch_ms
Bound: Keep the most recent 500 IDs per user
TTL: 30 days
# User real time feature profile for active users
Key: user_features:{user_id}
Type: Hash
Fields: embedding (128-dim float vector), interests, language, device_type
Note: Cold-user profiles are loaded from the offline feature store on demand
# Video feature profile for scoring and candidate generation
Key: video_features:{vid_id}
Type: Hash
Fields: completion_rate, like_rate, share_rate, category, embedding_version
Note: Canonical video embeddings live in the HNSW vector index rather than being duplicated in Redis
# Regional trending leaderboard for search, discovery, and viral notifications
Key: trending:{region}
Type: Sorted Set
Members: hashtag, sound, or topic identifier
Score: recent velocity plus acceleration score
Bound: Top 10,000 active trends per region with short retentionVector Database: Video Embeddings
The vector index owns the canonical 128-dimensional video embeddings used by ANN retrieval. Regional read replicas serve low latency candidate retrieval, while embeddings are updated asynchronously when videos become eligible for discovery and removed when content is deleted.
vector_index:
store: "HNSW vector database"
entity: "published video_id → 128-dimensional embedding"
distance: "Cosine similarity or dot product on L2-normalized vectors"
corpus: "Up to 10,000,000 active published videos"
ingestion: "Asynchronously upsert embeddings after transcoding and safety eligibility"
deletion: "Remove or tombstone the vector when video status becomes removed"
serving: "ANN retrieval returns the Stage 1 candidate set"MySQL: Following Graph Access Paths
The following feed needs both creator to follower and user to followee access patterns. Maintain denormalized relational tables with separate shard directions so creator fan out and a user's following list do not require cross shard scans. Commit the canonical follow mutation once, then use its durable event to maintain both projections asynchronously and repair drift when needed. A brief projection lag is acceptable because the relationship service remains the source of truth.
CREATE TABLE creator_followers_by_followee (
followee_id BIGINT NOT NULL,
follower_id BIGINT NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (followee_id, follower_id)
);
-- Shard by followee_id so fan out can fetch a creator's followers locally
CREATE TABLE user_followees_by_follower (
follower_id BIGINT NOT NULL,
followee_id BIGINT NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (follower_id, followee_id)
);
-- Shard by follower_id so a user's following list has a local read pathClickHouse: Creator Analytics
Asynchronous engagement events are aggregated into an analytical store for creator dashboards. Near real time aggregates serve current metrics, while daily batch reconciliation corrects late or duplicated telemetry.
creator_analytics_pipeline:
source: "Kafka engagement-events and social-events"
processing: "Flink for near real time aggregates plus Spark for daily reconciliation"
store: "ClickHouse"
dimensions: "creator_id, video_id, country, device_type, time_bucket"
metrics: "plays, watch_time, completion_rate, likes, comments, shares"
freshness_target: "Near real time aggregates within 1 minute and daily batch reconciliation for correctness"Notifications and Trending State
Notification delivery is asynchronous so engagement writes are not blocked by mobile push provider latency. The notification layer resolves recipients, applies per user preferences and rate limits, deduplicates repeated events, and retries provider failures. A regional trend store supports discovery and viral notification triggers without coupling trend computation to feed serving.
notification_and_trending:
notification_preferences:
store: "MySQL"
key: "user_id"
fields: "likes, comments, mentions, follows, viral_trends, quiet_hours"
delivery_state:
store: "MySQL"
dedupe_key: "(recipient_id, event_id)"
status: "queued | sent | failed"
provider_delivery:
providers: "APNs, FCM"
retry_policy: "Exponential backoff with bounded retries, then dead letter queue"
trending_state:
store: "Redis Sorted Sets"
key: "trending:{region}"
window: "5 minute sliding window"
score: "velocity plus acceleration"
output: "Regional top K trends and notification triggers"Cassandra: Engagement and Interaction Data
Wide-column storage handles high write throughput likes and chronological comment streams using hashed and time bucketed partitions so viral videos do not create unbounded hot partitions. Comment reads query the current and recent time buckets in parallel and merge results by descending comment ID.
CREATE TABLE video_likes (
video_id UUID,
bucket SMALLINT,
user_id UUID,
created_at TIMESTAMP,
PRIMARY KEY ((video_id, bucket), user_id)
);
-- bucket = hash(user_id) % 64 to distribute viral-video writes
CREATE TABLE video_comments (
video_id UUID,
day_bucket DATE,
comment_id TIMEUUID,
user_id UUID,
text TEXT,
PRIMARY KEY ((video_id, day_bucket), comment_id)
) WITH CLUSTERING ORDER BY (comment_id DESC);
-- day_bucket bounds each viral video's comment partition by timeFault Tolerance
| Concern | Solution |
|---|---|
| Transcoding failure | Retry up to 3 times with exponential backoff, routing persistent failures to a dead-letter queue for operator inspection |
| Recommendation cache miss | Fallback to regionally cached trending and popular video candidate pools while rebuilding personalization in the background |
| Feature store or vector index unavailable | Serve the retained Redis feed buffer or regional trending candidates, and rebuild missing feature or embedding replicas from durable sources |
| CDN cache miss | Perform origin pull from S3 with origin shielding, combined with proactive edge cache warming for predicted viral videos |
| Feed service outage | Clients cache the last 50 videos locally, enabling uninterrupted playback while degraded services recover |
| Content moderation false negative | Accumulated user reports automatically suppress feed ranking and escalate media to an expedited human review queue |
Concurrency and Data Consistency Race Conditions
This section covers publication state transitions and high frequency engagement counter mutations without degrading playback availability or dropping user interactions.
race_condition_1_premature_playback:
hazard: "A creator requests publication while transcoding workers are still generating video renditions. Another user discovers the video and encounters HTTP 404 playback errors if publication becomes visible too early."
resolution:
state_machine: "Video status remains 'processing' until all multiple bitrate HLS renditions are encoded and validated in object storage."
visibility_boundary: "Feed, recommendation, and search services query only videos with status = 'published'."
creator_feedback: "Creator profile displays an active 'Processing...' indicator until transcode completion."
race_condition_2_view_counter_hotspots:
hazard: "Viral videos accumulating tens of thousands of plays per second create write locks and row contention if written directly to relational databases."
resolution:
in_memory_accumulation: "Ingress telemetry executes atomic Redis INCR operations for real time counter rendering."
asynchronous_durable_flush: "Counts are flushed via Kafka and aggregated in micro-batches before updating MySQL/Vitess."
display_consistency: "Eventual consistency allows counts to lag by several seconds, which remains acceptable for high volume display counters."
recovery: "If Redis counter state is lost, use the durable MySQL aggregate as the baseline and replay retained Kafka events to rebuild recent deltas before serving recovered values."
fraud_prevention:
duplicate_delivery_rule: "Ignore retransmitted copies of the same beacon using event_id so client retries do not create duplicate events."
impression_dedup_window: "For impression counting, cap duplicate starts for the same user and video within 30 seconds while preserving watch duration and replay signals for ranking."
deduplication_command: "SET engagement_dedup:{event_id} 1 EX 86400 NX"
race_condition_3_post_publish_takedown:
hazard: "A video can be removed after recommendation, search, or feed buffers have already referenced its video ID or after a viewer has received a signed playback URL."
resolution:
authoritative_state: "MySQL transitions the video status to 'removed' and emits a durable removal event through Kafka."
downstream_invalidation: "Recommendation, search, and feed services consume the event and remove the video ID from candidate pools and cached feeds."
playback_boundary: "New signed playback URLs are denied for removed videos, and the CDN purges existing objects when the safety SLO requires it."
client_behavior: "Clients stop playback when authorization expires or a removal response is received and discard cached metadata."Additional Considerations
Related Problems and Core Concepts
Explore related problem architectures and core distributed systems concepts to understand how video streaming, recommendation engines, and high throughput content systems interconnect.
- Design a Recommendation System : Deep dive into multi-stage candidate retrieval, approximate nearest neighbor (ANN) indexing, feature stores, and cold start exploration strategies.
- Design a Video Recommendation Engine : Multi-task neural ranking, watch-time optimization, and session based sequence modeling.
- Design a Video Streaming Platform (YouTube / Netflix) : Multi-rendition video encoding, CMAF chunk packaging, adaptive bitrate streaming, and distributed origin shielding.
- Design a Content Moderation System : Multi-modal automated safety classifiers, human review routing queues, perceptual hash matching, and real time takedowns.
- Design a Real-Time Chat System : Direct user messaging, WebSocket gateway clusters, and presence synchronization.
- CDN and Edge Delivery : Edge caching hierarchies, origin shielding, cache invalidation, and byte range request routing.
- Sharding and Partitioning : Horizontal data partitioning, shard key selection, and scaling relational clusters with Vitess.
- Stream Processing Basics : Stateful stream processing with Apache Flink, event-time windowing, and real time feature extraction.
- System Design Interview Patterns : Structured frameworks for tackling high-concurrency systems and communicating architectural trade-offs.
Interview Walkthrough
- 25-minute cut
Focus on the decoupled video ingestion pipeline and two tower recommendation architecture before diving into CDN origin shielding and cold start exploration.
- Split the whiteboard early into two decoupled asynchronous paths: video delivery (upload, transcoding, CDN distribution) and the For You Page recommendation engine (feature extraction, retrieval, ranking) (5 min)
- Walk the media upload path sequentially from client presigned S3 upload through the transactional outbox, transcoding worker queues, and CDN edge propagation (6 min)
- Budget FYP latency across individual serving stages: 20ms for feature store lookups, 30ms for vector ANN retrieval, 50ms for neural ranker evaluation, and 10ms for business reranking, totaling approximately 110ms of service budget before network and client overhead (5 min)
- Precompute candidate pools in Redis refreshed every 30 minutes for active users to provide predictable low latency feed loading (5 min)
- For staff-level discussions, detail cold start exploration algorithms, content diversity guardrails, and real time engagement telemetry loops via Kafka (4 min)
- Split the problem into two distinct pipelines: an asynchronous upload and transcoding pipeline for media delivery, and an independent real time recommendation engine for the For You Page.
- Describe the upload path: the client uploads media directly to object storage via presigned URLs, emits an upload_ready event to the transcoding queue, and targets generation and validation of the required multiple bitrate HLS streams within 60 seconds on the normal successful path.
- Propose a two tower neural model where user and video embeddings allow Approximate Nearest Neighbor search to retrieve the Stage 1 pool of 10,000 candidates in a 30 millisecond target, followed by lightweight preranking and neural scoring.
- Pre-compute candidate pools in Redis refreshed every 30 minutes, and re-rank per request using real time session features such as recent completions and skips.
- Weight engagement signals by intent, prioritizing watch completion and multi loop rewatches above shares and likes, while treating quick scrolls past videos as strong negative feedback.
- Serve video segments directly from edge CDNs to achieve sub 100ms first-byte latency, keeping heavy feed assembly and metadata lookups within application servers.
- Budget total FYP latency across feature retrieval (20ms), ANN candidate search (30ms), GPU ranking (50ms), and business reranking (10ms), comfortably meeting the 500ms SLA.
- Avoid the common pitfall of attempting to score the entire catalog synchronously on every scroll, because without vector ANN recall and precomputed candidate buffers, a sub-500ms latency target is difficult to meet consistently.
Engineering Trade-offs
Candidate Retrieval: Two Tower Dual Encoder vs Deep Neural Ranking
This section evaluates architectural trade offs across retrieval algorithms, feed caching topologies, CDN edge delivery mechanics, and trust and safety gates.
This section evaluates the trade off between computational throughput and model expressiveness when filtering millions of video candidates within tight latency budgets.
candidate_retrieval_comparison:
two_tower_dual_encoder:
approach: "Computes user and video embeddings independently, scoring candidates via vector dot product"
latency: "Target under 30 milliseconds to retrieve the Stage 1 candidate set from 10,000,000 indexed videos using HNSW Approximate Nearest Neighbor search"
compute_efficiency: "High efficiency because video embeddings are precomputed offline during video upload"
strengths:
- "Enables sub linear nearest-neighbor search over massive catalog inventories"
- "User tower runs inference once per session and updates incrementally on interactions"
trade_offs:
- "Cannot capture real time cross features (such as how a specific user interacts with specific video traits)"
- "Less expressive than cross attention neural architectures for nuanced reranking"
verdict: "Primary choice for Stage 1 Candidate Retrieval (recalling top 10K candidates from 10M videos)"
single_tower_neural_ranker:
approach: "Feeds concatenated user, video, and context features directly into a deep neural network"
latency: "Approximately 50 milliseconds target to score up to 1,000 candidates on GPU"
compute_efficiency: "Requires O(N) GPU inference operations per request, making catalog-wide evaluation cost-prohibitive"
strengths:
- "Models deep non linear interactions between viewer preferences and multimodal video attributes"
- "Yields highly accurate multiple objective engagement predictions for P(complete), P(rewatch), and P(share)"
trade_offs:
- "Cannot precompute scores offline because inputs depend dynamically on real time viewer context"
- "Inference cost scales linearly with candidate pool size, requiring strict pre-filtering"
verdict: "Primary choice for Stage 2 Ranking after lightweight preranking reduces the 10,000 candidate pool to 1,000 neural ranking candidates"Feed Generation: Precomputed Redis Buffers vs On-Demand Synthesis
This section balances recommendation freshness against serving latency and infrastructure compute costs at 500M DAU scale.
feed_generation_comparison:
precomputed_feed_buffering:
architecture: "Background worker fleet periodically computes and caches top 200 video IDs in Redis per active user"
serving_latency: "< 50 milliseconds target for the Redis buffer read"
compute_profile: "Evenly distributed background batch compute, preventing peak-hour GPU saturation"
strengths:
- "Provides predictable low latency playback initiation for smooth infinite scroll"
- "Isolates viewer read requests from recommendation service outages or inference slowdowns"
trade_offs:
- "Feed staleness: Recent user interactions occurring within the trailing 30 minutes are not reflected immediately"
- "Wasted compute: Generates feeds for accounts that may not open the application before the next cycle"
optimization: "Gated recomputation: Recalculate when the user is active and the buffer is low, the feature state changes materially, or a scheduled refresh is due"
pure_on_demand_synthesis:
architecture: "Every feed request invokes the recommendation model pipeline to synthesize results dynamically"
serving_latency: "200 to 500 milliseconds (dominated by GPU queue depth and multi-stage inference)"
compute_profile: "Spiky workload that tracks active user concurrency, requiring substantial peak GPU provisioning"
strengths:
- "Maximum freshness: Incorporates the viewer's immediate previous swipe, skip, or like into the next item"
- "Minimizes wasted compute because recommendations are generated when a user is actively requesting content"
trade_offs:
- "Elevated latency degrades first-video playback start times, directly hurting session retention"
- "Vulnerable to cascading service degradation during viral events and global traffic surges"
hybrid_production_solution:
recommendation: "TikTok production model combining precomputed candidate buffers with on demand replenishment"
execution: "Serve the first 20 videos immediately from Redis cache, while triggering an asynchronous background replenishment job when the user consumes past item 150 of the 200-item buffer"CDN Edge Delivery: Serving 225 PB per Day
This section optimizes edge egress bandwidth through adaptive bitrate ladders, predictive prewarming for viral assets, and byte range request routing.
cdn_scale_and_egress_metrics:
daily_egress_volume: "225 PB per day"
average_egress_throughput: "2.6 TB per second"
peak_egress_throughput: "5.0+ TB per second"
optimization_strategies:
adaptive_bitrate_ladder:
encoding_renditions:
- resolution: "360p"
bitrate: "0.5 Mbps"
- resolution: "480p"
bitrate: "1.5 Mbps"
- resolution: "720p"
bitrate: "3.0 Mbps"
- resolution: "1080p"
bitrate: "5.0 Mbps"
client_heuristics: "Playback initializes on lower rendition for instant start, then upgrades dynamically based on network throughput"
bandwidth_savings: "Over 60% of views occur on mobile screens where 480p provides optimal perceived quality while saving significant bandwidth"
predictive_edge_prewarming:
trigger_condition: "Streaming Flink job detects view velocity exceeding viral acceleration thresholds"
execution: "Pushes HLS media segments to edge Points of Presence (PoPs) in target geographic regions before viewers request them"
benefit: "Prevents thundering herd origin pulls against central object storage when content goes viral"
cache_hit_ratio_economics:
asset_footprint: "A 30-second video encoded at 720p consumes approximately 5 MB of storage"
edge_capacity: "Illustrative: a CDN edge node with 1 TB of NVMe cache can retain approximately 200,000 5 MB video assets before metadata and filesystem overhead"
power_law_distribution: "Under the stated power-law assumption, the top 200,000 viral videos account for over 80% of global daily plays, supporting an 80%+ edge cache hit ratio"
byte_range_and_chunked_streaming:
segment_duration: "2-second Common Media Application Format (CMAF) / HLS chunk segments"
seek_behavior: "When a user seeks within a video, the client requests the relevant 2-second media segment. Byte range addressing may be used where supported to reduce Time-To-First-Byte (TTFB)"Content Moderation Gates: Blocking Upload Gate vs Asynchronous Pipeline
This section weighs immediate creator upload confirmation against the risk of policy violating media escaping into live recommendation pools.
content_moderation_gate_comparison:
synchronous_blocking_gate:
workflow: "Upload API blocks acknowledgment until deep multimodal classifiers and human review queues clear the video"
safety_guarantee: "Every video must clear the defined verification gate before it becomes discoverable or playable"
creator_experience: "Poor: Creators wait minutes or hours before receiving confirmation, discouraging frequent uploads"
infrastructure_impact: "Requires holding client connections open or polling extensively, increasing ingress server connection load"
asynchronous_two_phase_pipeline:
workflow: "Upload API acknowledges immediately (HTTP 202) while media enters 'processing' status during background AI classification"
safety_guarantee: "High: Known bad hashes are blocked within 5 seconds, while multimodal classifiers complete before status transitions to 'published'"
creator_experience: "Excellent: Upload completes in seconds with clear UI progress indications"
safeguards:
- "Videos remain invisible to recommendation and search services until status equals 'published'"
- "Post publish reporting thresholds are scenario policy: more than 10 flags deprioritizes media and more than 50 flags suppresses media"
verdict: "Recommended for this scale, balancing rapid creator feedback with layered safety gates"Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.