Interview Setup
Interview Prompt
Design a thumbnail generation service: on video/image upload, produce default thumbs, candidate frames, sprite sheets, and animated previews for CDN-backed feeds.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Video vs image upload ratio? |
|
| Who selects default thumbnail? | ML auto-pick vs creator manual changes UX and APIs. |
| Sprite sheet for seek scrubbing? | Extra FFmpeg pass + VTT per video. |
| Re-processing on same content_id? | Idempotent S3 keys prevent duplicate CDN objects. |
Scope
In scope
- Async job queue
- Video frame extraction
- Image resize/WebP
- Sprite+VTT
- CDN cache
- Idempotent generation
Out of scope (state explicitly)
- Full video transcoding pipeline
- Feed ranking
- Creator UI
Functional Requirements
Start by asking your interviewer whether thumbnails are pre-generated on upload or produced on-demand, as that single architectural decision drives queue sizing and CDN strategy. Thumbnail processing works in close coordination with the Video Transcoding Pipeline and the Image Processing Pipeline to deliver fast asset previews.
- Auto-generate thumbnails: From videos, images, PDFs, documents
- Multiple sizes: Small 150x150, medium 300x200, large 640x360
- Video thumbnails: Extract the "best" frame using ML scoring
- Sprite sheets: Grid of thumbnails for video seek preview
- Custom thumbnails: Upload or select from candidates
- A/B test thumbnails: Serve variants, measure CTR
- Animated thumbnails: Short WebP preview on hover
- Format optimization: Serve WebP/AVIF with JPEG fallback
Non-Functional Requirements
Thumbnails are immutable once generated, so lean into CDN cacheability as your primary latency strategy. The 30-second generation SLA is a write-path concern, whereas reads should be <50ms from edge cache with content-addressed URLs.
- Speed: Thumbnails ready within 30 seconds of upload
- Quality: Visually appealing, representative frames
- Scale: Process 50K+ videos/hour and 500K+ images/hour
- Cacheability: Highly cacheable on CDN (immutable URLs)
- Cost Efficient: Avoid unnecessary regeneration
- Availability: 99.9%
Capacity Estimations
Worker concurrency equals throughput multiplied by average job duration. Video jobs run 10 to 100 times longer than image resizes, so quoting a single worker count without separating queues will prompt pushback from your interviewer.
| Metric | Calculation | Value |
|---|---|---|
| Videos uploaded / hour | Given | 50K |
| Images uploaded / hour | Given | 500K |
| Thumbnails per video | Given | 6 (5 candidates + sprite) |
| Total thumbnails / hour | Given | 1.8M |
| Thumbnails / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 500 |
| Avg thumbnail size | Given | 20 KB |
| Thumbnail storage / day | 1.8M/hr x 24 x 20 KB | 864 GB |
| CDN bandwidth | Given | ~100 Gbps |
Architecture Diagram
In the room: clarify pre-generation on upload vs on-demand processing, because queue depth multiplied by processing time drives worker count rather than peak read QPS.
Walk through the trigger, queue, worker, and CDN publish loop. Upload completion events are the only entry point, while clients poll or receive webhooks as workers scale horizontally based on queue depth.
Content uploads trigger thumbnail jobs that run on a scaled worker pool. Finished assets are served from a CDN while clients poll or use webhooks for readiness.
Component Deep Dives
Worker Queue and Scaling Pipeline
Start with queue architecture and horizontal scaling to address throughput targets. ML frame scoring serves as a senior add-on that should not block the core worker pool explanation.
Throughput comes from horizontal workers rather than faster FFmpeg execution. Video and image jobs use separate queues so 500K image thumbnails per hour never queue behind 50K video extraction tasks.
Upload triggers an S3 event that publishes a message to Kafka with the content identifier and processing options. The worker pool auto-scales on queue depth to meet the target of processing within 30 seconds at p99.
Queue sizing: 500 thumbs/sec x 2 sec avg FFmpeg job = ~1,000 concurrent workers GPU workers for ML scoring: 50 nodes x 20 parallel = 1,000 inferences/sec Priority queue: user-facing uploads > batch backfill Worker lifecycle: 1. Pull job from queue (visibility timeout = 2x expected duration) 2. Download source from S3 to local /tmp (streaming for large videos) 3. FFmpeg extract frames → ML score → select best → resize → upload to S3 4. Write metadata to MySQL, invalidate Redis cache 5. ACK message; on failure: retry 3x → DLQ with alert Poison messages: corrupt video → skip segment, use adjacent frame; unrecoverable → mark failed, serve placeholder, notify uploader
Video Thumbnail Selection: Finding the Best Frame
Multi-criteria scoring: extract N candidate frames (uniform sampling + scene changes), score each on sharpness (25%), brightness (15%), contrast (10%), face presence (30%), aesthetic quality (20%). Select top 3-5 candidates.
Sprite Sheet Generation
Extract 1 frame per 5 seconds, resize to 160x90, arrange in grid (10 cols x 12 rows). Generate VTT metadata file for client-side seek preview. Single HTTP request vs 120 individual requests.
Animated Thumbnail (Hover Preview)
Select 3 interesting segments (2 seconds each), extract at 10fps, combine into animated WebP (~200-400 KB). Only generate for top 10% most-viewed videos.
A/B Testing Thumbnails
Generate 3 candidates. Consistent variant assignment via user hash. Measure CTR + watch time (avoid clickbait). Auto-promote winner when statistically significant (Chi-squared test, >10K impressions per variant).
Event Bus Design (Kafka)
Topic: thumbnail-requests-video
Partitions: 64 (CPU-heavy jobs, roughly 8s each)
Partition key: content_id
Retention: 3 days
Topic: thumbnail-requests-image
Partitions: 256 (lightweight, roughly 200ms each, 10:1 volume ratio vs video)
Partition key: content_id
Producer: upload service on S3 complete
Event: { content_id, s3_key, content_type, options: { count: 5, formats: ["webp"] } }
Consumer groups:
1. video-thumb-workers: FFmpeg frame extract + ML scene detection to yield 5 candidates
2. image-thumb-workers: libvips resize to generate WebP ladder
3. doc-thumb-workers: PDF/PPTX first-page render
4. metadata-writer: UPDATE thumbnails SET urls[], selected_index in MySQL + Redis
Sync path: upload ACK < 200ms, while thumbnail URLs appear asynchronously within 30s p99
Separate topics prevent image starvation behind video backlog
DLQ: thumbnail-requests-*-dlq, with auto-scaling triggered when consumer lag exceeds 60sAPI Design
Generate Thumbnails for Video
POST /api/v1/thumbnails/video
{
"video_id": "vid-uuid",
"s3_key": "originals/vid-uuid/video.mp4",
"options": {
"sizes": ["150x150", "300x200", "640x360"],
"candidates": 5,
"sprite_sheet": true,
"animated_preview": true,
"format": "webp"
}
}Get Thumbnails
GET /api/v1/thumbnails/{content_id}
Response: 200 OK
{
"content_id": "vid-uuid",
"thumbnails": {
"default": "https://cdn.example.com/thumbs/vid-uuid/default_640x360.webp",
"small": "https://cdn.example.com/thumbs/vid-uuid/small_150x150.webp"
},
"candidates": [
{"index": 0, "url": "https://cdn.example.com/thumbs/vid-uuid/candidate_0.webp", "score": 0.92}
],
"sprite_sheet": {
"url": "https://cdn.example.com/thumbs/vid-uuid/sprite.jpg",
"vtt_url": "https://cdn.example.com/thumbs/vid-uuid/sprite.vtt",
"columns": 10, "rows": 12
}
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
MySQL: Thumbnail Metadata
CREATE TABLE thumbnails (
thumbnail_id BIGINT PRIMARY KEY AUTO_INCREMENT,
content_id VARCHAR(36) NOT NULL,
content_type ENUM('video', 'image', 'document') NOT NULL,
variant_type ENUM('default', 'candidate', 'custom', 'sprite', 'animated') NOT NULL,
s3_key TEXT NOT NULL,
cdn_url TEXT NOT NULL,
format VARCHAR(10) DEFAULT 'webp',
width INT,
height INT,
file_size_bytes INT,
quality_score DECIMAL(4,3),
is_default BOOLEAN DEFAULT FALSE,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
INDEX idx_content (content_id, variant_type)
);S3 + Redis
S3 Bucket: thumbnails
/{content_id}/default_640x360.webp
/{content_id}/sprite_sheet.jpg
/{content_id}/animated_preview.webp
Redis: thumb:{content_id} → CDN URL (TTL: 86400)
thumb_ab:{content_id}:{user_hash} → variant_index (TTL: 7d)Fault Tolerance
| Concern | Solution |
|---|---|
| FFmpeg crash | Retry 3x with exponential backoff; DLQ |
| Corrupt video frame |
|
| ML model failure | Fall back to rule-based scoring |
| Missing thumbnail |
|
| Worker pool exhaustion | Auto-scale based on Kafka consumer lag |
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50 and arch75 depth unless interviewing for a staff-level role.
- Async job queue: upload triggers worker, while client polls or uses webhooks (5 min)
- FFmpeg input seeking at keyframes for rapid frame extraction (6 min)
- ML scoring to select the best representative frame (5 min)
- Perceptual hashing to deduplicate near-identical thumbnails (5 min)
- Content-addressable storage keyed by hash bytes (4 min)
- Position thumbnails as a latency-sensitive asynchronous job triggered on video upload, because the player needs a poster frame before full transcoding finishes.
- Walk through candidate extraction using FFmpeg input seeking at strategic timestamps, such as skipping the intro, sampling the midpoint, and capturing action peaks.
- Explain ML scoring to pick the best frame based on brightness, face detection, and motion blur, with a rule-based fallback if the scoring model is unavailable.
- Cover perceptual hashing to deduplicate near-identical candidates so the picker returns visually diverse options.
- Mention content-addressable storage by hashing image bytes into an immutable CDN URL with
Cache-Control: max-age=31536000, immutablefor near-100% cache hit rates. - Discuss sprite sheet generation as a batch follow-up where a single decode pass replaces seeking repeatedly across timestamps.
- Avoid the common pitfall of FFmpeg output seeking (which decodes from the beginning for every candidate), turning a 2-hour video extraction from seconds into minutes.
Content-Addressable Thumbnails
Hash thumbnail content: sha256(thumbnail_bytes) into an S3 key. Duplicate uploads map to the same hash with zero extra storage. An immutable URL with Cache-Control: max-age=31536000, immutable ensures near 100% CDN cache hit rates.
Perceptual Hashing for Dedup
Compute pHash for each candidate. Skip if Hamming distance < 5 bits from selected candidates. Ensures diverse, non-redundant thumbnail candidates.
Engineering Trade-offs
FFmpeg Seek: Input vs Output Seeking
Thumbnail systems trade seek accuracy against worker throughput. Choosing between input seeking, output seeking, and ML scoring depth represents key architectural forks.
Input seeking (-ss before -i): very fast (~50 ms), though it may land on keyframes off by up to 2 seconds. This approach is ideal for thumbnail generation.
Output seeking (-ss after -i): frame-exact but slow because it decodes from the start. Use input seeking for candidate thumbnails, and execute a single continuous decode pass for sprite sheets.
CDN Format Negotiation
URL-based format selection is strongly recommended over Accept header negotiation. The client requests the format-specific URL directly, allowing the CDN to cache each URL independently for near-100% cache hit rates without suffering Vary: Accept cache fragmentation.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.