System Design Problem

Design a Thumbnail Generation Service

Commonly Asked By:PinterestNetflixDropboxAWS

Interview Setup

Interview Prompt

Design a thumbnail generation service: on video/image upload, produce default thumbs, candidate frames, sprite sheets, and animated previews for CDN-backed feeds.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Video vs image upload ratio?
  • 50K videos/hr need FFmpeg extract
  • 500K images/hr need resize only.
Who selects default thumbnail?ML auto-pick vs creator manual changes UX and APIs.
Sprite sheet for seek scrubbing?Extra FFmpeg pass + VTT per video.
Re-processing on same content_id?Idempotent S3 keys prevent duplicate CDN objects.

Scope

In scope

  • Async job queue
  • Video frame extraction
  • Image resize/WebP
  • Sprite+VTT
  • CDN cache
  • Idempotent generation

Out of scope (state explicitly)

Functional Requirements

Start by asking your interviewer whether thumbnails are pre-generated on upload or produced on-demand, as that single architectural decision drives queue sizing and CDN strategy. Thumbnail processing works in close coordination with the Video Transcoding Pipeline and the Image Processing Pipeline to deliver fast asset previews.

  • Auto-generate thumbnails: From videos, images, PDFs, documents
  • Multiple sizes: Small 150x150, medium 300x200, large 640x360
  • Video thumbnails: Extract the "best" frame using ML scoring
  • Sprite sheets: Grid of thumbnails for video seek preview
  • Custom thumbnails: Upload or select from candidates
  • A/B test thumbnails: Serve variants, measure CTR
  • Animated thumbnails: Short WebP preview on hover
  • Format optimization: Serve WebP/AVIF with JPEG fallback

Non-Functional Requirements

Thumbnails are immutable once generated, so lean into CDN cacheability as your primary latency strategy. The 30-second generation SLA is a write-path concern, whereas reads should be <50ms from edge cache with content-addressed URLs.

  • Speed: Thumbnails ready within 30 seconds of upload
  • Quality: Visually appealing, representative frames
  • Scale: Process 50K+ videos/hour and 500K+ images/hour
  • Cacheability: Highly cacheable on CDN (immutable URLs)
  • Cost Efficient: Avoid unnecessary regeneration
  • Availability: 99.9%

Capacity Estimations

Worker concurrency equals throughput multiplied by average job duration. Video jobs run 10 to 100 times longer than image resizes, so quoting a single worker count without separating queues will prompt pushback from your interviewer.

MetricCalculationValue
Videos uploaded / hourGiven50K
Images uploaded / hourGiven500K
Thumbnails per videoGiven6 (5 candidates + sprite)
Total thumbnails / hourGiven1.8M
Thumbnails / secDerived from daily volume ÷ 86400 (+ peak factor)500
Avg thumbnail sizeGiven20 KB
Thumbnail storage / day1.8M/hr x 24 x 20 KB864 GB
CDN bandwidthGiven~100 Gbps

Architecture Diagram

In the room: clarify pre-generation on upload vs on-demand processing, because queue depth multiplied by processing time drives worker count rather than peak read QPS.

Walk through the trigger, queue, worker, and CDN publish loop. Upload completion events are the only entry point, while clients poll or receive webhooks as workers scale horizontally based on queue depth.

Content uploads trigger thumbnail jobs that run on a scaled worker pool. Finished assets are served from a CDN while clients poll or use webhooks for readiness.

Loading...

Component Deep Dives

Worker Queue and Scaling Pipeline

Start with queue architecture and horizontal scaling to address throughput targets. ML frame scoring serves as a senior add-on that should not block the core worker pool explanation.

Throughput comes from horizontal workers rather than faster FFmpeg execution. Video and image jobs use separate queues so 500K image thumbnails per hour never queue behind 50K video extraction tasks.

Upload triggers an S3 event that publishes a message to Kafka with the content identifier and processing options. The worker pool auto-scales on queue depth to meet the target of processing within 30 seconds at p99.

Queue sizing:
  500 thumbs/sec x 2 sec avg FFmpeg job = ~1,000 concurrent workers
  GPU workers for ML scoring: 50 nodes x 20 parallel = 1,000 inferences/sec
  Priority queue: user-facing uploads > batch backfill

Worker lifecycle:
  1. Pull job from queue (visibility timeout = 2x expected duration)
  2. Download source from S3 to local /tmp (streaming for large videos)
  3. FFmpeg extract frames → ML score → select best → resize → upload to S3
  4. Write metadata to MySQL, invalidate Redis cache
  5. ACK message; on failure: retry 3x → DLQ with alert

Poison messages: corrupt video → skip segment, use adjacent frame;
  unrecoverable → mark failed, serve placeholder, notify uploader

Video Thumbnail Selection: Finding the Best Frame

Multi-criteria scoring: extract N candidate frames (uniform sampling + scene changes), score each on sharpness (25%), brightness (15%), contrast (10%), face presence (30%), aesthetic quality (20%). Select top 3-5 candidates.

Sprite Sheet Generation

Extract 1 frame per 5 seconds, resize to 160x90, arrange in grid (10 cols x 12 rows). Generate VTT metadata file for client-side seek preview. Single HTTP request vs 120 individual requests.

Animated Thumbnail (Hover Preview)

Select 3 interesting segments (2 seconds each), extract at 10fps, combine into animated WebP (~200-400 KB). Only generate for top 10% most-viewed videos.

A/B Testing Thumbnails

Generate 3 candidates. Consistent variant assignment via user hash. Measure CTR + watch time (avoid clickbait). Auto-promote winner when statistically significant (Chi-squared test, >10K impressions per variant).

Event Bus Design (Kafka)

Topic: thumbnail-requests-video
  Partitions: 64 (CPU-heavy jobs, roughly 8s each)
  Partition key: content_id
  Retention: 3 days

Topic: thumbnail-requests-image
  Partitions: 256 (lightweight, roughly 200ms each, 10:1 volume ratio vs video)
  Partition key: content_id

Producer: upload service on S3 complete
  Event: { content_id, s3_key, content_type, options: { count: 5, formats: ["webp"] } }

Consumer groups:
  1. video-thumb-workers: FFmpeg frame extract + ML scene detection to yield 5 candidates
  2. image-thumb-workers: libvips resize to generate WebP ladder
  3. doc-thumb-workers: PDF/PPTX first-page render
  4. metadata-writer: UPDATE thumbnails SET urls[], selected_index in MySQL + Redis

Sync path: upload ACK < 200ms, while thumbnail URLs appear asynchronously within 30s p99
Separate topics prevent image starvation behind video backlog
DLQ: thumbnail-requests-*-dlq, with auto-scaling triggered when consumer lag exceeds 60s

API Design

Generate Thumbnails for Video

HTTP
POST /api/v1/thumbnails/video
{
  "video_id": "vid-uuid",
  "s3_key": "originals/vid-uuid/video.mp4",
  "options": {
    "sizes": ["150x150", "300x200", "640x360"],
    "candidates": 5,
    "sprite_sheet": true,
    "animated_preview": true,
    "format": "webp"
  }
}

Get Thumbnails

HTTP
GET /api/v1/thumbnails/{content_id}
Response: 200 OK
{
  "content_id": "vid-uuid",
  "thumbnails": {
    "default": "https://cdn.example.com/thumbs/vid-uuid/default_640x360.webp",
    "small": "https://cdn.example.com/thumbs/vid-uuid/small_150x150.webp"
  },
  "candidates": [
    {"index": 0, "url": "https://cdn.example.com/thumbs/vid-uuid/candidate_0.webp", "score": 0.92}
  ],
  "sprite_sheet": {
    "url": "https://cdn.example.com/thumbs/vid-uuid/sprite.jpg",
    "vtt_url": "https://cdn.example.com/thumbs/vid-uuid/sprite.vtt",
    "columns": 10, "rows": 12
  }
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpoint

Data Model

MySQL: Thumbnail Metadata

SQL
CREATE TABLE thumbnails (
    thumbnail_id    BIGINT PRIMARY KEY AUTO_INCREMENT,
    content_id      VARCHAR(36) NOT NULL,
    content_type    ENUM('video', 'image', 'document') NOT NULL,
    variant_type    ENUM('default', 'candidate', 'custom', 'sprite', 'animated') NOT NULL,
    s3_key          TEXT NOT NULL,
    cdn_url         TEXT NOT NULL,
    format          VARCHAR(10) DEFAULT 'webp',
    width           INT,
    height          INT,
    file_size_bytes INT,
    quality_score   DECIMAL(4,3),
    is_default      BOOLEAN DEFAULT FALSE,
    created_at      TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    INDEX idx_content (content_id, variant_type)
);

S3 + Redis

S3 Bucket: thumbnails
  /{content_id}/default_640x360.webp
  /{content_id}/sprite_sheet.jpg
  /{content_id}/animated_preview.webp

Redis: thumb:{content_id} → CDN URL (TTL: 86400)
       thumb_ab:{content_id}:{user_hash} → variant_index (TTL: 7d)

Fault Tolerance

ConcernSolution
FFmpeg crashRetry 3x with exponential backoff; DLQ
Corrupt video frame
  • Skip corrupt segment
  • use adjacent frame
ML model failureFall back to rule-based scoring
Missing thumbnail
  • CDN serves placeholder
  • queue regeneration
Worker pool exhaustionAuto-scale based on Kafka consumer lag

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50 and arch75 depth unless interviewing for a staff-level role.

    • Async job queue: upload triggers worker, while client polls or uses webhooks (5 min)
    • FFmpeg input seeking at keyframes for rapid frame extraction (6 min)
    • ML scoring to select the best representative frame (5 min)
    • Perceptual hashing to deduplicate near-identical thumbnails (5 min)
    • Content-addressable storage keyed by hash bytes (4 min)
  • Position thumbnails as a latency-sensitive asynchronous job triggered on video upload, because the player needs a poster frame before full transcoding finishes.
  • Walk through candidate extraction using FFmpeg input seeking at strategic timestamps, such as skipping the intro, sampling the midpoint, and capturing action peaks.
  • Explain ML scoring to pick the best frame based on brightness, face detection, and motion blur, with a rule-based fallback if the scoring model is unavailable.
  • Cover perceptual hashing to deduplicate near-identical candidates so the picker returns visually diverse options.
  • Mention content-addressable storage by hashing image bytes into an immutable CDN URL with Cache-Control: max-age=31536000, immutable for near-100% cache hit rates.
  • Discuss sprite sheet generation as a batch follow-up where a single decode pass replaces seeking repeatedly across timestamps.
  • Avoid the common pitfall of FFmpeg output seeking (which decodes from the beginning for every candidate), turning a 2-hour video extraction from seconds into minutes.

Content-Addressable Thumbnails

Hash thumbnail content: sha256(thumbnail_bytes) into an S3 key. Duplicate uploads map to the same hash with zero extra storage. An immutable URL with Cache-Control: max-age=31536000, immutable ensures near 100% CDN cache hit rates.

Perceptual Hashing for Dedup

Compute pHash for each candidate. Skip if Hamming distance < 5 bits from selected candidates. Ensures diverse, non-redundant thumbnail candidates.

Engineering Trade-offs

FFmpeg Seek: Input vs Output Seeking

Thumbnail systems trade seek accuracy against worker throughput. Choosing between input seeking, output seeking, and ML scoring depth represents key architectural forks.

Input seeking (-ss before -i): very fast (~50 ms), though it may land on keyframes off by up to 2 seconds. This approach is ideal for thumbnail generation.

Output seeking (-ss after -i): frame-exact but slow because it decodes from the start. Use input seeking for candidate thumbnails, and execute a single continuous decode pass for sprite sheets.

CDN Format Negotiation

URL-based format selection is strongly recommended over Accept header negotiation. The client requests the format-specific URL directly, allowing the CDN to cache each URL independently for near-100% cache hit rates without suffering Vary: Accept cache fragmentation.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...