Interview Setup
Interview Prompt
Design an image processing pipeline like Pinterest's backend: upload originals, generate variants, serve through CDN, support on-demand resize/format/crop.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Upload throughput vs on-demand transform QPS? | 1,200 uploads/sec vs 50K transforms/sec are different scaling problems. |
| Pre-generate variants or only on-demand? | 5 variants x 100M uploads = 450 TB/day before any custom sizes. |
| Is NSFW moderation required? | Public CDN URLs stay gated until ML/human approval. |
| Smart crop for thumbnails? | Face detection metadata powers crop=smart on read path. |
Scope
In scope
- Presigned upload
- Async variant generation
- On-demand transforms
- CDN
- NSFW moderation
- Webhooks
Out of scope (state explicitly)
- Feed ranking (see Video Recommendation Engine)
- Social graph
- In-app filters
Functional Requirements
Start by asking your interviewer whether on-the-fly URL transforms (resize on cache miss) are in scope or only pre-generated variants on upload. That split determines whether you need a transform worker pool alongside the async upload pipeline, which most interviewers expect.
- Upload images: Accept user image uploads via presigned S3 URLs (JPG, PNG, WebP, HEIC)
- Generate variants: Create standard sizes (thumbnail, small, medium, large, original)
- Format conversion: Convert to modern formats (WebP, AVIF) for bandwidth savings
- On-demand transformation: Dynamic resize/crop via URL parameters with CDN caching
- Content moderation: Automated NSFW detection before making images public
- Smart cropping: Face detection and saliency-aware cropping for thumbnails
- EXIF stripping: Remove sensitive metadata (GPS location, camera serial) for privacy
Non-Functional Requirements
Emphasize the read/write asymmetry: uploads are bursty and async-tolerant, but CDN cache misses on transforms need sub-200ms responses. Original durability (11 nines) is non-negotiable, because losing a source image is far worse than temporary variant generation delays.
- High Throughput: Handle 100M image uploads per day (~1,200/sec avg, 3,000/sec peak)
- Low Latency: On-demand transform < 200ms p99; upload confirmation < 200ms
- Durability: 11 nines (originals must NEVER be lost); variants can be re-generated
- Availability: 99.99% for image serving via CDN
- Cost Efficiency: Minimize storage (WebP/AVIF save 30-50%) and compute (cache aggressively)
- Scalability: Handle viral spikes with request coalescing (single-flight)
Capacity Estimations
Size your worker pool from upload rate multiplied by variants per image and processing time, rather than peak read QPS. The table below separates ingest volume from transform-on-miss load, because both drive different scaling knobs.
| Metric | Calculation | Value |
|---|---|---|
| Images uploaded / day | Given | 100M |
| Uploads / sec | 100M ÷ 86400 | ~1,200 |
| Avg original image size | Given | 3 MB |
| Upload storage / day | 100M x 3 MB | 300 TB |
| Variants per image | Given | 5 |
| Variant storage / day | 300 TB x ~1.5 (5 variants x ~30% of original) | ~450 TB |
| On-demand transforms / sec | Given (read-path cache misses) | 50K |
| CDN bandwidth | Given | ~500 Gbps |
Architecture Diagram
In the room: say presigned S3 upload before drawing workers, and never route 3 MB image bytes through API servers at 1,200 requests per second.
Draw the upload path and read path separately. Uploads land in object storage and fan out to async workers; reads hit immutable CDN URLs first, with transform workers only on cache miss. In the interview, state this out loud before drawing boxes to prevent the common mistake of placing image transformations on the synchronous upload response.
Uploads go straight to object storage; workers generate variants asynchronously while the read path serves immutable CDN URLs, with transform workers only on cache miss.
Component Deep Dives
We cover the async worker queue first because that is where throughput lives, then the hybrid pre-process vs on-demand strategy that keeps CDN hit rates high without pre-generating every possible size.
Writes are async and idempotent; reads are CDN-first. Moderation and smart crop are post-upload concerns that never block the presigned PUT.
Async Processing Queue: Kafka + Worker Pool
S3 upload completion event → Kafka topic image-processing (partition by image_id for ordering). Workers pull jobs, process, ACK on success.
Scale: 1,200 images/sec upload x 5 variants x 200ms avg = ~1,200 workers (I/O bound) Moderation ML adds GPU pool: 200K inferences/sec ÷ 4 batch = 50 GPUs Backpressure: Consumer lag > 60s → scale workers (K8s HPA on lag metric) Lag > 5 min → shed low-priority reprocessing, alert on-call Partial failure: Variant 3/5 fails → mark partial, retry failed variants only Idempotency key = image_id + variant_spec → safe to retry DLQ: after 3 retries → human review queue for corrupt uploads
Pre-Processing vs On-Demand: Hybrid Architecture
Pre-process core variants on upload: thumbnail 150x150, feed 1080x1080, profile 640x640 (cover 90% of requests). On-demand for the long tail (10%): unusual sizes, formats, crops processed on first request and cached. 70% less compute than pre-processing all variants.
Image Format Pipeline
WebP (30% smaller than JPEG, 96% browser support), AVIF (50% smaller, growing support), JPEG (universal fallback). Adaptive quality: complex images get quality 85, simple images get quality 70. Content negotiation at CDN via Accept header or URL-based format selection.
Smart Cropping
Face detection (OpenCV/MTCNN, ~50ms) leads to saliency or attention crop fallback, and ultimately the rule of thirds. Pre-compute crop coordinates per variant on upload and apply them at serving time. Lightweight face detection at 50K images/sec requires GPU acceleration.
Content Moderation ML Pipeline
Every image passes through: NSFW detection (ResNet/EfficientNet), violence detection, OCR + spam filter. At 50K images/sec: ~200K inferences/sec. Optimized with batch inference, model distillation (MobileNet for initial screening), TensorRT optimization. ~25 GPUs needed at scale.
Event Bus Design (Kafka)
Topic: image-processing
Partitions: 128 (scale worker pool horizontally)
Partition key: image_id (preserves per-image processing order)
Retention: 7 days (replay failed jobs)
Replication factor: 3, min.insync.replicas: 2
Producer: S3 upload-complete event (idempotent producer)
Event: { image_id, s3_key, user_id, content_type, upload_timestamp }
Consumer groups:
1. image-workers: validate → format convert → resize variants → moderation → S3 /processed
2. metadata-writer: UPDATE images SET status=ready, variants[] in MySQL
3. cdn-warmer: pre-warm CDN edge cache for thumb/sm/md/lg variants
Sync path: presigned PUT ACK < 200ms; return image_id immediately
Async path: 5 pre-gen variants + CDN redirect on read; workers idempotent by image_id
DLQ: image-processing-dlq after 3 retries; alert when consumer lag > 120sAPI Design
Upload Image
POST /api/v1/images/upload
{
"filename": "vacation.jpg",
"content_type": "image/jpeg"
}
Response: 200 OK
{
"image_id": "img-uuid",
"upload_url": "https://s3.amazonaws.com/originals/img-uuid?X-Amz-..."
}Get Processed Image
GET /api/v1/images/{image_id}?w=720&h=480&format=webp&crop=smart&quality=80
Response: 302 Redirect → CDN URL
Location: https://cdn.example.com/images/img-uuid/720x480_smart_q80.webpCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
202 Accepted: asynchronous job queued successfully, poll GET /jobs/{id} for completion status
408 Request Timeout: background job is still executing, continue polling status endpointData Model
MySQL: Image Metadata
CREATE TABLE images (
image_id VARCHAR(36) PRIMARY KEY,
user_id VARCHAR(36) NOT NULL,
original_s3_key TEXT NOT NULL,
original_format VARCHAR(10),
original_width INT,
original_height INT,
file_size_bytes INT,
exif_data JSON,
dominant_colors JSON,
faces_detected SMALLINT DEFAULT 0,
nsfw_score DECIMAL(4,3),
moderation_status ENUM('pending','approved','rejected','review') DEFAULT 'pending',
processing_status ENUM('uploaded','processing','completed','failed') DEFAULT 'uploaded',
smart_crop_data JSON,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
INDEX idx_user (user_id, created_at DESC)
);S3 Storage Layout
Bucket: image-originals (cross-region, never deleted)
/{image_id}/original.jpg
Bucket: image-processed (CDN-served)
/{image_id}/thumbnail_150x150.webp
/{image_id}/medium_640x640.webp
/{image_id}/large_1080x1080.webpFault Tolerance
| Concern | Solution |
|---|---|
| Original image lost |
|
| Processing worker crash |
|
| Corrupt image upload | Validate header + decode test before processing |
| ML moderation error | Human review queue for borderline scores |
| Decompression bomb |
|
| EXIF GPS leak |
|
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Presigned S3 upload and Kafka async pipeline (5 min)
- Worker path: validate, libvips resize, and 5 WebP variants (6 min)
- CDN redirect on read; on-demand transform on cache miss (5 min)
- ML moderation parallel step: never block resize path (5 min)
- Security: strip EXIF GPS and reject decompression bombs (4 min)
- Frame upload as fire-and-forget: store original to S3, publish a Kafka event, and return immediately because processing is asynchronous.
- Walk through the worker pipeline: validate header dimensions → decode with libvips (streaming, low memory) → generate WebP/JPEG variants → upload to CDN origin.
- Explain responsive serving via URL convention (
/images/{id}/w_{width}.webp) with long-lived CDN cache headers. - Cover ML moderation as a parallel step: scores above threshold go to human review, borderline cases never block the resize path.
- Mention security hardening: strip EXIF GPS from all public variants, reject decompression bombs before full decode.
- Discuss when to self-host (imgproxy + CDN) vs Cloudinary based on daily image volume and cost crossover.
- Common pitfall: using ImageMagick with full in-memory decode, where a 50 MB PNG ballooning to 2 GB RAM crashes the worker pool.
- Related systems: for dedicated preview assets generation, see Thumbnail Generation Service, and for content moderation workflows, see Content Moderation System.
Image Bomb Detection
Read the header to inspect dimensions before a full decode. Reject if width x height x 4 > 1 GB. Set strict resource limits. Use libvips (streaming, which does not load the full image in memory). Timeout after 30 seconds.
Responsive Image Serving
HTML srcset with multiple widths (300w, 600w, 1200w). Client Hints for automatic selection. URL convention: /images/{id}/w_{width}.{format}.
Engineering Trade-offs
Image pipelines trade memory, CPU, and storage cost; choosing libvips vs ImageMagick and pre-generation vs on-demand are the primary pivots.
ImageMagick vs libvips vs Pillow
libvips ⭐: 10x faster than ImageMagick, 10x less memory (20 MB RAM for resize). Streaming architecture. Recommended for production.
Sharp (Node.js): good for web backends.
Pillow: simple but single-threaded (GIL).
Self-Hosted vs SaaS (Cloudinary)
Under 1M images/day: Cloudinary (simpler). 1M to 100M: imgproxy + CDN (cost-effective). Over 100M: custom pipeline + imgproxy for the long tail. imgproxy signed URLs prevent abuse (HMAC).
Storage Cost Optimization
Format conversion (JPEG to WebP saves 30%), perceptual quality targeting (quality 85 vs 95 saves 40%), storage tiering (S3 Standard to Infrequent Access to Glacier), deduplication via perceptual hash, and progressive deletion of regenerable variants.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.