System Design Problem

Design an Ad Click Prediction System

Commonly Asked By:GoogleMetaTradeDeskByteDance

Interview Setup

Interview Prompt

Design an ad click prediction system serving 50M CTR predictions/sec with < 7ms total latency (2ms feature lookup + 5ms inference), using 100-200 features per prediction from a 2 TB online feature store.

Clarifying Questions (ask before designing)

QuestionWhy it matters
GBDT (LightGBM) or deep model (DCN) for serving?
  • LightGBM < 1ms inference
  • DCN ~5ms but learns feature crosses. Two-stage common.
Feature store: precomputed or computed at serving time?
  • 50M predictions/sec x 200 features = 10B feature lookups/sec
  • must be precomputed in Redis.
Model retraining frequency: daily batch or online learning?
  • 10 TB training data/day
  • daily retrain captures trends
  • online learning for hour-level adaptation.
Calibration required for bid optimization?Uncalibrated CTR → overbid/underbid; ECE < 0.01 required for RTB integration.

Scope

In scope

  • Feature engineering at serving time
  • Model serving latency
  • Click-through rate prediction
  • Bid optimization
  • Feedback loop
  • Online learning

Out of scope (state explicitly)

  • GPU cluster training and hyperparameter tuning
  • Content moderation of recommended items
  • Ad auction / sponsored placement ranking

Functional Requirements

Start by asking your interviewer for the latency budget inside the RTB auction (typically under 10ms total) and the prediction QPS at peak. Clarify whether calibrated probabilities are required for bid optimization and whether you own the feature store or consume an existing one.

  • Predict probability a user will click an ad (CTR prediction) in < 10ms
  • Feature engineering from user profile, ad creative, context (page, time, device)
  • Model training pipeline: daily retraining on latest click data
  • Online learning: model adapts to recent patterns within hours
  • A/B testing: compare model versions on live traffic
  • Feature store: consistent features between training and serving
  • Calibration: predicted probabilities must match actual click rates

Non-Functional Requirements

Your interviewer will stress-test the latency budget split (~2ms feature lookup, ~5ms inference, ~3ms overhead) and whether miscalibrated probabilities cause overbidding. They will also probe training-serving skew and how click feedback closes the loop within 24 hours.

  • Ultra-Low Latency: < 10ms p99 for prediction
  • High Throughput: 1M+ predictions/sec
  • Freshness: Model reflects behavior from last 24h
  • Accuracy: 0.1% CTR improvement = millions in revenue

Capacity Estimations

At 50M+ predictions/sec, Redis feature-store memory and inference fleet size dominate. Daily training on 10 TB of click logs justifies the offline/online split: never block the auction path on a Spark job.

MetricCalculationValue
Predictions / secDerived from daily volume ÷ 86400 (+ peak factor)50M
Feature lookup latency budgetGiven< 2ms
Model inference latency budgetGiven< 5ms
Features per predictionGiven100 to 200
Training data / day~10 TB ÷ 86400~10 TB
Model sizeGiven100 MB to 2 GB
Feature store (online)Given~2 TB (Redis cluster)
Training data (offline)Given~100 TB (S3/Parquet)

Architecture Diagram

Draw two paths on the whiteboard: synchronous predict and async feedback. The auction path must complete in under 7ms (spanning feature store lookup, LightGBM inference, and probability calibration) because the exchange will not wait for a batch Spark job.

Everything else is eventually consistent: impressions and clicks flow to Kafka, Flink updates real-time features, and a daily batch retrains on 10 TB of click logs. Model deploy follows canary with automatic rollback on AUC or ECE regression.

At 50M predictions/sec, the feature store is the bottleneck before GPU inference; use batch MGET per user and ad, alongside local caching on inference servers to absorb hot-key traffic.

Loading...

In the room

Ask whether you need calibrated probabilities for bidding: if yes, plan for isotonic regression and ECE monitoring from the start, not as an afterthought.

Component Deep Dives

Staff-level depth here is calibration and position bias, because models that rank well offline but overbid in production fail silently. Walk model choice, then calibration, then bias correction in that order.

Model Choice: Fast Trees vs Deep Architectures

Model selection establishes the hard latency ceiling. Comparing LightGBM and Deep & Cross Networks demonstrates why a two-stage funnel is standard practice at auction scale.

LightGBM (gradient boosted trees):
  ✅ Fast inference (< 1ms), interpretable, handles sparse features
  ❌ Can't learn complex interactions automatically

Deep & Cross Network (DCN):
  ✅ Automatically learns feature interactions (crossing layers)
  ❌ Slower inference (~5ms), needs GPU for serving

Practice: Two-stage
  Stage 1: LightGBM for candidate scoring (fast, high recall)
  Stage 2: DNN for final ranking (slow but accurate, only top 50 candidates)

Probability Calibration for Bidding

Calibration is a critical evaluation topic: a model that ranks candidates accurately offline but overestimates CTR by even 67% will systematically overbid and exhaust advertiser budgets.

Why: If model says P(click)=0.05 but actual CTR is 0.03 -> overbid by 67%

How: Isotonic regression or Platt scaling maps raw scores to calibrated probabilities

Monitoring: Expected calibration error (ECE) < 0.01
  Bucket predictions into deciles -> compare predicted vs actual CTR

Position Bias Correction

Position bias is a pervasive confounding variable in click logs: ads displayed in slot 1 receive clicks regardless of relevance, so training and serving pipelines must decouple position signals.

Problem: Ads in position 1 get clicked more regardless of relevance

Solution: Train on "position-aware" features, but serve WITHOUT position
  Training: include position as feature -> model learns position bias
  Serving: set position=1 for all -> model predicts "click if shown in position 1"

Alternative: IPW (Inverse Propensity Weighting)
  Weight each sample by 1/P(position), reducing position bias

Event Bus Design (Kafka)

Kafka closes the feedback loop between serving and learning: impression events record model predictions, while click events provide attribution labels over a 24-hour window.

Topic: ad-impressions
  Partitions: 256 (partition by user_id)
  Retention: 7 days (feature freshness + attribution window)
  Producers: Ad server at serve time (features used + model version logged)
  Consumers: Flink real-time feature updater, training label joiner

Topic: click-events
  Partitions: 256 (partition by user_id)
  Events: click, conversion (with 24h attribution window)
  Consumers: daily Spark retrain pipeline, online weight updater, PSI drift monitor

Topic: prediction-logs
  Sampling: 0.1% of predictions for offline calibration checks
  Payload: {user_id, ad_id, predicted_ctr, model_version, features_hash}

Serve path: feature lookup (Redis < 2ms) → LightGBM inference (< 5ms) → log impression async
  Click feedback loop: click-events → feature store update → next-day model retrain on 10 TB

API Design

Walk the predict endpoint as sitting inside the RTB auction: batch multiple ad candidates in one call to amortize feature lookup, and keep click feedback fully async so it never blocks the response.

# Real-time prediction (called by ad exchange in auction path)
POST /api/predict
{
  "user_id": "u_abc123",
  "ad_candidates": ["ad_001", "ad_002", "ad_003"],
  "context": { "page_url": "...", "device": "mobile", "geo": "US" }
}
-> {
  "predictions": [
    {"ad_id": "ad_001", "p_click": 0.032, "calibrated": true},
    {"ad_id": "ad_002", "p_click": 0.018, "calibrated": true}
  ],
  "latency_ms": 8
}

# Model management
POST /api/models/deploy      -> Deploy new model version (canary)
GET  /api/models/active       -> Current model version + metrics
POST /api/models/rollback     -> Revert to previous model

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Redis (Online Feature Store)

HSET user:features:{user_id}
  avg_ctr_7d "0.032"
  click_count_30d "145"
  top_category "electronics"
  session_depth "3"
  last_click_hours "2.5"
  _updated_at "2026-03-14T10:05:00Z"
EXPIRE user:features:{user_id} 86400

HSET ad:features:{ad_id}
  historical_ctr "0.025"
  category "travel"
  creative_size "300x250"

Kafka Topics

Topic: ad-impressions (partition by user_id, RF=3)
  { "user_id": "u_abc", "ad_id": "ad_001", "position": 2, "p_click": 0.032, "model_version": "v42" }

Topic: click-events (partition by user_id, RF=3)
  { "user_id": "u_abc", "ad_id": "ad_001", "click_timestamp": "..." }

Flink joins impressions + click-events → training labels with 24h attribution window

Fault Tolerance

ConcernSolution
Model serving failureFall back to simpler model (logistic regression) or last-known-good
Feature store unavailable
  • Use default features (population medians)
  • reduces accuracy, not availability
Stale features
  • TTL enforcement
  • degrade gracefully with confidence reduction
Training failure
  • Don't deploy
  • keep serving current model
  • alert ML team
Canary deployment
  • New model serves 5% traffic
  • monitor AUC, calibration -> promote or rollback
Redis shard failure
  • Redis Cluster auto-failover
  • missing features for ~1% of users during failover

Model Redundancy

Two models always warm in memory on every serving host:
  Champion (current production model, v47)
  Challenger (candidate model, v48, or last-known-good v46)

Routing: config flag determines active model
  Normal: champion serves 100%
  Canary: champion 95%, challenger 5%
  Rollback: instant config flip -> challenger active, <60 seconds

Both models score EVERY request (dual scoring):
  Active model's score -> used for auction
  Inactive model's score -> logged for offline comparison

Additional Considerations

Redis Cluster Architecture

Scale: 1B users x 2 KB features = 2 TB online feature store

Redis Cluster: 32 primary shards + 32 replicas = 64 nodes
Sharding: CRC16(user_id) % 16384 -> slot -> shard

Read path (per prediction):
  1. Hash user_id -> HGETALL user:features (0.5ms)
  2. Hash ad_id -> HGETALL ad:features (0.5ms)
  3. Context features in-process (0.1ms)
  4. Assemble 200-dim feature vector (0.2ms)
  Total: < 2ms (parallel Redis calls)

Hot-key problem: viral user -> millions of impressions
  Solution: rate-limit feature updates per user
  If user_id seen > 100 times in last minute -> skip update

Race Conditions in Online Feature Updates

Race 1: Flink writes while prediction reads
  HSET and HGETALL are serialized (Redis single-threaded)
  For multi-field: use MULTI/EXEC for atomic updates

Race 2: Two Flink workers updating same user
  Solution: Partition Kafka by user_id
  All events for same user -> same partition -> same Flink subtask -> single writer

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Separate online inference (p(click) in <10ms) from… (5 min)
    • feature store for consistent online/offline features (6 min)
    • logistic regression or GBDT as baseline before jumping… (5 min)
    • exploration vs exploitation (multi-armed bandit) for… (5 min)
    • model versioning and shadow deployment before… (4 min)
  • Separate online inference (p(click) in <10ms) from offline training (batch on historical logs).
  • Explain feature store for consistent online/offline features, because training-serving skew degrades model quality.
  • Cover logistic regression or GBDT as baseline before jumping to deep learning, demonstrating pragmatic engineering choices.
  • Discuss exploration vs exploitation (multi-armed bandit) for new ads without click history.
  • Mention model versioning and shadow deployment before promoting a new model to production traffic.
  • Scope boundary: CTR prediction and calibration for bid input is in scope, whereas full RTB auction mechanics are handled separately in Real-Time Bidding System.
  • Common pitfall: training on click labels without correcting for position bias, where top slots receive clicks regardless of relevance.

Engineering Trade-offs

LightGBM vs Deep Neural Network

LightGBM:
  Inference: < 1ms (single CPU core)
  Training: 2 hours on 1B samples
  AUC: 0.785
  Best for: Candidate scoring (fast, high recall)

Deep & Cross Network (DCN-v2):
  Inference: 3-5ms (GPU-accelerated)
  Training: 12 hours on 1B samples (8 GPUs)
  AUC: 0.795 (+1% -> millions in revenue)
  Best for: Final ranking (accurate, only top 50)

Two-stage: LightGBM scores 200 candidates (1ms) -> DCN re-ranks top 50 (5ms) = 6ms total

Negative Downsampling

Problem: CTR ~1% -> 99% negative samples
  Model sees 99 negatives per positive -> predicts 0 for everything

Solution: Downsample negatives
  Keep ALL positives, keep 1-in-10 negatives (1:10 ratio)
  Calibration correction: p_calibrated = p_model x rate / (p_model x rate + (1-p_model))

Industry standard: 1:10 to 1:20 ratio

Online Learning vs Daily Batch Retraining

Daily Batch: stable but 24-48h behind
Online Learning: adapts within hours but risk of catastrophic forgetting

Hybrid (recommended):
  Base model: daily batch retrain (stable foundation)
  Delta model: hourly online updates (captures recent trends)
  Final score = 0.7 x base_score + 0.3 x delta_score
  Weekly: merge delta into base -> fresh start for delta

If delta model degrades -> coefficient automatically reduces to 0 -> graceful fallback

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...