System Design Problem

Design an Ecommerce Recommendation System

Commonly Asked By:AmazonGoogleNetflixByteDance

Interview Setup

Interview Prompt

Design a recommendation system for an ecommerce platform that suggests products to users. Support personalized recommendations for logged in users, reasonable defaults for new users (cold start), and balance exploitation (show what we know works) with exploration (discover new preferences). Handle 100M users and 10M products.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What surfaces need recommendations, such as the homepage, product pages, or promotional emails?The homepage requires broad category coverage and personalization, whereas a product page can emphasize similar products, accessories, and compatible items, which requires different candidate signals.
Are recommendations personalized in real time or updated periodically in batch?Session based recommendations respond to real time product views and clicks within seconds, whereas daily batch workflows compute collaborative filtering or retrieval embeddings offline with fundamentally different latency budgets.
What interaction signals are available, such as clicks, purchases, ratings, or dwell time?Purchase signals are sparse but high intent, while click events are dense but noisy, which directly determines the training loss formulation and positive label definitions.
How do we measure success, such as click through rate, conversion rate, revenue, or catalog diversity?Optimizing click through rate in isolation creates narrow filter bubbles, requiring multi objective ranking to balance engagement, revenue, and content diversity.

Scope

In scope

  • Two tower model for retrieval
  • ANN (Approximate Nearest Neighbor) candidate generation
  • Ranking model (LTR) with feature store
  • Catalog and inventory eligibility signals
  • Cold start for new users and new items
  • Explore/exploit balance (multi armed bandit)
  • Real time and batch feature pipeline

Out of scope (state explicitly)

  • Short video For You Page feeds and media upload or CDN distribution covered in TikTok (Short Video Platform)
  • Full model training pipeline (GPU cluster, hyperparameter tuning)
  • Content moderation of recommended items
  • Ad auction / sponsored recommendations

Functional Requirements

Recommendation systems solve a multi stage retrieval and ranking problem rather than scoring every catalog item on demand. When scoping the architecture with an interviewer, confirm the candidate generation boundaries, ranking criteria, cold start expectations, and explore-vs-exploit diversity requirements.

In an interview, introduce the multi stage funnel early, highlighting approximate nearest neighbor retrieval in roughly 10 milliseconds followed by a representative 40 millisecond ranking target over approximately 500 candidates. The separate 50 millisecond GPU benchmark covers a larger 5,000 item scoring batch, so the figures describe different capacity targets.

  • Personalized recommendations: Generate relevant item suggestions tailored to each user interaction history and preference profile.
  • Multiple surfaces: Deliver contextual recommendations across distinct product surfaces, including the homepage feed, product similarity modules, trending modules, category pages, and promotional surfaces.
  • Real time signals: Ingest and incorporate recent user actions such as views, likes, and skips into session level recommendation features within seconds, with slower aggregate features following their documented freshness tier.
  • Cold start handling: Provide engaging recommendations for newly registered users without historical data as well as newly published catalog items.
  • Catalog diversity: Prevent repetitive filter bubbles by ensuring recommended lists expose users to varied product categories, brands, and price bands.
  • Recommendation explainability: Present clear attribution cues alongside recommendations, such as "Similar to products you viewed" or "Popular in your area".

Non-Functional Requirements

The system must target under 100 milliseconds for p95 recommendation responses and under 200 milliseconds for p99 responses while filtering millions of catalog items down to dozens of high relevance items. The representative component budget is approximately 110 milliseconds before additional platform overhead, so the serving tier must rely on parallelism, caching, and efficient execution paths to meet the end to end SLO.

  • Low latency: Deliver online recommendation responses in under 100 milliseconds for p95 requests and under 200 milliseconds for p99 requests.
  • High throughput: Sustain peak loads exceeding 100,000 recommendation requests per second across global regions.
  • Feature and catalog freshness: Ingest real time engagement signals within seconds, and surface newly published catalog items within 24 hours of ingestion.
  • Scalability: Scale horizontally to support over 300 million daily active users and a catalog of more than 100 million items.
  • Hybrid offline and online processing: Decouple heavy batch model retraining schedules from low latency online inference and real time streaming feature pipelines.
  • Rigorous experimentation: Support concurrent A/B testing infrastructure to validate model updates and ranking variants against live engagement guardrails.

Capacity Estimations

User volume, catalog size, and peak query throughput dictate whether the approximate nearest neighbor index fits within memory and how many candidates the ranking stage can evaluate. Because evaluating 100 million items on every request is impossible, the design sizes a multi stage funnel that retrieves roughly 3,000 candidates, applies cheap eligibility filters to select roughly 500 candidates for ranking, then applies authoritative eligibility checks before returning the top 20. At the 100K requests per second planning peak, ranking 500 candidates produces roughly 50 million candidate scores per second. Precomputed candidates remain an acceleration and fallback layer for active users rather than the final source of truth.

MetricCalculationValue
Baseline usersGiven in prompt100M
Baseline catalogGiven in prompt10M products
Baseline user item interactions / dayGiven in prompt scale1B
Planning DAUGrowth planning assumption300M
Recommendation requests / secPlanning peak target100K
Planning recommendation requests / day100K x 86,4008.64B
Planning recommendation requests / DAU / day8.64B ÷ 300M~28.8
Planning catalog sizeGrowth planning assumption100M items
Planning user item interactions / dayGrowth planning assumption10B
Average planning interaction ingest rate10B ÷ 86,400~115.7K events/sec
Model artifact sizeGiven planning range, excluding large embedding tables10-50 GB
Feature store planning scaleGiven planning scale300M users + 100M items
Hot user embedding working set1M active users x 512B~512 MB raw
Ranking candidate score throughput100K requests/sec x 500 ranked candidates50M candidate scores/sec
Illustrative GPU worker capacity5,000 scores ÷ 0.05 sec100K scores/sec per benchmark worker
Illustrative ranking workers before headroom50M ÷ 100K~500 workers
Candidate generation latencyGiven target< 50 ms
Ranking latency budgetGiven target< 100 ms

Architecture Diagram

Architecture Overview

When presenting this architecture, outline the multi stage funnel before detailing individual components. Candidate generation retrieves roughly 3,000 candidates and selects roughly 500 eligible candidates in the retrieval stage. The vector index layer targets under 10 milliseconds, while the representative end to end candidate generation budget is ~30 milliseconds. The ranking stage uses a cross feature scoring model with a representative ~40 millisecond target and a separate ~50 millisecond GPU batch benchmark for up to 5,000 scored items. Supporting workloads such as model training, feature backfills, and A/B test telemetry operate asynchronously through Kafka.

User and item embeddings are trained offline, while catalog vectors are indexed in a vector store using FAISS and HNSW. Recommendations are served through a retrieve then rank pipeline backed by Redis for active user features and candidate caching. The candidate pool can combine two tower ANN retrieval with popularity, content based, similar item, or policy driven candidate sources before deduplication and ranking. Authoritative catalog and inventory systems provide current item metadata, price, stock, and eligibility signals at serving time. The synchronous read path avoids scoring the full catalog, instead recalling hundreds of candidates and ranking dozens within the stated latency targets.

If an interviewer asks about collaborative filtering, clarify that matrix factorization operates as an offline retrieval or feature generation step, while online serving relies on approximate nearest neighbor retrieval followed by a scoring ranker and business eligibility checks.

Loading...

Three Stage Pipeline

The recommendation workflow cascades through candidate generation, deep neural network scoring, and business rule filtering to balance low latency with precision:

1. Candidate Generation (ANN Retrieval): ~3,000 candidates in ~30ms
   Two-Tower Model: dot_product(user_embedding, item_embeddings) via HNSW in Faiss
   Scale: Filters 100M catalog items down to top candidates
   Latency note: The vector index lookup itself can complete in ~5-10ms, while ~30ms represents the end-to-end candidate-generation budget including feature preparation, filtering, and service overhead

2. Ranking (Deep Neural Network): ~500 items ranked in ~40ms representative target via GPU batch inference
   Serving capacity: The ranking tier can score batches of up to ~5,000 items in the stated benchmark, whereas this request path ranks the selected 500 candidates returned after retrieval and eligibility filtering
   Scoring: Evaluates ~200 real time and batch features per (user, item) pair

3. Re-ranking (Business Rules): Top 20 items selected in ~10ms
   Enforcement: Deduplication, category diversity, exploration injection, and eligibility checks for inventory, region, policy, and other catalog constraints

Component Deep Dives

Two Tower Embedding Generation

Candidate generation maps user interaction histories and item metadata into a unified embedding space, enabling approximate nearest neighbor retrieval at catalog scale:

YAML
tower_1_user_encoder:
  inputs:
    - user interaction history sequence
    - user demographics (country, age group, language)
    - explicit category preferences
    - request context such as device and current session
  output: user_embedding (128 dimensional float vector)
  architecture: Transformer encoder over the user interaction sequence

tower_2_item_encoder:
  inputs:
    - item metadata (category, tags, title, description, price)
    - aggregate engagement and conversion statistics
  output: item_embedding (128 dimensional float vector)
  architecture: Multi-layer perceptron (MLP) over concatenated dense features

training:
  objective: Contrastive loss with in-batch negatives
  positive_label: Product purchase, add to cart, or strong engagement such as high dwell time
  negative_label: Unclicked eligible impressions for supervised negatives, and sampled eligible catalog items for contrastive negatives

online_inference:
  query_vector: user_embedding generated in real time
  catalog_lookup: dot_product(user_embedding, item_embeddings)
  index: HNSW approximate nearest neighbor search in Faiss
  performance: Top 3,000 candidates retrieved in under 10ms at the index layer
  note: HNSW offers sublinear approximate search behavior in practice, though exact latency and recall depend on index parameters and workload

Cold Start: New Users and New Items

Cold start handling addresses the data deficit when new accounts join or new catalog items are published, preventing empty recommendation responses through tiered fallback strategies:

YAML
new_user_strategies:
  demographics_based:
    description: Recommend popular products filtered by approved broad demographic, language, and geographic segments
  onboarding_survey:
    description: Capture initial category, brand, and price range selections during signup to seed baseline preferences
  exploration_allocation:
    description: Allocate a high exploration budget across varied categories and price bands, learning quickly from the first 10 user interactions
  multi_armed_bandits:
    description: Deploy multi armed bandit policies to select among diverse candidate pools while balancing conversion quality and retention

anonymous_user_strategies:
  session_context:
    description: Use current session views, clicks, cart actions, device context, and regional popularity without requiring a persistent user profile
  popularity_fallback:
    description: Serve regional and category popularity baselines until sufficient session signals accumulate

new_item_strategies:
  content_based_features:
    description: Predict candidate similarity from category, tags, descriptions, price, and other catalog attributes via the item tower
  seller_brand_reputation_signal:
    description: Boost new item visibility when seller or brand history demonstrates strong engagement or conversion quality
  guaranteed_exposure:
    description: Target at least 1,000 eligible impressions for each sufficiently trafficked new item through exploration slots when sufficient qualified traffic exists
  freshness_boost:
    description: Apply an explicit score multiplier in the re ranking stage during the item's first 48 hours

Event Bus Design (Kafka)

The event streaming backbone decouples real time user interaction ingestion from asynchronous model training and feature store updates:

YAML
topic: user-events
partitions: 256
partition_key: user_id for logged in users, session_id for anonymous users # preserves per-viewer ordering. Downstream Flink jobs can re-key when item aggregates are needed
retention_days: 14     # supports offline training windows and streaming feature updates
replication_factor: 3
min_insync_replicas: 2
producer_acks: all

producers:
  source: API gateway and client telemetry services
  triggers: User actions including impression, view, click, rating, add_to_cart, purchase, wishlist, and skip
  payload:
    event_id: string
    user_id: string | null
    session_id: string | null
    item_id: string
    action: impression | view | rating | click | add_to_cart | purchase | wishlist | skip
    engagement_duration_ms: integer | null
    context: structured request context
    occurred_at: ISO8601 timestamp
    schema_version: string
  idempotency: event_id is the stable deduplication key for retried telemetry

consumer_groups:
  feature_updater:
    engine: Apache Flink
    action: Updates Redis user profile, session features, and rolling item aggregates
  training_etl:
    engine: Apache Spark
    action: Batches events into Amazon S3 Parquet files for offline retrieval and ranking model retraining
  embedding_trainer:
    schedule: daily
    action: Refreshes collaborative filtering and retrieval embeddings

routing:
  online_read_path: GET /recommendations reads the Redis online feature store directly without a Kafka dependency and can use the batch candidate cache as an acceleration or fallback path
  offline_train_path: Raw events to batch training to versioned model artifacts and canary deployment
  dead_letter_queue:
    topic: user-events-dlq
    alert_condition: feature_updater consumer lag exceeds 120 seconds
    policy: Retry bounded transient failures first, sending only malformed or persistently unprocessable events to the DLQ

Catalog and Inventory Eligibility

Recommendation scores are not authoritative for availability. The serving path validates current inventory, region, price, policy, and catalog status against authoritative services before returning the final top 20 items:

YAML
eligibility_pipeline:
  prefilter: Apply cheap stable request constraints such as region, category, and policy before expensive ranking when possible
  ranking: Score candidates using the feature store and model features
  authoritative_check: Revalidate inventory, price, region, policy, and catalog active status immediately before response
  stale_data_policy: Never use a stale availability cache as the final purchase eligibility authority
  purchase_authority: Checkout or order placement must revalidate stock, price, and other purchase constraints because recommendation-time eligibility can become stale after the response
  fallback: If authoritative eligibility checks fail, remove affected items and fill from a validated fallback candidate pool

API Design

Domain Types and Service Interfaces

The recommendation service defines structured TypeScript contracts for personalized request identity, contextual client parameters, opaque pagination cursors, and asynchronous interaction events:

TYPESCRIPT
type UserId = string;
type AnonymousSessionId = string;
type ItemId = string;
type EventId = string;
type Cursor = string;
type ISO8601Timestamp = string;
type RecommendationCount = number;
type RecommendationScore = number;

type RecommendationSurface = "home" | "detail" | "trending" | "category" | "promotion";
type DeviceType = "web" | "ios" | "android" | "other";
type RequestId = string;
type ExperimentId = string;
type ModelVersion = string;
type ImpressionId = string;
type RecommendationReasonCode =
  | "similar_to_recent_view"
  | "popular_in_category"
  | "trending_in_region"
  | "exploration";
type InteractionAction =
  | "impression"
  | "view"
  | "rating"
  | "click"
  | "add_to_cart"
  | "purchase"
  | "wishlist"
  | "skip";
type ViewerIdentity =
  | { userId: UserId }
  | { sessionId: AnonymousSessionId };

export interface RecommendationItem {
  itemId: ItemId;
  title: string;
  score: RecommendationScore;
  reasonCode: RecommendationReasonCode;
}

export interface GetRecommendationsRequest {
  viewer: ViewerIdentity;
  surface: RecommendationSurface;
  count?: RecommendationCount; // validated by the service, for example 1 to 100
  cursor?: Cursor; // opaque cursor bound to viewer, surface, and the serving snapshot or version
  context?: {
    lastViewedItemId?: ItemId;
    deviceType?: DeviceType;
  };
}

export interface GetRecommendationsResponse {
  items: RecommendationItem[];
  nextCursor?: Cursor;
  requestId: RequestId;
  modelVersion: ModelVersion;
  experimentId?: ExperimentId;
}

export interface InteractionEvent {
  eventId: EventId;
  viewer: ViewerIdentity;
  itemId: ItemId;
  action: InteractionAction;
  engagementDurationMs?: number;
  position?: number;
  surface?: RecommendationSurface;
  requestId?: RequestId;
  impressionId?: ImpressionId;
  experimentId?: ExperimentId;
  modelVersion?: ModelVersion;
  occurredAt: ISO8601Timestamp;
}

export interface RecommendationService {
  // Retrieve personalized top-K recommendations for a target surface
  getRecommendations(req: GetRecommendationsRequest): Promise<GetRecommendationsResponse>;

  // Asynchronously ingest user interaction telemetry for real time feature updates
  recordEvent(event: InteractionEvent): Promise<{ accepted: boolean; eventId: EventId }>;
}

Recommendation and Event Ingestion Endpoints

Clients retrieve ranked recommendations synchronously with opaque cursor based pagination and publish idempotent user interaction telemetry asynchronously:

HTTP
GET /api/v1/recommendations?surface=home&count=20&cursor=eyJvZmZzZXQiOjIwLCJ2IjoxfQ HTTP/1.1
Host: api.recservice.internal
Authorization: Bearer <jwt-token>
# Anonymous requests omit Authorization and send a server issued or validated session token:
# X-Session-Id: <anonymous-session-id>

HTTP/1.1 200 OK
Content-Type: application/json

{
  "items": [
    {
      "item_id": "item-123",
      "title": "Noise Cancelling Headphones",
      "score": 0.95,
      "reason_code": "similar_to_recent_view"
    }
  ],
  "next_cursor": "eyJvZmZzZXQiOjQwLCJ2IjoxfQ",
  "request_id": "req-9ab2",
  "model_version": "v15",
  "experiment_id": "exp-42"
}

POST /api/v1/events HTTP/1.1
Host: api.recservice.internal
Content-Type: application/json
Idempotency-Key: evt-7f3a
Authorization: Bearer <jwt-token>
# Anonymous requests omit Authorization and send a server issued or validated session token:
# X-Session-Id: <anonymous-session-id>

{
  "event_id": "evt-7f3a",
  "user_id": "u1",
  "item_id": "item-123",
  "action": "view",
  "engagement_duration_ms": 3600,
  "surface": "home",
  "request_id": "req-9ab2",
  "impression_id": "imp-31cd",
  "position": 1,
  "experiment_id": "exp-42",
  "model_version": "v15",
  "occurred_at": "2026-09-06T12:00:00Z"
}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "status": "queued",
  "event_id": "evt-7f3a"
}

Common Error Responses

The API gateway returns standardized error codes when clients provide invalid pagination cursors, exceed rate limits, or encounter backend timeouts:

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
504 Gateway Timeout: search index shard responded slowly, narrow query parameters or retry

Data Model

Feature Store (Redis)

Redis functions as the low latency online feature store, serving batch materialized features, selected hot item vectors, user attributes, and real time interaction counters in under 5 milliseconds:

YAML
# User Feature Entities
user:{uid}:embedding:
  type: Binary
  size: 512 bytes (128 floats * 4 bytes)
  scope: Hot active users only
  description: Dense user preference vector generated by the user tower for the active serving profile
user:{uid}:history:
  type: List
  scope: Hot active users only
  capacity: Last 100 interacted item IDs
  description: Recent product views, clicks, and purchases for session feature generation
user:{uid}:profile:
  type: Hash
  fields: { country: string, age_group: string, language: string, signup_date: string }
  description: Broad approved demographic segments for cold start and segment filtering

# Item Feature Entities
item:{iid}:embedding:
  type: Binary
  size: 512 bytes (128 floats * 4 bytes)
  description: Dense item attribute vector for hot catalog items. Full catalog vectors live in the ANN index
item:{iid}:stats:
  type: Hash
  fields: { conversion_rate: float, add_to_cart_rate: float, view_count: integer }
  description: Rolling engagement and conversion statistics updated asynchronously

# Ephemeral Operational Keys
recent_items:{uid}:
  type: Set
  ttl: 30 days
  description: Recently interacted item IDs used for recommendation deduplication
recommendation_candidates:{uid}:
  type: List
  ttl: 30 minutes
  description: Precomputed top candidate item IDs for active users

Model Artifacts (Amazon S3)

Amazon S3 stores versioned retrieval and ranking model artifacts, frozen neural graph checkpoints, and precomputed collaborative filtering embeddings produced by offline training pipelines:

s3://models/retrieval-two-tower/v42/
  ├── user_embeddings.npy     (300M x 128 x 4 = ~150 GB)
  ├── item_embeddings.npy     (100M x 128 x 4 = ~50 GB)
  └── model_metadata.json

s3://models/collaborative-filtering/v42/
  └── matrix_factorization_embeddings.npy

s3://models/ranking-dnn/v15/
  ├── saved_model.pb
  └── model_config.json

Vector Index (Faiss / Milvus)

The candidate generation layer maintains a versioned approximate nearest neighbor index over catalog embeddings to support sub 10 millisecond index level similarity search:

YAML
vector_index_configuration:
  engine: Faiss / Milvus (HNSW index)
  catalog_size: 100M item embeddings (128 dimensional vectors)
  build_time: ~2 hours on a batch cluster (illustrative planning benchmark)
  search_latency: ~5ms for top-3,000 nearest neighbors at the index layer (illustrative benchmark)
  memory_footprint: ~65-75 GB per full index replica before replication, depending on HNSW parameters, allocator overhead, and metadata
  update_strategy: Nightly batch rebuild with streaming incremental inserts and updates for new or materially changed catalog items
  freshness_and_deletion: Versioned rebuilds, tombstones for removed items, and atomic index version swap with rollback to the previous validated version
  compatibility: Retrieval model version and ANN index version are validated as a pair before rollout or rollback

Fault Tolerance

Failure Modes and Mitigations

Production recommendation architectures isolate offline and online dependencies so failures in background training, feature stores, or vector indexes degrade recommendation quality without causing user facing outages:

ConcernSolution
Model serving failureFallback to the previous validated model artifact through blue green deployment routing
Feature store downServe candidates from an independently replicated or local cache when available, then degrade gracefully to popularity and catalog based recommendations. Final eligibility checks still validate stock, region, price, and policy.
ANN index unavailable or staleUse the last validated index snapshot when healthy, accept fresh catalog items through incremental inserts, fall back to cached candidate sets and popularity retrieval when necessary, and monitor index version, coverage, and freshness
Training data poisoningFilter automated bot, spam, anomalous, and policy violating interactions before training dataset compilation
Cold startUse regional popularity, onboarding preferences, content and catalog attributes, and exploration slots
Filter bubbleReserve 10% of recommendation slots for randomized exploration and novelty injection
User deletion or privacy requestPropagate deletion to online features and vector indexes, tombstone affected offline interactions, and rebuild or retrain derived artifacts when policy requires it

1. Feature Staleness

When a user interacts with catalog items during an active session, streaming feature pipelines must ingest signals rapidly to prevent serving stale recommendations:

YAML
failure_scenario:
  event: User views "Noise Cancelling Headphones" at T=0
  action: Refreshes feed at T=5s while stream processing pipeline has not completed
  impact: Online feature store still reflects older preferences

mitigation_strategies:
  1_client_context:
    strategy: Client passes last_viewed_item_id directly in request context
    advantage: Ranking model consumes immediate interaction context without stream ingestion lag
  2_session_reranking:
    strategy: Re-ranking stage boosts items sharing categories, brands, or price bands with recently viewed products
    advantage: Instant adaptation to fast session interest shifts
  3_bounded_freshness_sla:
    strategy: Target streaming propagation under 30 seconds via Apache Flink
    advantage: Stale intervals remain brief for the majority of active sessions

2. Popularity Bias and Exploration

Without active exploration, recommendation models reinforce historical popularity, creating rich get richer dynamics that starve new catalog items of user interactions:

YAML
exploration_budget:
  exploitation_allocation: "90% of feed slots assigned to model scored recommendations maximizing expected reward"
  exploration_allocation: "10% of feed slots reserved for exploratory items from underexposed catalog tiers"

bandit_strategy:
  algorithm: Thompson Sampling with Bayesian multi armed bandits
  state: Each eligible exploration arm or item tracks a Beta distribution parameterized by observed successes and failures
  selection: Draw random sample from each item posterior distribution and rank by sampled value
  outcome: Balances exploration of uncertain items against exploitation of proven high performers

Additional Considerations

Batch vs Real Time Recommendations

Balancing computation cost and personalization freshness requires a hybrid split between batch candidate generation and real time rescoring:

YAML
hybrid_serving_tier:
  batch_generation: Precomputes candidate sets of 500 items every 30 minutes for active users or high traffic segments as a fast path and fallback
  real_time_scoring: Re-scores candidates using fresh contextual features on every request
  business_reranking: Applies diversity, deduplication, inventory, region, price, and policy eligibility filters per request

performance_profile:
  total_serving_latency: ~150 ms for an uncached representative planning path including service and network overhead, while production SLO is enforced separately
  component_latency_budget: ~110 ms for an illustrative component benchmark of feature fetch, ANN retrieval, GPU ranking, and business re ranking
  feature_freshness: Ranking consumes real time signals including last viewed item, session interaction history, and time of day

Model Serving Infrastructure: GPU Inference

Evaluating complex neural ranking models across hundreds of candidates benefits from hardware acceleration and vectorized batch evaluation:

YAML
inference_batching:
  strategy: Vectorized tensor scoring of 5,000 candidate items in a single forward pass
  gpu_parallelism: 5,000 item scores completed in ~50ms (compared to sequential CPU evaluation requiring ~5s, stated as a benchmark assumption)

benchmark_latency_breakdown:
  redis_feature_fetch: 20ms
  ann_candidate_generation: 30ms
  gpu_batch_ranking: 50ms
  business_reranking: 10ms
  total_benchmark_path: ~110ms
  note: This 110ms figure is a separate large batch benchmark, not the production p95 end to end SLO

Real Time Feature Pipeline

User interactions stream through Apache Kafka and Apache Flink to update feature states incrementally, improving responsiveness in subsequent requests:

YAML
stream_processing_pipeline:
  event_ingestion: User action published to Kafka topic
  stream_computation: Apache Flink consumes event and executes three incremental updates:
    1_embedding_update: Adjusts lightweight user session embedding features in under 5ms and writes to Redis
    2_counter_increment: After atomic event_id deduplication and mutation, increments user:{uid}:category_interaction_count:electronics once per accepted event
    3_session_append: After atomic event_id deduplication and mutation, appends item_id to user:{uid}:session_history once per accepted event

replay_safety:
  deduplication_key: event_id
  strategy: For Flink, keep deduplication state and derived feature updates in checkpointed operator state where possible. For Redis, apply the deduplication marker and dependent counter or list mutation in one atomic Lua function or transaction so a crash cannot mark an event processed before its mutation succeeds
  external_write_rule: Do not claim exactly once for a non transactional external sink unless the mutation is idempotent or atomically coupled to deduplication. Kafka retries must not double count or duplicate session history
  dedupe_retention: Retain deduplication state for at least the event replay or retry window, and extend it when the source retention policy permits long replay

downstream_impact:
  subsequent_request: If the pipeline remains degraded until T+1 minute, the query service still reads the previous Redis embeddings and falls back to request context and cached candidates
  candidate_retrieval: Vector search retrieves candidates related to recent category and brand interactions
  ranking_execution: Ranker consumes updated interaction counts to adjust product scores

feature_freshness_tiers:
  real_time: Sub-minute latency for last viewed items, session history, and request time
  near_real_time: Sub-30-minute latency for updated user embeddings and rolling category preferences
  batch: Daily offline refreshes for broad demographics and long term historical aggregates

Offline Evaluation and Release Gates

Offline evaluation catches retrieval, ranking, calibration, coverage, and slice regressions before production traffic is exposed to a new model or vector index:

YAML
offline_evaluation:
  retrieval_metrics: Recall@100, Recall@500, Recall@3,000
  ranking_metrics: NDCG@20, conversion rate, revenue per session, calibration
  catalog_metrics: coverage, novelty, category diversity, eligibility coverage
  validation_set: Time based holdout with point in time feature reconstruction
  leakage_checks: Verify that features and labels only use information available before the prediction timestamp
  slice_checks: Region, device, user lifecycle stage, category, and traffic cohort performance
  policy_checks: Monitor recommendation coverage and outcome disparities across approved cohorts where applicable

release_gates:
  schema_compatibility: Feature schemas and model inputs must match the trained contract
  model_index_pair: Retrieval model version and ANN index version must be validated together
  shadow_validation: Compare candidate model outputs with production outputs before exposure
  canary: Start with a small traffic percentage and verify quality, latency, error rate, and guardrails
  rollback: Keep the previous validated model and index versions immediately available

Privacy and Data Lifecycle

Recommendation systems must propagate deletion and privacy decisions across online features, vector indexes, interaction logs, and derived model artifacts rather than deleting only the user profile:

YAML
privacy_lifecycle:
  online_features: Delete user embeddings, history, profile attributes, recent interaction keys, and cached candidates for the affected identity
  vector_index: Tombstone deleted catalog items immediately where applicable, while removing user embeddings from online feature stores and rebuilding derived user or item artifacts when required
  offline_events: Apply the approved deletion or tombstone policy to interaction logs and exclude deleted identities from future training datasets
  derived_models: Rebuild or retrain artifacts when policy requires removal of deleted user influence
  data_minimization: Use only policy approved demographic and behavioral features for recommendation decisions
  access_control: Restrict raw interaction data and feature stores to the services and operators that require access

Serving Observability

Production observability must connect infrastructure health to recommendation quality so that a low error rate does not hide degraded retrieval, eligibility, or personalization:

YAML
serving_observability:
  latency: p50, p95, p99 for feature fetch, ANN retrieval, ranking, reranking, and total request
  retrieval_quality: Sampled or offline Recall@K, candidate fill rate, ANN timeout rate, index version, and index freshness
  eligibility: Rejection rate by inventory, region, price, policy, and unavailable catalog metadata
  personalization: Feature freshness, fallback rate, cache hit rate, model score drift, and segment coverage
  experimentation: Traffic assignment integrity, guardrail regressions, conversion lift, diversity, and retention
  incident_alerts: Trigger alerts on sustained latency, feature staleness, index corruption, high fallback rate, or abnormal cohort behavior

Related Problems and Core Concepts

Recommendation pipelines connect closely with large scale streaming, search ranking, caching, and experimentation architectures:

  • TikTok (Short Video Platform): Applies the same retrieval and ranking funnel to a high scale personalized feed with real time behavioral signals.
  • Video Recommendation Engine: Deep dive into candidate generation, collaborative filtering at scale, and streaming interaction feature stores.
  • Search Ranking (Learning to Rank): Explores gradient boosted decision trees, point in time feature generation, and offline to online evaluation frameworks.
  • Music Streaming Platform: Demonstrates collaborative filtering matrix factorization and personalized recommendations integrated with large scale item delivery.
  • Stream Processing Basics: Core primitives for stateful stream computing, windowing, and incremental state aggregation using Apache Flink.
  • Caching Patterns and Invalidation: Strategies for low latency Redis feature caching, TTL management, and active user recommendation buffering.
  • Sharding and Partitioning: Partitioning strategies for scaling vector indexes, Redis feature stores, and Kafka event topics across clusters.
  • System Design Interview Patterns: Foundational frameworks for scoping requirements, structuring funnels, and managing latency budgets during senior interviews.

Interview Walkthrough

Follow this structured pacing to guide the conversation smoothly across candidate retrieval, neural ranking, real time feature streaming, business eligibility, and edge cases:

  • 25 minute cut

    Skip arch50 and arch75 depth unless the interview targets staff level.

    • Frame the challenge as a machine learning system design where retrieval narrows millions of items to hundreds and ranking orders the survivors (5 minutes).
    • Illustrate the two tower funnel, walking through user and item embedding generation, approximate nearest neighbor vector recall, and cross feature ranking (6 minutes).
    • Explain the dual role of the feature store in uniting offline batch feature tables with real time streaming signals from Kafka (5 minutes).
    • Address cold start strategies, explaining popularity fallbacks for new accounts and catalog attribute embeddings for newly published products (5 minutes).
    • For staff level discussions, detail exploration vs exploitation bandits, A/B experimentation frameworks, and diversity constraints during re ranking (4 minutes).
  • Structure the architecture as a multi stage funnel that progresses from candidate retrieval narrowing millions of items down to hundreds, through neural ranking, and finally to business rule re ranking.
  • Generate candidates with two tower embeddings and approximate nearest neighbor search using FAISS or HNSW, explaining that brute force scoring across the entire catalog violates the strict latency budget.
  • Precompute the top 500 candidates in periodic batch jobs every 30 minutes, and rescore candidates on every request using real time features from Redis such as recently viewed items and session interaction history.
  • Serve ranking inference on GPU clusters using request batching. At the planning peak of 100,000 requests per second and 500 ranked candidates per request, the system evaluates roughly 50 million candidate scores per second. The 5,000 score benchmark at roughly 50 milliseconds represents about 100,000 candidate scores per second per benchmark worker, implying roughly 500 workers before headroom, batching inefficiency, and failover capacity.
  • Inject exploration and exploitation balance during re ranking by reserving 10% to 15% of presentation slots for new products to resolve cold start and gather vital engagement data.
  • Stream user actions through Kafka to Apache Flink to update dynamic session features and rolling category counters within 1 minute of an observed interaction event. The 10 billion daily interaction planning volume averages roughly 115.7 thousand events per second before peak factor and retry overhead, which guides Kafka partition and Flink capacity planning.
  • State the end to end latency budget explicitly: allocate 20ms for feature fetching, 30ms for vector search, 50ms for GPU ranking inference, and 10ms for business re ranking, totaling approximately 110 milliseconds.
  • Avoid the common pitfall of proposing a single monolithic model that scores the full catalog on each request, because interviewers expect a clear separation between lightweight retrieval, deep ranking, and business eligibility.

Engineering Trade-offs

Collaborative vs Content Based vs Hybrid Filtering

Candidate generation architectures navigate fundamental trade offs between serendipitous discovery, behavioral coverage, and item metadata quality:

ApproachHowProsCons
Collaborative"Users like you also liked item X"Discovers unexpected interests across disparate categoriesSuffers from cold start for new users and exhibits strong popularity bias
Content Based"Similar attributes to items you previously engaged with"Handles new items when usable catalog attributes exist and provides transparent explainabilityLimited serendipity with a tendency toward overspecialization
HybridMultiple candidate generation towers feeding a unified rankerCombines diverse candidate pools with high personalization accuracyHigher operational complexity across multiple model training and feature pipelines

A/B Testing: How to Validate Model Changes

Validating algorithmic updates in production requires strict statistical discipline to ensure model changes improve business outcomes and long term user value rather than short term click behavior:

YAML
experiment_setup:
  control_group: 50% routed to current production model (v14)
  treatment_group: 50% routed to new candidate ranking model (v15)
  traffic_split: Deterministic hashing where hash(experiment_id + assignment_key) % 100 < 50
  duration: Minimum 7 days to account for day of week seasonality
  statistical_significance: p < 0.05 with minimum detectable effect of 0.5%
  assignment_key: user_id for logged in users or anonymous_session_id for anonymous users
  assignment: Sticky across sessions and supported devices for logged in users. For anonymous users, sticky within the anonymous session or an approved first party identifier
  integrity_checks: Sample ratio mismatch, assignment stickiness, bot and spam exclusion, complete exposure logging, and model version attribution
  analysis_plan: Predefine attribution windows and stopping rules. Use sequential testing controls if results are inspected before the planned horizon

metric_framework:
  primary_metric: Conversion rate per eligible recommendation session
  secondary_metric: Click-through rate across recommendation surfaces
  guardrail_metric: Next-day user retention rate
  counter_metric: Catalog category diversity entropy

pitfall_mitigation:
  risk: Engagement trap where model over-optimizes for clickbait
  solution: Optimize for conversion quality, dwell time, explicit satisfaction, revenue, and retention rather than raw clicks

Feedback Loops and Position Bias

Production recommendation systems continuously influence user behavior, creating systematic biases and feedback loops that can degrade model accuracy over time:

YAML
bias_and_feedback_challenges:
  position_bias:
    observation: In a measured cohort, items displayed at slot 1 receive 10x more clicks than slot 5 even at similar relevance
    mitigation: Log exposure position and serving propensity, use counterfactual or inverse propensity weighting during training, and exclude displayed position from online predictive features unless the same feature semantics are intentionally reproduced at serving time

  popularity_feedback_loop:
    observation: Highly ranked items receive disproportionate exposure, accumulating engagement that inflates future scores
    mitigation: Deploy an exploration budget alongside inverse propensity weighting to normalize training sample weights

  rabbit_hole_effect:
    observation: A single niche interaction causes the model to dominate recommendations with similar products
    mitigation: Enforce catalog diversity constraints during re ranking, limiting consecutive same-category items to 3 and capping single-category representation at 40% of the feed

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...