Interview Setup
Interview Prompt
Design a recommendation system for an ecommerce platform that suggests products to users. Support personalized recommendations for logged in users, reasonable defaults for new users (cold start), and balance exploitation (show what we know works) with exploration (discover new preferences). Handle 100M users and 10M products.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What surfaces need recommendations, such as the homepage, product pages, or promotional emails? | The homepage requires broad category coverage and personalization, whereas a product page can emphasize similar products, accessories, and compatible items, which requires different candidate signals. |
| Are recommendations personalized in real time or updated periodically in batch? | Session based recommendations respond to real time product views and clicks within seconds, whereas daily batch workflows compute collaborative filtering or retrieval embeddings offline with fundamentally different latency budgets. |
| What interaction signals are available, such as clicks, purchases, ratings, or dwell time? | Purchase signals are sparse but high intent, while click events are dense but noisy, which directly determines the training loss formulation and positive label definitions. |
| How do we measure success, such as click through rate, conversion rate, revenue, or catalog diversity? | Optimizing click through rate in isolation creates narrow filter bubbles, requiring multi objective ranking to balance engagement, revenue, and content diversity. |
Scope
In scope
- Two tower model for retrieval
- ANN (Approximate Nearest Neighbor) candidate generation
- Ranking model (LTR) with feature store
- Catalog and inventory eligibility signals
- Cold start for new users and new items
- Explore/exploit balance (multi armed bandit)
- Real time and batch feature pipeline
Out of scope (state explicitly)
- Short video For You Page feeds and media upload or CDN distribution covered in TikTok (Short Video Platform)
- Full model training pipeline (GPU cluster, hyperparameter tuning)
- Content moderation of recommended items
- Ad auction / sponsored recommendations
Functional Requirements
Recommendation systems solve a multi stage retrieval and ranking problem rather than scoring every catalog item on demand. When scoping the architecture with an interviewer, confirm the candidate generation boundaries, ranking criteria, cold start expectations, and explore-vs-exploit diversity requirements.
In an interview, introduce the multi stage funnel early, highlighting approximate nearest neighbor retrieval in roughly 10 milliseconds followed by a representative 40 millisecond ranking target over approximately 500 candidates. The separate 50 millisecond GPU benchmark covers a larger 5,000 item scoring batch, so the figures describe different capacity targets.
- Personalized recommendations: Generate relevant item suggestions tailored to each user interaction history and preference profile.
- Multiple surfaces: Deliver contextual recommendations across distinct product surfaces, including the homepage feed, product similarity modules, trending modules, category pages, and promotional surfaces.
- Real time signals: Ingest and incorporate recent user actions such as views, likes, and skips into session level recommendation features within seconds, with slower aggregate features following their documented freshness tier.
- Cold start handling: Provide engaging recommendations for newly registered users without historical data as well as newly published catalog items.
- Catalog diversity: Prevent repetitive filter bubbles by ensuring recommended lists expose users to varied product categories, brands, and price bands.
- Recommendation explainability: Present clear attribution cues alongside recommendations, such as "Similar to products you viewed" or "Popular in your area".
Non-Functional Requirements
The system must target under 100 milliseconds for p95 recommendation responses and under 200 milliseconds for p99 responses while filtering millions of catalog items down to dozens of high relevance items. The representative component budget is approximately 110 milliseconds before additional platform overhead, so the serving tier must rely on parallelism, caching, and efficient execution paths to meet the end to end SLO.
- Low latency: Deliver online recommendation responses in under 100 milliseconds for p95 requests and under 200 milliseconds for p99 requests.
- High throughput: Sustain peak loads exceeding 100,000 recommendation requests per second across global regions.
- Feature and catalog freshness: Ingest real time engagement signals within seconds, and surface newly published catalog items within 24 hours of ingestion.
- Scalability: Scale horizontally to support over 300 million daily active users and a catalog of more than 100 million items.
- Hybrid offline and online processing: Decouple heavy batch model retraining schedules from low latency online inference and real time streaming feature pipelines.
- Rigorous experimentation: Support concurrent A/B testing infrastructure to validate model updates and ranking variants against live engagement guardrails.
Capacity Estimations
User volume, catalog size, and peak query throughput dictate whether the approximate nearest neighbor index fits within memory and how many candidates the ranking stage can evaluate. Because evaluating 100 million items on every request is impossible, the design sizes a multi stage funnel that retrieves roughly 3,000 candidates, applies cheap eligibility filters to select roughly 500 candidates for ranking, then applies authoritative eligibility checks before returning the top 20. At the 100K requests per second planning peak, ranking 500 candidates produces roughly 50 million candidate scores per second. Precomputed candidates remain an acceleration and fallback layer for active users rather than the final source of truth.
| Metric | Calculation | Value |
|---|---|---|
| Baseline users | Given in prompt | 100M |
| Baseline catalog | Given in prompt | 10M products |
| Baseline user item interactions / day | Given in prompt scale | 1B |
| Planning DAU | Growth planning assumption | 300M |
| Recommendation requests / sec | Planning peak target | 100K |
| Planning recommendation requests / day | 100K x 86,400 | 8.64B |
| Planning recommendation requests / DAU / day | 8.64B ÷ 300M | ~28.8 |
| Planning catalog size | Growth planning assumption | 100M items |
| Planning user item interactions / day | Growth planning assumption | 10B |
| Average planning interaction ingest rate | 10B ÷ 86,400 | ~115.7K events/sec |
| Model artifact size | Given planning range, excluding large embedding tables | 10-50 GB |
| Feature store planning scale | Given planning scale | 300M users + 100M items |
| Hot user embedding working set | 1M active users x 512B | ~512 MB raw |
| Ranking candidate score throughput | 100K requests/sec x 500 ranked candidates | 50M candidate scores/sec |
| Illustrative GPU worker capacity | 5,000 scores ÷ 0.05 sec | 100K scores/sec per benchmark worker |
| Illustrative ranking workers before headroom | 50M ÷ 100K | ~500 workers |
| Candidate generation latency | Given target | < 50 ms |
| Ranking latency budget | Given target | < 100 ms |
Architecture Diagram
Architecture Overview
When presenting this architecture, outline the multi stage funnel before detailing individual components. Candidate generation retrieves roughly 3,000 candidates and selects roughly 500 eligible candidates in the retrieval stage. The vector index layer targets under 10 milliseconds, while the representative end to end candidate generation budget is ~30 milliseconds. The ranking stage uses a cross feature scoring model with a representative ~40 millisecond target and a separate ~50 millisecond GPU batch benchmark for up to 5,000 scored items. Supporting workloads such as model training, feature backfills, and A/B test telemetry operate asynchronously through Kafka.
User and item embeddings are trained offline, while catalog vectors are indexed in a vector store using FAISS and HNSW. Recommendations are served through a retrieve then rank pipeline backed by Redis for active user features and candidate caching. The candidate pool can combine two tower ANN retrieval with popularity, content based, similar item, or policy driven candidate sources before deduplication and ranking. Authoritative catalog and inventory systems provide current item metadata, price, stock, and eligibility signals at serving time. The synchronous read path avoids scoring the full catalog, instead recalling hundreds of candidates and ranking dozens within the stated latency targets.
If an interviewer asks about collaborative filtering, clarify that matrix factorization operates as an offline retrieval or feature generation step, while online serving relies on approximate nearest neighbor retrieval followed by a scoring ranker and business eligibility checks.
Three Stage Pipeline
The recommendation workflow cascades through candidate generation, deep neural network scoring, and business rule filtering to balance low latency with precision:
1. Candidate Generation (ANN Retrieval): ~3,000 candidates in ~30ms Two-Tower Model: dot_product(user_embedding, item_embeddings) via HNSW in Faiss Scale: Filters 100M catalog items down to top candidates Latency note: The vector index lookup itself can complete in ~5-10ms, while ~30ms represents the end-to-end candidate-generation budget including feature preparation, filtering, and service overhead 2. Ranking (Deep Neural Network): ~500 items ranked in ~40ms representative target via GPU batch inference Serving capacity: The ranking tier can score batches of up to ~5,000 items in the stated benchmark, whereas this request path ranks the selected 500 candidates returned after retrieval and eligibility filtering Scoring: Evaluates ~200 real time and batch features per (user, item) pair 3. Re-ranking (Business Rules): Top 20 items selected in ~10ms Enforcement: Deduplication, category diversity, exploration injection, and eligibility checks for inventory, region, policy, and other catalog constraints
Component Deep Dives
Two Tower Embedding Generation
Candidate generation maps user interaction histories and item metadata into a unified embedding space, enabling approximate nearest neighbor retrieval at catalog scale:
tower_1_user_encoder:
inputs:
- user interaction history sequence
- user demographics (country, age group, language)
- explicit category preferences
- request context such as device and current session
output: user_embedding (128 dimensional float vector)
architecture: Transformer encoder over the user interaction sequence
tower_2_item_encoder:
inputs:
- item metadata (category, tags, title, description, price)
- aggregate engagement and conversion statistics
output: item_embedding (128 dimensional float vector)
architecture: Multi-layer perceptron (MLP) over concatenated dense features
training:
objective: Contrastive loss with in-batch negatives
positive_label: Product purchase, add to cart, or strong engagement such as high dwell time
negative_label: Unclicked eligible impressions for supervised negatives, and sampled eligible catalog items for contrastive negatives
online_inference:
query_vector: user_embedding generated in real time
catalog_lookup: dot_product(user_embedding, item_embeddings)
index: HNSW approximate nearest neighbor search in Faiss
performance: Top 3,000 candidates retrieved in under 10ms at the index layer
note: HNSW offers sublinear approximate search behavior in practice, though exact latency and recall depend on index parameters and workloadCold Start: New Users and New Items
Cold start handling addresses the data deficit when new accounts join or new catalog items are published, preventing empty recommendation responses through tiered fallback strategies:
new_user_strategies:
demographics_based:
description: Recommend popular products filtered by approved broad demographic, language, and geographic segments
onboarding_survey:
description: Capture initial category, brand, and price range selections during signup to seed baseline preferences
exploration_allocation:
description: Allocate a high exploration budget across varied categories and price bands, learning quickly from the first 10 user interactions
multi_armed_bandits:
description: Deploy multi armed bandit policies to select among diverse candidate pools while balancing conversion quality and retention
anonymous_user_strategies:
session_context:
description: Use current session views, clicks, cart actions, device context, and regional popularity without requiring a persistent user profile
popularity_fallback:
description: Serve regional and category popularity baselines until sufficient session signals accumulate
new_item_strategies:
content_based_features:
description: Predict candidate similarity from category, tags, descriptions, price, and other catalog attributes via the item tower
seller_brand_reputation_signal:
description: Boost new item visibility when seller or brand history demonstrates strong engagement or conversion quality
guaranteed_exposure:
description: Target at least 1,000 eligible impressions for each sufficiently trafficked new item through exploration slots when sufficient qualified traffic exists
freshness_boost:
description: Apply an explicit score multiplier in the re ranking stage during the item's first 48 hoursEvent Bus Design (Kafka)
The event streaming backbone decouples real time user interaction ingestion from asynchronous model training and feature store updates:
topic: user-events
partitions: 256
partition_key: user_id for logged in users, session_id for anonymous users # preserves per-viewer ordering. Downstream Flink jobs can re-key when item aggregates are needed
retention_days: 14 # supports offline training windows and streaming feature updates
replication_factor: 3
min_insync_replicas: 2
producer_acks: all
producers:
source: API gateway and client telemetry services
triggers: User actions including impression, view, click, rating, add_to_cart, purchase, wishlist, and skip
payload:
event_id: string
user_id: string | null
session_id: string | null
item_id: string
action: impression | view | rating | click | add_to_cart | purchase | wishlist | skip
engagement_duration_ms: integer | null
context: structured request context
occurred_at: ISO8601 timestamp
schema_version: string
idempotency: event_id is the stable deduplication key for retried telemetry
consumer_groups:
feature_updater:
engine: Apache Flink
action: Updates Redis user profile, session features, and rolling item aggregates
training_etl:
engine: Apache Spark
action: Batches events into Amazon S3 Parquet files for offline retrieval and ranking model retraining
embedding_trainer:
schedule: daily
action: Refreshes collaborative filtering and retrieval embeddings
routing:
online_read_path: GET /recommendations reads the Redis online feature store directly without a Kafka dependency and can use the batch candidate cache as an acceleration or fallback path
offline_train_path: Raw events to batch training to versioned model artifacts and canary deployment
dead_letter_queue:
topic: user-events-dlq
alert_condition: feature_updater consumer lag exceeds 120 seconds
policy: Retry bounded transient failures first, sending only malformed or persistently unprocessable events to the DLQCatalog and Inventory Eligibility
Recommendation scores are not authoritative for availability. The serving path validates current inventory, region, price, policy, and catalog status against authoritative services before returning the final top 20 items:
eligibility_pipeline:
prefilter: Apply cheap stable request constraints such as region, category, and policy before expensive ranking when possible
ranking: Score candidates using the feature store and model features
authoritative_check: Revalidate inventory, price, region, policy, and catalog active status immediately before response
stale_data_policy: Never use a stale availability cache as the final purchase eligibility authority
purchase_authority: Checkout or order placement must revalidate stock, price, and other purchase constraints because recommendation-time eligibility can become stale after the response
fallback: If authoritative eligibility checks fail, remove affected items and fill from a validated fallback candidate poolAPI Design
Domain Types and Service Interfaces
The recommendation service defines structured TypeScript contracts for personalized request identity, contextual client parameters, opaque pagination cursors, and asynchronous interaction events:
type UserId = string;
type AnonymousSessionId = string;
type ItemId = string;
type EventId = string;
type Cursor = string;
type ISO8601Timestamp = string;
type RecommendationCount = number;
type RecommendationScore = number;
type RecommendationSurface = "home" | "detail" | "trending" | "category" | "promotion";
type DeviceType = "web" | "ios" | "android" | "other";
type RequestId = string;
type ExperimentId = string;
type ModelVersion = string;
type ImpressionId = string;
type RecommendationReasonCode =
| "similar_to_recent_view"
| "popular_in_category"
| "trending_in_region"
| "exploration";
type InteractionAction =
| "impression"
| "view"
| "rating"
| "click"
| "add_to_cart"
| "purchase"
| "wishlist"
| "skip";
type ViewerIdentity =
| { userId: UserId }
| { sessionId: AnonymousSessionId };
export interface RecommendationItem {
itemId: ItemId;
title: string;
score: RecommendationScore;
reasonCode: RecommendationReasonCode;
}
export interface GetRecommendationsRequest {
viewer: ViewerIdentity;
surface: RecommendationSurface;
count?: RecommendationCount; // validated by the service, for example 1 to 100
cursor?: Cursor; // opaque cursor bound to viewer, surface, and the serving snapshot or version
context?: {
lastViewedItemId?: ItemId;
deviceType?: DeviceType;
};
}
export interface GetRecommendationsResponse {
items: RecommendationItem[];
nextCursor?: Cursor;
requestId: RequestId;
modelVersion: ModelVersion;
experimentId?: ExperimentId;
}
export interface InteractionEvent {
eventId: EventId;
viewer: ViewerIdentity;
itemId: ItemId;
action: InteractionAction;
engagementDurationMs?: number;
position?: number;
surface?: RecommendationSurface;
requestId?: RequestId;
impressionId?: ImpressionId;
experimentId?: ExperimentId;
modelVersion?: ModelVersion;
occurredAt: ISO8601Timestamp;
}
export interface RecommendationService {
// Retrieve personalized top-K recommendations for a target surface
getRecommendations(req: GetRecommendationsRequest): Promise<GetRecommendationsResponse>;
// Asynchronously ingest user interaction telemetry for real time feature updates
recordEvent(event: InteractionEvent): Promise<{ accepted: boolean; eventId: EventId }>;
}Recommendation and Event Ingestion Endpoints
Clients retrieve ranked recommendations synchronously with opaque cursor based pagination and publish idempotent user interaction telemetry asynchronously:
GET /api/v1/recommendations?surface=home&count=20&cursor=eyJvZmZzZXQiOjIwLCJ2IjoxfQ HTTP/1.1
Host: api.recservice.internal
Authorization: Bearer <jwt-token>
# Anonymous requests omit Authorization and send a server issued or validated session token:
# X-Session-Id: <anonymous-session-id>
HTTP/1.1 200 OK
Content-Type: application/json
{
"items": [
{
"item_id": "item-123",
"title": "Noise Cancelling Headphones",
"score": 0.95,
"reason_code": "similar_to_recent_view"
}
],
"next_cursor": "eyJvZmZzZXQiOjQwLCJ2IjoxfQ",
"request_id": "req-9ab2",
"model_version": "v15",
"experiment_id": "exp-42"
}
POST /api/v1/events HTTP/1.1
Host: api.recservice.internal
Content-Type: application/json
Idempotency-Key: evt-7f3a
Authorization: Bearer <jwt-token>
# Anonymous requests omit Authorization and send a server issued or validated session token:
# X-Session-Id: <anonymous-session-id>
{
"event_id": "evt-7f3a",
"user_id": "u1",
"item_id": "item-123",
"action": "view",
"engagement_duration_ms": 3600,
"surface": "home",
"request_id": "req-9ab2",
"impression_id": "imp-31cd",
"position": 1,
"experiment_id": "exp-42",
"model_version": "v15",
"occurred_at": "2026-09-06T12:00:00Z"
}
HTTP/1.1 202 Accepted
Content-Type: application/json
{
"status": "queued",
"event_id": "evt-7f3a"
}Common Error Responses
The API gateway returns standardized error codes when clients provide invalid pagination cursors, exceed rate limits, or encounter backend timeouts:
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 504 Gateway Timeout: search index shard responded slowly, narrow query parameters or retry
Data Model
Feature Store (Redis)
Redis functions as the low latency online feature store, serving batch materialized features, selected hot item vectors, user attributes, and real time interaction counters in under 5 milliseconds:
# User Feature Entities
user:{uid}:embedding:
type: Binary
size: 512 bytes (128 floats * 4 bytes)
scope: Hot active users only
description: Dense user preference vector generated by the user tower for the active serving profile
user:{uid}:history:
type: List
scope: Hot active users only
capacity: Last 100 interacted item IDs
description: Recent product views, clicks, and purchases for session feature generation
user:{uid}:profile:
type: Hash
fields: { country: string, age_group: string, language: string, signup_date: string }
description: Broad approved demographic segments for cold start and segment filtering
# Item Feature Entities
item:{iid}:embedding:
type: Binary
size: 512 bytes (128 floats * 4 bytes)
description: Dense item attribute vector for hot catalog items. Full catalog vectors live in the ANN index
item:{iid}:stats:
type: Hash
fields: { conversion_rate: float, add_to_cart_rate: float, view_count: integer }
description: Rolling engagement and conversion statistics updated asynchronously
# Ephemeral Operational Keys
recent_items:{uid}:
type: Set
ttl: 30 days
description: Recently interacted item IDs used for recommendation deduplication
recommendation_candidates:{uid}:
type: List
ttl: 30 minutes
description: Precomputed top candidate item IDs for active usersModel Artifacts (Amazon S3)
Amazon S3 stores versioned retrieval and ranking model artifacts, frozen neural graph checkpoints, and precomputed collaborative filtering embeddings produced by offline training pipelines:
s3://models/retrieval-two-tower/v42/ ├── user_embeddings.npy (300M x 128 x 4 = ~150 GB) ├── item_embeddings.npy (100M x 128 x 4 = ~50 GB) └── model_metadata.json s3://models/collaborative-filtering/v42/ └── matrix_factorization_embeddings.npy s3://models/ranking-dnn/v15/ ├── saved_model.pb └── model_config.json
Vector Index (Faiss / Milvus)
The candidate generation layer maintains a versioned approximate nearest neighbor index over catalog embeddings to support sub 10 millisecond index level similarity search:
vector_index_configuration:
engine: Faiss / Milvus (HNSW index)
catalog_size: 100M item embeddings (128 dimensional vectors)
build_time: ~2 hours on a batch cluster (illustrative planning benchmark)
search_latency: ~5ms for top-3,000 nearest neighbors at the index layer (illustrative benchmark)
memory_footprint: ~65-75 GB per full index replica before replication, depending on HNSW parameters, allocator overhead, and metadata
update_strategy: Nightly batch rebuild with streaming incremental inserts and updates for new or materially changed catalog items
freshness_and_deletion: Versioned rebuilds, tombstones for removed items, and atomic index version swap with rollback to the previous validated version
compatibility: Retrieval model version and ANN index version are validated as a pair before rollout or rollbackFault Tolerance
Failure Modes and Mitigations
Production recommendation architectures isolate offline and online dependencies so failures in background training, feature stores, or vector indexes degrade recommendation quality without causing user facing outages:
| Concern | Solution |
|---|---|
| Model serving failure | Fallback to the previous validated model artifact through blue green deployment routing |
| Feature store down | Serve candidates from an independently replicated or local cache when available, then degrade gracefully to popularity and catalog based recommendations. Final eligibility checks still validate stock, region, price, and policy. |
| ANN index unavailable or stale | Use the last validated index snapshot when healthy, accept fresh catalog items through incremental inserts, fall back to cached candidate sets and popularity retrieval when necessary, and monitor index version, coverage, and freshness |
| Training data poisoning | Filter automated bot, spam, anomalous, and policy violating interactions before training dataset compilation |
| Cold start | Use regional popularity, onboarding preferences, content and catalog attributes, and exploration slots |
| Filter bubble | Reserve 10% of recommendation slots for randomized exploration and novelty injection |
| User deletion or privacy request | Propagate deletion to online features and vector indexes, tombstone affected offline interactions, and rebuild or retrain derived artifacts when policy requires it |
1. Feature Staleness
When a user interacts with catalog items during an active session, streaming feature pipelines must ingest signals rapidly to prevent serving stale recommendations:
failure_scenario:
event: User views "Noise Cancelling Headphones" at T=0
action: Refreshes feed at T=5s while stream processing pipeline has not completed
impact: Online feature store still reflects older preferences
mitigation_strategies:
1_client_context:
strategy: Client passes last_viewed_item_id directly in request context
advantage: Ranking model consumes immediate interaction context without stream ingestion lag
2_session_reranking:
strategy: Re-ranking stage boosts items sharing categories, brands, or price bands with recently viewed products
advantage: Instant adaptation to fast session interest shifts
3_bounded_freshness_sla:
strategy: Target streaming propagation under 30 seconds via Apache Flink
advantage: Stale intervals remain brief for the majority of active sessions2. Popularity Bias and Exploration
Without active exploration, recommendation models reinforce historical popularity, creating rich get richer dynamics that starve new catalog items of user interactions:
exploration_budget:
exploitation_allocation: "90% of feed slots assigned to model scored recommendations maximizing expected reward"
exploration_allocation: "10% of feed slots reserved for exploratory items from underexposed catalog tiers"
bandit_strategy:
algorithm: Thompson Sampling with Bayesian multi armed bandits
state: Each eligible exploration arm or item tracks a Beta distribution parameterized by observed successes and failures
selection: Draw random sample from each item posterior distribution and rank by sampled value
outcome: Balances exploration of uncertain items against exploitation of proven high performersAdditional Considerations
Batch vs Real Time Recommendations
Balancing computation cost and personalization freshness requires a hybrid split between batch candidate generation and real time rescoring:
hybrid_serving_tier:
batch_generation: Precomputes candidate sets of 500 items every 30 minutes for active users or high traffic segments as a fast path and fallback
real_time_scoring: Re-scores candidates using fresh contextual features on every request
business_reranking: Applies diversity, deduplication, inventory, region, price, and policy eligibility filters per request
performance_profile:
total_serving_latency: ~150 ms for an uncached representative planning path including service and network overhead, while production SLO is enforced separately
component_latency_budget: ~110 ms for an illustrative component benchmark of feature fetch, ANN retrieval, GPU ranking, and business re ranking
feature_freshness: Ranking consumes real time signals including last viewed item, session interaction history, and time of dayModel Serving Infrastructure: GPU Inference
Evaluating complex neural ranking models across hundreds of candidates benefits from hardware acceleration and vectorized batch evaluation:
inference_batching:
strategy: Vectorized tensor scoring of 5,000 candidate items in a single forward pass
gpu_parallelism: 5,000 item scores completed in ~50ms (compared to sequential CPU evaluation requiring ~5s, stated as a benchmark assumption)
benchmark_latency_breakdown:
redis_feature_fetch: 20ms
ann_candidate_generation: 30ms
gpu_batch_ranking: 50ms
business_reranking: 10ms
total_benchmark_path: ~110ms
note: This 110ms figure is a separate large batch benchmark, not the production p95 end to end SLOReal Time Feature Pipeline
User interactions stream through Apache Kafka and Apache Flink to update feature states incrementally, improving responsiveness in subsequent requests:
stream_processing_pipeline:
event_ingestion: User action published to Kafka topic
stream_computation: Apache Flink consumes event and executes three incremental updates:
1_embedding_update: Adjusts lightweight user session embedding features in under 5ms and writes to Redis
2_counter_increment: After atomic event_id deduplication and mutation, increments user:{uid}:category_interaction_count:electronics once per accepted event
3_session_append: After atomic event_id deduplication and mutation, appends item_id to user:{uid}:session_history once per accepted event
replay_safety:
deduplication_key: event_id
strategy: For Flink, keep deduplication state and derived feature updates in checkpointed operator state where possible. For Redis, apply the deduplication marker and dependent counter or list mutation in one atomic Lua function or transaction so a crash cannot mark an event processed before its mutation succeeds
external_write_rule: Do not claim exactly once for a non transactional external sink unless the mutation is idempotent or atomically coupled to deduplication. Kafka retries must not double count or duplicate session history
dedupe_retention: Retain deduplication state for at least the event replay or retry window, and extend it when the source retention policy permits long replay
downstream_impact:
subsequent_request: If the pipeline remains degraded until T+1 minute, the query service still reads the previous Redis embeddings and falls back to request context and cached candidates
candidate_retrieval: Vector search retrieves candidates related to recent category and brand interactions
ranking_execution: Ranker consumes updated interaction counts to adjust product scores
feature_freshness_tiers:
real_time: Sub-minute latency for last viewed items, session history, and request time
near_real_time: Sub-30-minute latency for updated user embeddings and rolling category preferences
batch: Daily offline refreshes for broad demographics and long term historical aggregatesOffline Evaluation and Release Gates
Offline evaluation catches retrieval, ranking, calibration, coverage, and slice regressions before production traffic is exposed to a new model or vector index:
offline_evaluation:
retrieval_metrics: Recall@100, Recall@500, Recall@3,000
ranking_metrics: NDCG@20, conversion rate, revenue per session, calibration
catalog_metrics: coverage, novelty, category diversity, eligibility coverage
validation_set: Time based holdout with point in time feature reconstruction
leakage_checks: Verify that features and labels only use information available before the prediction timestamp
slice_checks: Region, device, user lifecycle stage, category, and traffic cohort performance
policy_checks: Monitor recommendation coverage and outcome disparities across approved cohorts where applicable
release_gates:
schema_compatibility: Feature schemas and model inputs must match the trained contract
model_index_pair: Retrieval model version and ANN index version must be validated together
shadow_validation: Compare candidate model outputs with production outputs before exposure
canary: Start with a small traffic percentage and verify quality, latency, error rate, and guardrails
rollback: Keep the previous validated model and index versions immediately availablePrivacy and Data Lifecycle
Recommendation systems must propagate deletion and privacy decisions across online features, vector indexes, interaction logs, and derived model artifacts rather than deleting only the user profile:
privacy_lifecycle:
online_features: Delete user embeddings, history, profile attributes, recent interaction keys, and cached candidates for the affected identity
vector_index: Tombstone deleted catalog items immediately where applicable, while removing user embeddings from online feature stores and rebuilding derived user or item artifacts when required
offline_events: Apply the approved deletion or tombstone policy to interaction logs and exclude deleted identities from future training datasets
derived_models: Rebuild or retrain artifacts when policy requires removal of deleted user influence
data_minimization: Use only policy approved demographic and behavioral features for recommendation decisions
access_control: Restrict raw interaction data and feature stores to the services and operators that require accessServing Observability
Production observability must connect infrastructure health to recommendation quality so that a low error rate does not hide degraded retrieval, eligibility, or personalization:
serving_observability:
latency: p50, p95, p99 for feature fetch, ANN retrieval, ranking, reranking, and total request
retrieval_quality: Sampled or offline Recall@K, candidate fill rate, ANN timeout rate, index version, and index freshness
eligibility: Rejection rate by inventory, region, price, policy, and unavailable catalog metadata
personalization: Feature freshness, fallback rate, cache hit rate, model score drift, and segment coverage
experimentation: Traffic assignment integrity, guardrail regressions, conversion lift, diversity, and retention
incident_alerts: Trigger alerts on sustained latency, feature staleness, index corruption, high fallback rate, or abnormal cohort behaviorRelated Problems and Core Concepts
Recommendation pipelines connect closely with large scale streaming, search ranking, caching, and experimentation architectures:
- TikTok (Short Video Platform): Applies the same retrieval and ranking funnel to a high scale personalized feed with real time behavioral signals.
- Video Recommendation Engine: Deep dive into candidate generation, collaborative filtering at scale, and streaming interaction feature stores.
- Search Ranking (Learning to Rank): Explores gradient boosted decision trees, point in time feature generation, and offline to online evaluation frameworks.
- Music Streaming Platform: Demonstrates collaborative filtering matrix factorization and personalized recommendations integrated with large scale item delivery.
- Stream Processing Basics: Core primitives for stateful stream computing, windowing, and incremental state aggregation using Apache Flink.
- Caching Patterns and Invalidation: Strategies for low latency Redis feature caching, TTL management, and active user recommendation buffering.
- Sharding and Partitioning: Partitioning strategies for scaling vector indexes, Redis feature stores, and Kafka event topics across clusters.
- System Design Interview Patterns: Foundational frameworks for scoping requirements, structuring funnels, and managing latency budgets during senior interviews.
Interview Walkthrough
Follow this structured pacing to guide the conversation smoothly across candidate retrieval, neural ranking, real time feature streaming, business eligibility, and edge cases:
- 25 minute cut
Skip arch50 and arch75 depth unless the interview targets staff level.
- Frame the challenge as a machine learning system design where retrieval narrows millions of items to hundreds and ranking orders the survivors (5 minutes).
- Illustrate the two tower funnel, walking through user and item embedding generation, approximate nearest neighbor vector recall, and cross feature ranking (6 minutes).
- Explain the dual role of the feature store in uniting offline batch feature tables with real time streaming signals from Kafka (5 minutes).
- Address cold start strategies, explaining popularity fallbacks for new accounts and catalog attribute embeddings for newly published products (5 minutes).
- For staff level discussions, detail exploration vs exploitation bandits, A/B experimentation frameworks, and diversity constraints during re ranking (4 minutes).
- Structure the architecture as a multi stage funnel that progresses from candidate retrieval narrowing millions of items down to hundreds, through neural ranking, and finally to business rule re ranking.
- Generate candidates with two tower embeddings and approximate nearest neighbor search using FAISS or HNSW, explaining that brute force scoring across the entire catalog violates the strict latency budget.
- Precompute the top 500 candidates in periodic batch jobs every 30 minutes, and rescore candidates on every request using real time features from Redis such as recently viewed items and session interaction history.
- Serve ranking inference on GPU clusters using request batching. At the planning peak of 100,000 requests per second and 500 ranked candidates per request, the system evaluates roughly 50 million candidate scores per second. The 5,000 score benchmark at roughly 50 milliseconds represents about 100,000 candidate scores per second per benchmark worker, implying roughly 500 workers before headroom, batching inefficiency, and failover capacity.
- Inject exploration and exploitation balance during re ranking by reserving 10% to 15% of presentation slots for new products to resolve cold start and gather vital engagement data.
- Stream user actions through Kafka to Apache Flink to update dynamic session features and rolling category counters within 1 minute of an observed interaction event. The 10 billion daily interaction planning volume averages roughly 115.7 thousand events per second before peak factor and retry overhead, which guides Kafka partition and Flink capacity planning.
- State the end to end latency budget explicitly: allocate 20ms for feature fetching, 30ms for vector search, 50ms for GPU ranking inference, and 10ms for business re ranking, totaling approximately 110 milliseconds.
- Avoid the common pitfall of proposing a single monolithic model that scores the full catalog on each request, because interviewers expect a clear separation between lightweight retrieval, deep ranking, and business eligibility.
Engineering Trade-offs
Collaborative vs Content Based vs Hybrid Filtering
Candidate generation architectures navigate fundamental trade offs between serendipitous discovery, behavioral coverage, and item metadata quality:
| Approach | How | Pros | Cons |
|---|---|---|---|
| Collaborative | "Users like you also liked item X" | Discovers unexpected interests across disparate categories | Suffers from cold start for new users and exhibits strong popularity bias |
| Content Based | "Similar attributes to items you previously engaged with" | Handles new items when usable catalog attributes exist and provides transparent explainability | Limited serendipity with a tendency toward overspecialization |
| Hybrid | Multiple candidate generation towers feeding a unified ranker | Combines diverse candidate pools with high personalization accuracy | Higher operational complexity across multiple model training and feature pipelines |
A/B Testing: How to Validate Model Changes
Validating algorithmic updates in production requires strict statistical discipline to ensure model changes improve business outcomes and long term user value rather than short term click behavior:
experiment_setup:
control_group: 50% routed to current production model (v14)
treatment_group: 50% routed to new candidate ranking model (v15)
traffic_split: Deterministic hashing where hash(experiment_id + assignment_key) % 100 < 50
duration: Minimum 7 days to account for day of week seasonality
statistical_significance: p < 0.05 with minimum detectable effect of 0.5%
assignment_key: user_id for logged in users or anonymous_session_id for anonymous users
assignment: Sticky across sessions and supported devices for logged in users. For anonymous users, sticky within the anonymous session or an approved first party identifier
integrity_checks: Sample ratio mismatch, assignment stickiness, bot and spam exclusion, complete exposure logging, and model version attribution
analysis_plan: Predefine attribution windows and stopping rules. Use sequential testing controls if results are inspected before the planned horizon
metric_framework:
primary_metric: Conversion rate per eligible recommendation session
secondary_metric: Click-through rate across recommendation surfaces
guardrail_metric: Next-day user retention rate
counter_metric: Catalog category diversity entropy
pitfall_mitigation:
risk: Engagement trap where model over-optimizes for clickbait
solution: Optimize for conversion quality, dwell time, explicit satisfaction, revenue, and retention rather than raw clicksFeedback Loops and Position Bias
Production recommendation systems continuously influence user behavior, creating systematic biases and feedback loops that can degrade model accuracy over time:
bias_and_feedback_challenges:
position_bias:
observation: In a measured cohort, items displayed at slot 1 receive 10x more clicks than slot 5 even at similar relevance
mitigation: Log exposure position and serving propensity, use counterfactual or inverse propensity weighting during training, and exclude displayed position from online predictive features unless the same feature semantics are intentionally reproduced at serving time
popularity_feedback_loop:
observation: Highly ranked items receive disproportionate exposure, accumulating engagement that inflates future scores
mitigation: Deploy an exploration budget alongside inverse propensity weighting to normalize training sample weights
rabbit_hole_effect:
observation: A single niche interaction causes the model to dominate recommendations with similar products
mitigation: Enforce catalog diversity constraints during re ranking, limiting consecutive same-category items to 3 and capping single-category representation at 40% of the feedReview
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.