Interview Setup
Interview Prompt
Design an ML feature store serving 1M online feature retrievals/sec from a 2 TB Redis cluster, with 100 TB offline training data in S3 and point-in-time correct joins for 10K feature definitions.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Online store latency budget: 1ms or 5ms? |
|
| Point-in-time correctness required for training? | Without it, training uses future data → inflated offline metrics, poor online performance. |
| Batch features (daily aggregates) or streaming (real-time)? |
|
| Feature sharing across teams: governance model? | 10K definitions across teams needs registry, versioning, and access control. |
Scope
In scope
- Online vs offline store
- Feature freshness
- Point-in-time correctness
- Feature serving latency
- Capacity estimation with shown math
Out of scope (state explicitly)
- GPU cluster training and hyperparameter tuning
- Content moderation of recommended items
- Ad auction / sponsored placement ranking
Functional Requirements
Start by asking your interviewer whether teams share features across models and whether point-in-time correctness is required for training joins. A feature store is the contract between data scientists and production: clarify online vs offline serving paths before listing registry capabilities.
- Register feature definitions (name, type, entity, description, owner, SLA)
- Ingest features from batch (Spark/Hive) and streaming (Flink/Kafka) sources
- Serve features at low latency for online inference (< 5ms p99)
- Serve features in batch for model training (point-in-time correct joins)
- Feature versioning: track schema changes, backward compatibility
- Feature sharing: discover and reuse features across teams
- Point-in-time correctness: no data leakage
- Feature monitoring: drift detection, freshness alerts
- Feature lineage: trace from raw data source → transformation → feature
Non-Functional Requirements
Your interviewer will stress-test the tension between online latency (<5ms p99) and offline correctness, where point-in-time joins prevent data leakage while identical transformation logic in batch and stream pipelines prevents training-serving skew. They will also probe freshness SLAs per feature type.
- Ultra-Low Latency (Online): < 5ms p99 for feature retrieval during inference
- High Throughput (Online): 1M+ feature lookups/sec
- Training Correctness: Point-in-time joins must be exact (no future data leakage)
- Freshness: Online features reflect latest state (< 1 minute for streaming)
- Consistency: Online and offline features must compute identically (no training-serving skew)
- Scalability: 10K+ feature definitions, 1B+ entity rows
Capacity Estimations
Size the online store for 1M+ lookups/sec and the offline store for point-in-time training joins over billions of rows. Redis memory for hot features and S3/Parquet for historical snapshots are usually separate line items.
| Metric | Calculation | Value |
|---|---|---|
| Feature definitions | Given | 10,000 |
| Entities (users, items, etc.) | Given | 1B |
| Features per entity | Given | 50 to 200 |
| Online feature retrievals / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 1M |
| Online store size | Given | 2 TB (fits in Redis cluster) |
| Offline store size | Given | ~100 TB (S3/Parquet) |
| Batch ingestion (daily) | Given | 2 TB |
| Streaming ingestion | Given | 100K events/sec → feature updates |
Architecture Diagram
Draw the dual-write path: batch materialization from Spark into S3/Parquet for training, and stream materialization from Flink into Redis for online inference. Both paths read from the same feature definitions in the registry, because that shared contract is the core purpose of a feature store.
The online store serves 1M retrievals/sec at <5ms via MGET on Redis. The offline store holds 100 TB of historical snapshots for point-in-time joins, because training must fetch feature values as they existed at label time rather than the latest value.
Materialization pipelines sync offline to online on a schedule; streaming features update within minutes via Kafka. Feature monitoring (PSI, drift, freshness alerts) catches skew before it silently degrades model accuracy.
In the room
Ask whether point-in-time correctness is required, because without it, offline AUC looks great but the model fails in production when training data leaks future signals.
Component Deep Dives
Lead with training-serving skew as the primary production ML failure mode, followed by point-in-time correctness to prevent future data in labels, and materialization pipelines that keep batch and stream logic identical.
Training-Serving Skew
Training-serving skew is the most common production ML bug. Highlight it early because interviewers expect you to know why identical feature logic matters more than store choice.
What it is: Features computed differently in training vs serving leads to degraded model performance. Example: Training: avg_purchase computed with pandas (Python float64) Serving: avg_purchase computed in Java (Java double, different rounding) Result: 0.1% difference → model accuracy degrades Prevention: 1. Single transformation definition: same code for batch AND streaming 2. Validation job: compute features both ways, compare → alert on divergence 3. Feature logging: log online features at serving time → use THESE for training
Point-in-Time Correctness
Point-in-time correctness prevents data leakage, ensuring training labels at time T only join features computed at or before T rather than the latest value.
Wrong approach:
Join training data (March 1 prediction) with current features (March 14)
→ Model sees "future" information → fails in production
Correct approach:
For training example at time T:
feature_value = latest(feature WHERE timestamp <= T)
Implementation:
1. Store features as (entity_id, timestamp, value) triples
2. Training join: AS OF JOIN on timestamp
3. Delta Lake / Apache Iceberg support time-travel queries nativelyFeature Freshness: Batch vs Streaming vs On-Demand
Not every feature needs real-time freshness. Walk through batch, streaming, and on-demand compute, explaining how to assign SLAs per feature based on business impact.
Batch (Spark, daily/hourly): user_age_bucket, user_avg_purchase_value_30d: recomputed daily ✓ Cheap ✗ Stale: feature value may be 24 hours old Streaming (Flink, near real-time): user_clicks_last_5min, user_cart_total: updated every event ✓ Fresh (seconds latency) ✗ More infrastructure, higher cost On-demand (at inference time): "Does this user follow the seller of this product?" ✓ Always fresh ✗ Adds latency to inference pipeline Prioritization: Historical purchase categories: weekly batch Average order value last 30 days: daily batch Recent search queries: 5-minute streaming Current cart contents: on-demand (< 100ms)
Event Bus Design (Kafka)
Kafka connects ingestion to materialization, allowing streaming events to update Redis in near real time while batch snapshots land in S3 for training joins.
Topic: feature-ingestion-stream
Partitions: 128 (partition by entity_id: user_id / item_id)
Retention: 24h (streaming materialization buffer)
Producers: Flink/Spark Streaming jobs, Kafka Connect from warehouses
Consumers: Online store materializer (Redis/DynamoDB HSET updates)
Topic: feature-definition-changes
Partitions: 8 (low volume registry updates)
Events: feature_registered, schema_version_bumped, pipeline_deployed
Consumers: serving nodes cache invalidation, training job version pins
Topic: batch-feature-snapshots
Daily partition by feature_group + date (S3 sink via Kafka Connect)
Consumers: offline store (Parquet), point-in-time join for training
Online serve: GET /features/{entity} → Redis lookup only (< 5ms)
Streaming path: source event → Flink transform → feature-ingestion-stream → RedisAPI Design
Present online retrieval as a batch GET by entity ID, emphasizing MGET for <5ms p99 and local SDK caching. Offline APIs return historical snapshots keyed by event timestamp so training joins stay leakage-safe.
Feature Management and Retrieval Endpoints
# Feature definition
POST /api/features
{
"name": "user_avg_purchase_30d",
"entity": "user", "value_type": "FLOAT",
"description": "Average purchase amount in last 30 days",
"source": "transactions_table",
"freshness_sla": "1h"
}
# Online serving
GET /api/features/online?entity=user&entity_id=123,456&features=avg_purchase_30d,click_rate_7d
# Training data generation (batch)
POST /api/features/historical
{
"entity": "user",
"entity_ids_with_timestamps": [
{"entity_id": "123", "timestamp": "2026-03-01T10:00:00Z"}
],
"features": ["avg_purchase_30d", "click_rate_7d"]
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Online Store (Redis)
HSET user:123 avg_purchase_30d "45.20" click_rate_7d "0.032" last_login_hours "2.5" _updated_at "2026-03-14T10:05:00Z" EXPIRE user:123 86400
Offline Store (S3 / Delta Lake / Parquet)
Path: s3://feature-store/user/avg_purchase_30d/ Partitioned by date: date=2026-03-14/part-00000.parquet Schema: entity_id, timestamp, value, created_at
Feature Registry (PostgreSQL)
CREATE TABLE features (
feature_id UUID PRIMARY KEY,
name TEXT UNIQUE NOT NULL,
entity_type TEXT NOT NULL,
value_type TEXT NOT NULL,
description TEXT,
transformation_code TEXT,
freshness_sla INTERVAL,
owner_team TEXT,
deprecated BOOLEAN DEFAULT FALSE
);Fault Tolerance
Online Store Failure
Redis cluster failure → feature retrieval fails → inference fails Mitigations: 1. Redis Cluster with replicas (6-node: 3 masters + 3 replicas) 2. Local feature cache on inference servers (LRU, 5-min TTL) 3. Default values: if Redis unavailable, use population median for each feature (degrades accuracy, preserves availability) 4. Feature importance: top 10 features provide 80% of model accuracy → cache only critical features locally
Additional Considerations
Popular Feature Store Systems
When You Need a Feature Store
DON'T need: When managing fewer than 10 features for a single model and team, inline feature computation is sufficient. For batch-only machine learning, standard SQL views are typically enough.
DO need: Real-time serving under 10ms latency, multiple teams sharing features across models, ongoing training-serving skew issues, strict point-in-time correctness, or continuous streaming features.
Feast API Example
# features.py: define feature view
from feast import Entity, FeatureView, Field, FileSource
from feast.types import Float32
user = Entity(name="user_id", join_keys=["user_id"])
user_stats = FeatureView(
name="user_purchase_stats",
entities=[user],
schema=[Field(name="total_purchases_7d", dtype=Float32)],
source=FileSource(path="s3://features/user_stats.parquet"),
ttl=timedelta(days=1),
)
# Training (offline: point-in-time join)
training_df = store.get_historical_features(
entity_df=labels_df, # (user_id, event_timestamp, label)
features=["user_purchase_stats:total_purchases_7d"],
).to_df()
# Serving (online: sub-5ms Redis lookup)
online_features = store.get_online_features(
features=["user_purchase_stats:total_purchases_7d"],
entity_rows=[{"user_id": "usr_123"}],
).to_dict()Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- offline store and online store duality (5 min)
- point-in-time correct joins (6 min)
- feature versioning and lineage for reproducible training (5 min)
- materialization pipeline from batch compute to online key-value serving (5 min)
- training-serving skew detection and monitoring (4 min)
- Explain the duality between the offline store for batch training on historical snapshots and the online store for low-latency serving.
- Cover point-in-time correct joins, ensuring features reflect state at prediction time rather than the latest value.
- Discuss feature versioning and lineage for reproducible training runs and governance.
- Mention the materialization pipeline: running batch computations and pushing feature records to an online key-value store for serving.
- Cover training-serving skew detection as a first-class monitoring concern with drift and stability metrics.
- Common pitfall: online serving reads the latest feature value while training used historical snapshots, causing silent model degradation.
- Downstream integration: features produced here feed directly into real-time prediction pipelines such as Ad Click Prediction and Real-Time Bidding System.
Engineering Trade-offs
Online Store: Redis vs DynamoDB vs Aerospike
Redis (most common choice): ✓ Sub-millisecond latency (memory-only, < 0.5ms p99) ✓ Flexible data types (Hash, String, List) ✓ TTL natively (features expire automatically) ✗ Memory-only: 10M users x 50 features x 50B = 25 GB RAM Best for: < 100M active users DynamoDB (AWS managed): ✓ Auto-scaling, zero ops overhead ✗ Latency: 5-10ms without DAX Best for: AWS ecosystem, medium scale Aerospike (Uber, Criteo, Twitter): ✓ NVMe SSD-optimized, < 1ms latency ✓ 10-100x cheaper than Redis for same data volume ✗ More complex to operate Best for: 1B+ entities where RAM is too expensive
Training-Serving Skew: The Silent Killer
Common causes of training-serving skew: 1. Different code paths: Training: SQL in Spark Serving: Python application code with different time window Fix: Feature Store shares ONE feature definition 2. Timestamp handling: Training: uses order creation time Serving: uses current time Fix: point-in-time joins (as-of join on timestamp) 3. Missing value handling: Training: NaN → fill with 0 Serving: missing key → return None (not 0) Fix: feature store enforces consistent default values 4. Type mismatches: Training: integer feature [1, 2, 3] Serving: string feature ["1", "2", "3"] Fix: feature store enforces schema with type validation Monitoring: Log feature values at serving time → compare distribution to training KS test between training and serving distributions → alert on significant shift
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.