System Design Problem

Design an ML Feature Store

Commonly Asked By:UberFeastTectonAirbnbGoogle

Interview Setup

Interview Prompt

Design an ML feature store serving 1M online feature retrievals/sec from a 2 TB Redis cluster, with 100 TB offline training data in S3 and point-in-time correct joins for 10K feature definitions.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Online store latency budget: 1ms or 5ms?
  • 1M retrievals/sec
  • 5ms budget allows batch fetch of 50 to 200 features per entity.
Point-in-time correctness required for training?Without it, training uses future data → inflated offline metrics, poor online performance.
Batch features (daily aggregates) or streaming (real-time)?
  • 100K streaming events/sec for real-time
  • 2 TB/day batch for historical aggregates.
Feature sharing across teams: governance model?10K definitions across teams needs registry, versioning, and access control.

Scope

In scope

  • Online vs offline store
  • Feature freshness
  • Point-in-time correctness
  • Feature serving latency
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • GPU cluster training and hyperparameter tuning
  • Content moderation of recommended items
  • Ad auction / sponsored placement ranking

Functional Requirements

Start by asking your interviewer whether teams share features across models and whether point-in-time correctness is required for training joins. A feature store is the contract between data scientists and production: clarify online vs offline serving paths before listing registry capabilities.

  • Register feature definitions (name, type, entity, description, owner, SLA)
  • Ingest features from batch (Spark/Hive) and streaming (Flink/Kafka) sources
  • Serve features at low latency for online inference (< 5ms p99)
  • Serve features in batch for model training (point-in-time correct joins)
  • Feature versioning: track schema changes, backward compatibility
  • Feature sharing: discover and reuse features across teams
  • Point-in-time correctness: no data leakage
  • Feature monitoring: drift detection, freshness alerts
  • Feature lineage: trace from raw data source → transformation → feature

Non-Functional Requirements

Your interviewer will stress-test the tension between online latency (<5ms p99) and offline correctness, where point-in-time joins prevent data leakage while identical transformation logic in batch and stream pipelines prevents training-serving skew. They will also probe freshness SLAs per feature type.

  • Ultra-Low Latency (Online): < 5ms p99 for feature retrieval during inference
  • High Throughput (Online): 1M+ feature lookups/sec
  • Training Correctness: Point-in-time joins must be exact (no future data leakage)
  • Freshness: Online features reflect latest state (< 1 minute for streaming)
  • Consistency: Online and offline features must compute identically (no training-serving skew)
  • Scalability: 10K+ feature definitions, 1B+ entity rows

Capacity Estimations

Size the online store for 1M+ lookups/sec and the offline store for point-in-time training joins over billions of rows. Redis memory for hot features and S3/Parquet for historical snapshots are usually separate line items.

MetricCalculationValue
Feature definitionsGiven10,000
Entities (users, items, etc.)Given1B
Features per entityGiven50 to 200
Online feature retrievals / secDerived from daily volume ÷ 86400 (+ peak factor)1M
Online store sizeGiven2 TB (fits in Redis cluster)
Offline store sizeGiven~100 TB (S3/Parquet)
Batch ingestion (daily)Given2 TB
Streaming ingestionGiven100K events/sec → feature updates

Architecture Diagram

Draw the dual-write path: batch materialization from Spark into S3/Parquet for training, and stream materialization from Flink into Redis for online inference. Both paths read from the same feature definitions in the registry, because that shared contract is the core purpose of a feature store.

The online store serves 1M retrievals/sec at <5ms via MGET on Redis. The offline store holds 100 TB of historical snapshots for point-in-time joins, because training must fetch feature values as they existed at label time rather than the latest value.

Materialization pipelines sync offline to online on a schedule; streaming features update within minutes via Kafka. Feature monitoring (PSI, drift, freshness alerts) catches skew before it silently degrades model accuracy.

Loading...

In the room

Ask whether point-in-time correctness is required, because without it, offline AUC looks great but the model fails in production when training data leaks future signals.

Component Deep Dives

Lead with training-serving skew as the primary production ML failure mode, followed by point-in-time correctness to prevent future data in labels, and materialization pipelines that keep batch and stream logic identical.

Training-Serving Skew

Training-serving skew is the most common production ML bug. Highlight it early because interviewers expect you to know why identical feature logic matters more than store choice.

What it is: Features computed differently in training vs serving leads to degraded model performance.

Example:
  Training: avg_purchase computed with pandas (Python float64)
  Serving: avg_purchase computed in Java (Java double, different rounding)
  Result: 0.1% difference → model accuracy degrades

Prevention:
  1. Single transformation definition: same code for batch AND streaming
  2. Validation job: compute features both ways, compare → alert on divergence
  3. Feature logging: log online features at serving time → use THESE for training

Point-in-Time Correctness

Point-in-time correctness prevents data leakage, ensuring training labels at time T only join features computed at or before T rather than the latest value.

Wrong approach:
  Join training data (March 1 prediction) with current features (March 14)
  → Model sees "future" information → fails in production

Correct approach:
  For training example at time T:
    feature_value = latest(feature WHERE timestamp <= T)

Implementation:
  1. Store features as (entity_id, timestamp, value) triples
  2. Training join: AS OF JOIN on timestamp
  3. Delta Lake / Apache Iceberg support time-travel queries natively

Feature Freshness: Batch vs Streaming vs On-Demand

Not every feature needs real-time freshness. Walk through batch, streaming, and on-demand compute, explaining how to assign SLAs per feature based on business impact.

Batch (Spark, daily/hourly):
  user_age_bucket, user_avg_purchase_value_30d: recomputed daily
  ✓ Cheap
  ✗ Stale: feature value may be 24 hours old

Streaming (Flink, near real-time):
  user_clicks_last_5min, user_cart_total: updated every event
  ✓ Fresh (seconds latency)
  ✗ More infrastructure, higher cost

On-demand (at inference time):
  "Does this user follow the seller of this product?"
  ✓ Always fresh
  ✗ Adds latency to inference pipeline

Prioritization:
  Historical purchase categories: weekly batch
  Average order value last 30 days: daily batch
  Recent search queries: 5-minute streaming
  Current cart contents: on-demand (< 100ms)

Event Bus Design (Kafka)

Kafka connects ingestion to materialization, allowing streaming events to update Redis in near real time while batch snapshots land in S3 for training joins.

Topic: feature-ingestion-stream
  Partitions: 128 (partition by entity_id: user_id / item_id)
  Retention: 24h (streaming materialization buffer)
  Producers: Flink/Spark Streaming jobs, Kafka Connect from warehouses
  Consumers: Online store materializer (Redis/DynamoDB HSET updates)

Topic: feature-definition-changes
  Partitions: 8 (low volume registry updates)
  Events: feature_registered, schema_version_bumped, pipeline_deployed
  Consumers: serving nodes cache invalidation, training job version pins

Topic: batch-feature-snapshots
  Daily partition by feature_group + date (S3 sink via Kafka Connect)
  Consumers: offline store (Parquet), point-in-time join for training

Online serve: GET /features/{entity} → Redis lookup only (< 5ms)
  Streaming path: source event → Flink transform → feature-ingestion-stream → Redis

API Design

Present online retrieval as a batch GET by entity ID, emphasizing MGET for <5ms p99 and local SDK caching. Offline APIs return historical snapshots keyed by event timestamp so training joins stay leakage-safe.

Feature Management and Retrieval Endpoints

HTTP
# Feature definition
POST /api/features
{
  "name": "user_avg_purchase_30d",
  "entity": "user", "value_type": "FLOAT",
  "description": "Average purchase amount in last 30 days",
  "source": "transactions_table",
  "freshness_sla": "1h"
}

# Online serving
GET /api/features/online?entity=user&entity_id=123,456&features=avg_purchase_30d,click_rate_7d

# Training data generation (batch)
POST /api/features/historical
{
  "entity": "user",
  "entity_ids_with_timestamps": [
    {"entity_id": "123", "timestamp": "2026-03-01T10:00:00Z"}
  ],
  "features": ["avg_purchase_30d", "click_rate_7d"]
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Online Store (Redis)

HSET user:123
  avg_purchase_30d "45.20"
  click_rate_7d "0.032"
  last_login_hours "2.5"
  _updated_at "2026-03-14T10:05:00Z"
EXPIRE user:123 86400

Offline Store (S3 / Delta Lake / Parquet)

Path: s3://feature-store/user/avg_purchase_30d/
Partitioned by date: date=2026-03-14/part-00000.parquet
Schema: entity_id, timestamp, value, created_at

Feature Registry (PostgreSQL)

SQL
CREATE TABLE features (
    feature_id    UUID PRIMARY KEY,
    name          TEXT UNIQUE NOT NULL,
    entity_type   TEXT NOT NULL,
    value_type    TEXT NOT NULL,
    description   TEXT,
    transformation_code TEXT,
    freshness_sla INTERVAL,
    owner_team    TEXT,
    deprecated    BOOLEAN DEFAULT FALSE
);

Fault Tolerance

Online Store Failure

Redis cluster failure → feature retrieval fails → inference fails

Mitigations:
1. Redis Cluster with replicas (6-node: 3 masters + 3 replicas)
2. Local feature cache on inference servers (LRU, 5-min TTL)
3. Default values: if Redis unavailable, use population median for each feature
   (degrades accuracy, preserves availability)
4. Feature importance: top 10 features provide 80% of model accuracy
   → cache only critical features locally

Additional Considerations

Popular Feature Store Systems

SystemTypeKey Trait
FeastOpen sourceMost popular OSS, Redis/DynamoDB + S3/BigQuery
TectonSaaSFounded by Uber Michelangelo team
HopsworksOpen sourceFeature pipeline as first-class citizen
Vertex AIGCP managedIntegrated with Vertex AI
SageMakerAWS managedIntegrated with SageMaker

When You Need a Feature Store

DON'T need: When managing fewer than 10 features for a single model and team, inline feature computation is sufficient. For batch-only machine learning, standard SQL views are typically enough.

DO need: Real-time serving under 10ms latency, multiple teams sharing features across models, ongoing training-serving skew issues, strict point-in-time correctness, or continuous streaming features.

Feast API Example

PYTHON
# features.py: define feature view
from feast import Entity, FeatureView, Field, FileSource
from feast.types import Float32

user = Entity(name="user_id", join_keys=["user_id"])

user_stats = FeatureView(
    name="user_purchase_stats",
    entities=[user],
    schema=[Field(name="total_purchases_7d", dtype=Float32)],
    source=FileSource(path="s3://features/user_stats.parquet"),
    ttl=timedelta(days=1),
)

# Training (offline: point-in-time join)
training_df = store.get_historical_features(
    entity_df=labels_df,  # (user_id, event_timestamp, label)
    features=["user_purchase_stats:total_purchases_7d"],
).to_df()

# Serving (online: sub-5ms Redis lookup)
online_features = store.get_online_features(
    features=["user_purchase_stats:total_purchases_7d"],
    entity_rows=[{"user_id": "usr_123"}],
).to_dict()

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • offline store and online store duality (5 min)
    • point-in-time correct joins (6 min)
    • feature versioning and lineage for reproducible training (5 min)
    • materialization pipeline from batch compute to online key-value serving (5 min)
    • training-serving skew detection and monitoring (4 min)
  • Explain the duality between the offline store for batch training on historical snapshots and the online store for low-latency serving.
  • Cover point-in-time correct joins, ensuring features reflect state at prediction time rather than the latest value.
  • Discuss feature versioning and lineage for reproducible training runs and governance.
  • Mention the materialization pipeline: running batch computations and pushing feature records to an online key-value store for serving.
  • Cover training-serving skew detection as a first-class monitoring concern with drift and stability metrics.
  • Common pitfall: online serving reads the latest feature value while training used historical snapshots, causing silent model degradation.
  • Downstream integration: features produced here feed directly into real-time prediction pipelines such as Ad Click Prediction and Real-Time Bidding System.

Engineering Trade-offs

Online Store: Redis vs DynamoDB vs Aerospike

Redis (most common choice):
  ✓ Sub-millisecond latency (memory-only, < 0.5ms p99)
  ✓ Flexible data types (Hash, String, List)
  ✓ TTL natively (features expire automatically)
  ✗ Memory-only: 10M users x 50 features x 50B = 25 GB RAM
  Best for: < 100M active users

DynamoDB (AWS managed):
  ✓ Auto-scaling, zero ops overhead
  ✗ Latency: 5-10ms without DAX
  Best for: AWS ecosystem, medium scale

Aerospike (Uber, Criteo, Twitter):
  ✓ NVMe SSD-optimized, < 1ms latency
  ✓ 10-100x cheaper than Redis for same data volume
  ✗ More complex to operate
  Best for: 1B+ entities where RAM is too expensive

Training-Serving Skew: The Silent Killer

Common causes of training-serving skew:

1. Different code paths:
   Training: SQL in Spark
   Serving: Python application code with different time window
   Fix: Feature Store shares ONE feature definition

2. Timestamp handling:
   Training: uses order creation time
   Serving: uses current time
   Fix: point-in-time joins (as-of join on timestamp)

3. Missing value handling:
   Training: NaN → fill with 0
   Serving: missing key → return None (not 0)
   Fix: feature store enforces consistent default values

4. Type mismatches:
   Training: integer feature [1, 2, 3]
   Serving: string feature ["1", "2", "3"]
   Fix: feature store enforces schema with type validation

Monitoring: Log feature values at serving time → compare distribution to training
KS test between training and serving distributions → alert on significant shift

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...