System Design Problem

Design an A/B Testing and Experimentation Platform

Commonly Asked By:NetflixMetaOptimizelyGoogle

Interview Setup

Interview Prompt

Design an A/B testing platform for 500M users running 1,000 concurrent experiments, with deterministic variant assignment at 500K calls/sec and 50B experiment events/day for statistical analysis.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Sticky assignment per user or per session? Can users switch variants mid-experiment?
  • Deterministic hashing on user_id ensures consistent experience
  • session-based breaks statistical validity.
Real-time significance or batch daily reports?
  • 50B events/day → 5 TB
  • real-time needs streaming aggregation
  • batch is simpler and cheaper.
How many concurrent experiments can overlap on the same user?1,000 experiments with interaction effects need factorial design or mutual exclusion groups.
Guardrail metrics: auto-stop bad variants or alert only?A checkout experiment degrading revenue 5% needs auto-pause within minutes, not next day's report.

Scope

In scope

  • Experiment assignment (deterministic hashing)
  • Metric pipeline
  • Statistical significance engine
  • Interaction effects
  • Guardrail metrics
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Detailed frontend/UI pixel implementation
  • Org structure, staffing, and hiring plan

Functional Requirements

Start by asking your interviewer what belongs in scope. For an A/B platform, confirm experiment assignment, guardrail metrics, and whether statistical significance testing is in-band for this round.

  • Create experiments: Define experiment with name, hypothesis, variants, and traffic allocation
  • User assignment: Deterministically assign users to experiment variants
  • Feature flags: Toggle features on and off for gradual rollouts, integrating with our Feature Flag System.
  • Metric tracking: Track conversion rates, revenue, engagement metrics per variant
  • Statistical analysis: Compute p-value, confidence interval, statistical significance
  • Mutual exclusion: Prevent conflicting experiments from overlapping
  • Experiment lifecycle: Draft → Running → Paused → Completed → Archived
  • Guardrail metrics: Auto-stop experiment if key metrics degrade
  • Segmentation: Run experiments on specific user segments

Non-Functional Requirements

Your interviewer will care most about sticky assignment and zero production impact. Variant lookup must stay under 5 ms with no Kafka on the read path; analytics lag is acceptable, but incorrect bucketing is not.

  • Low Latency: Variant assignment in < 5 ms
  • Consistency: Same user always sees same variant
  • Scale: 1000+ concurrent experiments, 500M+ users, 50B+ events/day
  • Statistical Rigor: Correct p-values; account for multiple comparisons
  • No Impact on Production: Experiment infrastructure must not slow down main services
  • Availability: 99.99% for assignment; analytics can tolerate minutes of lag

Capacity Estimations

Run this math before you split assignment from analytics. Assignment QPS and events per day tell you why the read path is pure CPU hashing while Kafka handles 50B events/day asynchronously.

MetricCalculationValue
Concurrent experimentsGiven1,000
UsersGiven500M
Variant assignment calls / secDerived from daily volume ÷ 86400 (+ peak factor)500K
Experiment events / dayGiven50B
Event data storage / day50B events x ~100 B5 TB

Architecture Diagram

In the room: draw assignment (sync hash, <5ms) separately from analytics (async Kafka), ensuring the hot path never blocks on event logging.

Walk your interviewer through the assignment vs analytics split first. Every page load triggers a synchronous variant lookup: the client SDK hashes user_id + experiment_id to a sticky bucket and returns config in under 5 ms without Kafka on the read path. Exposure and conversion events fire asynchronously into Kafka, where Flink pre-aggregates before ClickHouse stores them for significance testing and guardrail monitoring.

In the interview, draw these two paths separately: assignment must never block production traffic, while the analytics pipeline can tolerate minutes of lag. At 500K assignment calls/sec and 50B events/day (~5 TB), the diagram below shows how config flows down and events flow up without conflating the hot and cold paths.

Loading...

Component Deep Dives

Next we examine each component in the architecture. We start with deterministic assignment because sticky bucketing serves as the foundation of statistical validity; if users flip variants on refresh, the experiment becomes meaningless.

The platform splits into a read path (variant assignment on every page load) and a write path (event ingestion for metrics). Assignment is pure CPU computation through a deterministic hash with no network round-trip on the critical path, while the event pipeline handles 50B records/day through Kafka and Flink. The sections below walk through stickiness guarantees, experiment layering, statistical rigor, and the async bus that feeds guardrail auto-pause.

Deterministic User Assignment

Sticky assignment is the foundation of statistical validity: if the same user sees different variants on refresh, conversion lift becomes meaningless. MurmurHash3 on experiment_id:user_id gives uniform, reproducible bucketing in under 100 ns without per-request randomness or assignment database lookups.

PYTHON
User must ALWAYS see the same variant. No randomness per request.

Algorithm: hash-based assignment

  def get_variant(user_id, experiment_id, variants, traffic_percent):
      hash_input = f"{experiment_id}:{user_id}"
      # Gate: is this user in the experiment at all?
      if murmurhash3(hash_input) % 10000 >= traffic_percent * 100:
          return None
      
      # Independent hash for the variant split so it's uniform
      # across the in-experiment population (not just the gated range).
      variant_hash = murmurhash3(hash_input + ":variant") % 10000
      cumulative = 0
      for variant_name, weight in variants:
          cumulative += weight * 100
          if variant_hash < cumulative:
              return variant_name
      
      return variants[-1][0]

Properties:
  - Deterministic: same user + experiment -> same variant
  - Uniform: murmurhash3 is uniformly distributed
  - Independent: adding/removing experiments doesn't change other assignments
  - Fast: murmurhash3 is < 100 ns

Mutual Exclusion (Experiment Layers)

With 1,000 concurrent experiments, overlapping tests on the same surface corrupt attribution, making it impossible to determine whether the green button or the new layout drove the conversion lift. Experiment layers partition the hash space so each user enters at most one experiment per layer, while independent layers (UI vs backend) can run in parallel without interaction.

Problem: Experiment A tests checkout button color. Experiment B tests checkout page layout.
If same user is in both: which caused the conversion improvement?

Solution: Experiment Layers (Google's Overlapping Experiment Infrastructure)

  Layer 1 (UI experiments):
    Experiment A: button color (control: blue, treatment: green)
    Experiment B: header layout (control: v1, treatment: v2)
    Within a layer: user is in AT MOST one experiment
    
  Layer 2 (Backend experiments):
    Experiment C: recommendation algorithm (control: v1, treatment: v2)
    Independent of Layer 1

Implementation:
  Each experiment belongs to a layer.
  Assignment: hash(user_id + layer_id) % total_traffic
  Each experiment "owns" a non-overlapping range of hash values.

Statistical Analysis Engine

Raw event counts are not experiment results; product teams need lift, confidence intervals, and guardrail checks before shipping. An hourly batch job on ClickHouse aggregates runs two-proportion z-tests, applies multiple-comparison correction, and auto-pauses variants that degrade crash rate or latency beyond 2σ.

1. Conversion rate per variant
2. Relative lift = (treatment_rate - control_rate) / control_rate
3. Statistical significance (two-proportion z-test):
   p_pooled = total_conversions / total_impressions
   se = sqrt(p_pooled * (1-p_pooled) * (1/n_control + 1/n_treatment))
   z = (treatment_rate - control_rate) / se
   p_value = 2 * (1 - norm_cdf(abs(z)))
4. Confidence interval (95%): CI = diff ± 1.96 * se
5. Multiple comparison correction: Bonferroni or Benjamini-Hochberg
6. Guardrail metrics: auto-pause experiment if degraded > 2 std dev

Event Bus Design (Kafka)

The event bus decouples fire-and-forget SDK telemetry from analytics compute, leveraging the stream processing patterns detailed in Stream Processing Basics. Partitioning by user_idco-locates all events for one user so Flink can deduplicate exposures and conversions within tumbling windows before inserting pre-aggregated rows into ClickHouse, ensuring that raw 50B events/day never hit the OLAP store directly.

Topic: experiment-events
  Partitions: 640 (50B events/day to scale Flink horizontally)
  Partition key: user_id (co-locate all events for one user)
  Retention: 7 days (replay for ClickHouse backfill)
  Replication factor: 3, min.insync.replicas: 2

Producer: Client SDK (idempotent producer, fire-and-forget)
  Event: { event_id, user_id, experiment_id, variant, event_type: "exposure"|"conversion", value, timestamp }

Consumer groups:
  1. metric-aggregator: Flink dedup by (user_id, event_type, experiment_id, timestamp_bucket)
     -> tumbling 1-min windows -> ClickHouse materialized views
  2. guardrail-monitor: crash_rate, p99_latency, error_rate per variant -> auto-pause if > 2σ
  3. srm-detector: chi-squared test on assignment counts (expected 50/50 vs actual)

Sync path: variant assignment is deterministic hash; no Kafka on read path (< 5ms)
Async path: 50B events/day (~500K/sec avg) -> Flink pre-aggregate before ClickHouse insert
DLQ: experiment-events-dlq; alert when consumer lag > 60s

API Design

HTTP
POST /api/v1/experiments
{
  "name": "checkout_button_color",
  "hypothesis": "Green button increases conversion by 5%",
  "layer": "checkout_ui",
  "traffic_percent": 20,
  "variants": [
    { "name": "control", "weight": 50, "config": { "button_color": "blue" } },
    { "name": "treatment", "weight": 50, "config": { "button_color": "green" } }
  ],
  "primary_metric": "checkout_conversion_rate",
  "guardrail_metrics": ["crash_rate", "p95_latency"]
}

GET /api/v1/experiments/{id}/assignment?user_id=user-uuid
-> { "variant": "treatment", "config": { "button_color": "green" } }

POST /api/v1/experiments/{id}/events
{ "user_id": "user-uuid", "event_type": "conversion", "value": 1 }

GET /api/v1/experiments/{id}/results
-> {
  "status": "running", "days_running": 7,
  "variants": {
    "control": { "impressions": 50000, "conversions": 2500, "rate": 0.0500 },
    "treatment": { "impressions": 49800, "conversions": 2750, "rate": 0.0552 }
  },
  "lift": 0.104, "p_value": 0.0023, "significant": true,
  "confidence_interval": [0.032, 0.176],
  "recommended_action": "Ship treatment (statistically significant improvement)"
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

PostgreSQL: Experiment Configuration

SQL
CREATE TABLE experiments (
    experiment_id UUID PRIMARY KEY, name VARCHAR(100), hypothesis TEXT,
    layer VARCHAR(50), traffic_percent DECIMAL(5,2),
    primary_metric VARCHAR(100), guardrail_metrics JSONB,
    status ENUM('draft','running','paused','completed','archived'),
    started_at TIMESTAMPTZ, ended_at TIMESTAMPTZ,
    owner VARCHAR(100), created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE experiment_variants (
    variant_id UUID PRIMARY KEY, experiment_id UUID NOT NULL,
    name VARCHAR(50), weight INT, config JSONB
);

Redis: Fast Assignment

exp_config:{experiment_id}  -> JSON (experiment config)
TTL: 60 (refreshed from PostgreSQL)

active_experiments:{layer}  -> LIST of experiment_ids
TTL: 60

ClickHouse: Event Analytics

SQL
CREATE TABLE experiment_events (
    experiment_id UUID, variant String, user_id UUID,
    event_type String, value Float64,
    timestamp DateTime, date Date MATERIALIZED toDate(timestamp)
) ENGINE = MergeTree() PARTITION BY toYYYYMM(timestamp)
  ORDER BY (experiment_id, variant, timestamp);

Fault Tolerance

ConcernSolution
Assignment service down
  • SDK caches last assignment locally
  • fall back to control variant
Event pipeline lag
  • ClickHouse backfill from Kafka replay
  • results delayed but not lost
Experiment degrades metrics
  • Guardrail auto-pause
  • manual kill switch
Hash collision causing uneven split
  • Chi-squared test on assignment counts
  • alert if > 2% deviation
Novelty effect
  • Run experiments for minimum 2 weeks
  • track metrics over time

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Hash user_id to sticky bucket for deterministic assignment (8 min)
    • Assignment sync path in < 5ms without Kafka on the read path (9 min)
    • Exposure and conversion events async to Kafka (8 min)
  • Split the problem into two paths: synchronous variant assignment (< 5 ms, must never block production) and asynchronous event analytics (50B+ events/day via Kafka to ClickHouse).
  • Lead with hash-based deterministic assignment: murmurhash3(experiment_id:user_id) guarantees the same user always sees the same variant with no per-request randomness.
  • Introduce experiment layers early to handle mutual exclusion, ensuring that overlapping UI tests on checkout button color vs page layout do not share the same user pool.
  • Walk through the statistical pipeline: conversion rate, relative lift, two-proportion z-test, and Bonferroni correction for multiple comparisons.
  • Cover guardrail metrics with auto-pause: if crash_rate or p95_latency degrades beyond 2 standard deviations, stop the experiment before it damages production.
  • Detect Sample Ratio Mismatch (SRM) with chi-squared tests on assignment counts, because an uneven split invalidates the entire experiment.
  • Common pitfall: checking p-values daily and stopping when significant, where the peeking problem inflates false positives from 5% to 20% to 30%.

Engineering Trade-offs

Peeking Problem

A/B platforms trade assignment stickiness against experiment velocity, balancing deterministic hashing against dynamic re-bucketing.

Problem: checking p-value daily and stopping when p < 0.05.
  This inflates false positive rate from 5% to 20-30%!
  
Why? Statistical tests assume you look at data ONCE at predetermined sample size.

Solution:
  1. Pre-determine sample size. Don't look until reached.
  2. Sequential testing with adjusted boundaries (O'Brien-Fleming).
  3. Bayesian approach: compute posterior probability continuously.

Sample Ratio Mismatch (SRM) Detection

Experiment configured for 50/50 split. After 1 week:
  Control: 502,000 users. Treatment: 498,000 users.

Chi-squared test:
  Expected: 500,000 each. Observed: 502,000 / 498,000.
  chi2 = 16.0, p-value < 0.001 -> SIGNIFICANT MISMATCH

Causes: bot traffic filtered differently, treatment causes more crashes,
  redirect-based experiment with slower redirect, hash function not uniform.

Impact: SRM invalidates the experiment. Results cannot be trusted.
Action: investigate root cause. Fix. Re-run experiment.

Network Effects: When User Independence Breaks

A/B tests assume: user A's behavior is independent of user B's assignment.
This breaks with network effects (social features).

Solutions:
  1. Cluster randomization: assign entire friend clusters to same variant
  2. Geo-based randomization: assign entire cities to variants
  3. Time-based (switchback): alternate treatment by time period

For social platforms: cluster or geo randomization is necessary.
For independent features (checkout UI): standard user-level randomization is fine.

Feature Flags vs A/B Tests

While both control feature delivery, feature flags focus on operational release management, as detailed in our Feature Flag System, whereas A/B testing platforms prioritize statistical inference and experiment rigor.

Feature flag: binary (on/off). Used for gradual rollout, kill switches.
  No statistical analysis needed. Just monitoring.

A/B test: controlled experiment with statistical rigor.
  Requires control group, sufficient sample size, statistical analysis.

In practice: same infrastructure serves both.
  Feature flag = experiment with 1 variant + 100% traffic + no metrics.
  A/B test = experiment with 2+ variants + metrics + statistical analysis.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...