System Design Problem

Design a Feature Flag System

Commonly Asked By:LaunchDarklyNetflixMetaGoogle

Interview Setup

Interview Prompt

Design a feature flag system serving 2,000 active flags across 100,000 SDK instances with ~12K evaluations/sec per instance, supporting percentage rollouts, user segmentation, and instant kill switches.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Server-side evaluation or client-side SDK?
  • Client-side = zero latency but stale config
  • server-side = fresh but adds network hop per eval.
Percentage rollout or explicit user targeting?
  • hash(user_id) % 100 for sticky percentage
  • explicit lists for beta testers.
Kill switch propagation SLA — seconds or minutes?A bad deploy needs flag flip to reach all 100K instances within seconds.
Flag dependencies (flag B requires flag A enabled)?2K active flags with dependencies need evaluation order and conflict detection.

Scope

In scope

  • Flag evaluation engine
  • Percentage rollouts
  • User segmentation
  • Kill switches
  • Flag dependency graph
  • Audit trail

Out of scope (state explicitly)

  • Detailed frontend/UI pixel implementation
  • Org structure, staffing, and hiring plan

Functional Requirements

Start by asking your interviewer what flag types and targeting rules are in scope. Confirm kill-switch latency, percentage rollout, and whether SDK offline resilience is required.

  • Create, update, and delete feature flags (boolean, string, number, JSON variants)
  • Target flags by user ID, user attributes (country, plan, device), percentage rollout
  • Gradual rollout: 1% → 5% → 25% → 50% → 100% (with rollback at any point)
  • A/B testing integration: assign users to experiment variants deterministically
  • Kill switch: instantly disable a feature globally in < 5 seconds
  • Flag dependencies: flag B requires flag A to be enabled
  • Audit log: who changed what flag, when, and why
  • SDK support: server-side (Java, Go, Python) and client-side (JS, iOS, Android)
  • Environment separation: dev, staging, prod with independent flag states
  • Scheduled flags: auto-enable at a specific time (launch events)

Non-Functional Requirements

Your interviewer will care most about sub-microsecond local evaluation and kill-switch propagation under 5 seconds. Flag checks run on every request — they must never add a network hop on the hot path.

  • Ultra-Low Latency: Flag evaluation in < 1µs (local in-memory, no network call)
  • High Availability: 99.999%: flag evaluation must never fail (default fallback)
  • Consistency: Flag change propagated to all servers within 10 seconds
  • Scalability: 10K+ flags, 1B+ evaluations/day
  • Resilience: SDK works offline / when flag service is down (cached state)
  • Zero Performance Impact: No measurable overhead in the hot path

Capacity Estimations

Run this math before you size relay servers. Flag count and evaluations per day are cheap; SSE connection fan-out from config pushes drives relay cluster sizing.

MetricCalculationValue
Feature flags (total)Given10,000
Active flagsGiven2,000
Flag evaluations / secDerived from daily volume ÷ 86400 (+ peak factor)~12K/sec per instance x 100 instances
Flag changes / day~50 ÷ 86400~50
SDK instancesGiven100,000
Flag definition payloadGiven~50 KB
Streaming update bandwidthGiven~1.7 MB/sec

Architecture Diagram

In the room: SDK local evaluation in <1ms — relay pushes config via SSE; never hit the database per flag check on the hot path.

Walk your interviewer through the local-eval vs push-refresh split. SDKs evaluate flags locally from an in-memory config cache, refreshed via SSE push from regional relay servers when flag definitions change in the admin API — I draw evaluation as a pure in-process hash lookup with no network on the hot path.

Loading...

Component Deep Dives

Next we walk through each box on the diagram. I start with the flag relay because 100K SDKs polling the database directly would kill it — relay multiplexes one Kafka consumer into SSE pushes.

The relay is the scaling bottleneck for config propagation — explain regional deployment and ~10K SSE connections per relay instance.

Flag Relay Service

Why a dedicated relay? Direct DB polling from 100K SDKs = 100K queries/sec → DB dies. Relay multiplexes: 1 Kafka consumer → 100K SSE pushes.

Scaling: 1 relay handles ~10K SSE connections. 10 relays per region.

Regional deployment: us-east, eu-west, ap-south → low latency push.

Deterministic MurmurHash rollout is the core algorithm — users must never flip between on and off when you increase percentage.

Deterministic Percentage Rollout (Key Algorithm)

bucket = murmurhash3(flag_key + ":" + user_id) % 100

rollout_percent = 25 → bucket < 25 → ON

Monotonic increase:
  25% → 50%: users 0-24 STILL ON, users 25-49 NOW ON
  Nobody loses access — only gains

Why MurmurHash3: fast (~5ns), uniform distribution, deterministic

Multi-Variant Experiments

Flag: "checkout_layout"
  Variant A: "single_page" (33%), Variant B: "multi_step" (33%), Variant C: "wizard" (34%)

Mutual exclusion via experiment layers:
  Layer 1 (checkout): hash1(user_id) % 100 < 50
  Layer 2 (pricing): hash2(user_id) % 100 >= 50
  Different hash seed per layer ensures non-correlation

Event Bus Design (Kafka)

Topic: flag-changes
  Partitions: 16 (low volume ~50 flag edits/day; headroom for bursts)
  Partition key: flag_key (all variants for one flag stay ordered)
  Retention: 30 days (audit + rollback replay)
  Producers: Flag Management API after PostgreSQL commit
  Consumers: Flag Relay Service (SSE push to 100K SDKs), analytics pipeline, audit log

Topic: flag-evaluation-samples
  Partitions: 64 (partition by flag_key)
  Sampling: 1% of evaluations for experimentation dashboards
  Consumers: A/B metrics aggregator (variant exposure counts)

Admin edit path: validate rules → PostgreSQL → invalidate Redis → publish flag-changes → 200
  Relay pushes full flag snapshot to connected SDKs; disconnected SDKs poll every 30s

API Design

POST   /api/flags              → Create flag with targeting rules
GET    /api/flags               → List all flags (paginated)
GET    /api/flags/{key}         → Flag details with all environments
PUT    /api/flags/{key}         → Update flag (creates audit entry)
DELETE /api/flags/{key}         → Archive flag (soft delete)
POST   /api/flags/{key}/toggle  → Kill switch enable/disable
GET    /api/flags/{key}/audit   → Audit log with diffs
POST   /api/flags/{key}/schedule → Schedule future enable/disable

# SDK endpoints
GET    /api/sdk/flags?env=prod   → Full flag definitions
GET    /api/sdk/stream?env=prod  → SSE stream of changes
POST   /api/sdk/evaluate         → Server-side evaluation for client SDKs

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

PostgreSQL: Source of Truth

Why PostgreSQL? ACID for metadata, JSONB for flexible rules, reliable for low-write workload (~50 writes/day).

SQL
CREATE TABLE feature_flags (
    flag_id     UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    flag_key    TEXT UNIQUE NOT NULL,
    name        TEXT NOT NULL,
    description TEXT,
    flag_type   TEXT NOT NULL CHECK (flag_type IN ('boolean','string','number','json')),
    created_by  UUID, created_at TIMESTAMPTZ DEFAULT NOW(),
    archived    BOOLEAN DEFAULT FALSE
);

CREATE TABLE flag_environments (
    flag_id         UUID REFERENCES feature_flags(flag_id),
    environment     TEXT NOT NULL,
    enabled         BOOLEAN DEFAULT FALSE,
    default_variant JSONB,
    fallthrough     JSONB,
    rules           JSONB,
    version         BIGINT DEFAULT 1,
    updated_at      TIMESTAMPTZ DEFAULT NOW(),
    updated_by      UUID,
    PRIMARY KEY (flag_id, environment)
);

CREATE TABLE flag_audit_log (
    audit_id   UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    flag_id    UUID NOT NULL,
    flag_key   TEXT NOT NULL,
    environment TEXT,
    action     TEXT NOT NULL,
    old_value  JSONB,
    new_value  JSONB,
    changed_by UUID NOT NULL,
    reason     TEXT,
    changed_at TIMESTAMPTZ DEFAULT NOW()
);

Redis Cache

SET flags:production '{...all flags...}' EX 300
SET flag_version:production 42

SDK In-Memory (Atomic Pointer Swap)

JAVA
class FlagStore {
    volatile Map<String, FlagDef> flags;
    FlagDef get(String key) { return flags.get(key); }
    void update(Map<String, FlagDef> n) { this.flags = Map.copyOf(n); }
}

Fault Tolerance

TechniqueApplication
SDK disk cache
  • Survives restart
  • fallback when relay unreachable
Default valuesFlag not found → hardcoded default
Streaming + pollingSSE primary, 30s poll backup, disk cache last resort
PostgreSQL replicasRead replicas for API reads
Kafka RF=3Change events survive broker failure
Relay redundancy
  • Multiple instances per region
  • SDK reconnects on failure

SDK Resilience

Startup: disk cache → relay fetch → defaults. Runtime: streaming → polling → cache. Evaluation NEVER makes a network call.

Concurrent Flag Updates

Optimistic locking: UPDATE ... WHERE version = $expected. Conflict → 409, UI shows diff.

Relay Failure

SDKs detect disconnect → reconnect to another relay → send last_seen_version → receive delta. If ALL relays down → poll API. If API down → use cached flags.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • SDK local eval < 1ms (8 min)
    • Config push via SSE from relay (9 min)
    • Disk cache survives relay restart (8 min)
  • Separate admin plane (flag definitions in PostgreSQL) from data plane (SDK local evaluation) — relay pushes snapshots; the hot path never hits the DB per request.
  • Deterministic percentage rollout: murmurhash3(flag_key + ":" + user_id) % 100 — monotonic increase means users never lose access when rollout expands from 25% to 50%.
  • SDK evaluates flags locally with zero network calls on the hot path — relay/SSE pushes updates; disk cache → relay → defaults is the fallback chain.
  • Flag relay multiplexes 100K SDK connections through Kafka consumers — direct DB polling from every SDK instance kills the database.
  • Experiment layers with independent hash seeds prevent correlated flag assignments across overlapping experiments.
  • Optimistic locking on flag updates (WHERE version = $expected) prevents silent overwrites during concurrent admin edits.
  • Common pitfall: random per-request rollout — users flip between variants on every page load, breaking experiments and eroding trust in gradual releases.

Engineering Trade-offs

Feature flags trade evaluation speed against config freshness — local SDK cache vs server-side evaluation.

Full Snapshot Push vs Delta Updates

Full Snapshot (this design, LaunchDarkly approach):
  Every update: relay sends complete set of all flag definitions (~50 KB)
  SDK replaces entire in-memory map atomically
  ✓ Always consistent
  ✓ Recovery is trivial
  ✗ Bandwidth: 100K SDKs x 50 KB x 50 updates/day = 250 GB/day

Delta Updates:
  Only send the changed flag definition (~1 KB)
  ✓ 50x less bandwidth
  ✗ Ordering matters: missed delta → state diverges
  ✗ Need sequence numbers and gap detection

Recommendation: Full snapshot (simplicity + consistency > bandwidth)
  50 KB compressed to ~15 KB with gzip → 75 GB/day → trivial at scale

In-Process SDK vs Remote Evaluation API

In-Process SDK (recommended):
  Latency: ~100 ns, Availability: 100%, works offline
  ✗ Flag definitions exposed, SDK must be updated

Remote Evaluation API:
  Latency: 5-20 ms, Availability: depends on service
  ✓ Secure, no SDK per language
  ✗ Network failure → flags stop working

Production pattern:
  Server-side: in-process SDK → 0ms latency
  Client-side: server evaluates on behalf of client

Deterministic Hash vs Server-Assigned Cohort

Deterministic Hash (MurmurHash3):
  ✓ Stateless, no storage, monotonic rollout
  ✗ Can't manually override specific users

Server-Assigned Cohort:
  ✓ Full control, stable assignment
  ✗ Requires DB lookup, storage expensive

Recommendation: Deterministic hash for most flags.
  Server-assigned only for formal A/B experiments.

Consistency Model: Bounded Eventual Consistency

Timeline of a flag change:
  T+0s:     Admin clicks "Enable flag"
  T+0.1s:   API writes to PostgreSQL
  T+0.2s:   API publishes to Kafka
  T+0.5s:   Relay Service consumes from Kafka
  T+1s:     Relay pushes update via SSE
  T+5s:     All healthy SDKs have received update
  T+10s:    Reconnecting SDKs have received update (poll fallback)

CAP analysis: AP system by design.
  Network partition → SDKs serve stale cached flags (availability over consistency)
  Flags are NOT a source of truth for critical business logic.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...