Interview Setup
Interview Prompt
Design a feature flag system serving 2,000 active flags across 100,000 SDK instances with ~12K evaluations/sec per instance, supporting percentage rollouts, user segmentation, and instant kill switches.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Server-side evaluation or client-side SDK? |
|
| Percentage rollout or explicit user targeting? |
|
| Kill switch propagation SLA — seconds or minutes? | A bad deploy needs flag flip to reach all 100K instances within seconds. |
| Flag dependencies (flag B requires flag A enabled)? | 2K active flags with dependencies need evaluation order and conflict detection. |
Scope
In scope
- Flag evaluation engine
- Percentage rollouts
- User segmentation
- Kill switches
- Flag dependency graph
- Audit trail
Out of scope (state explicitly)
- Detailed frontend/UI pixel implementation
- Org structure, staffing, and hiring plan
Functional Requirements
Start by asking your interviewer what flag types and targeting rules are in scope. Confirm kill-switch latency, percentage rollout, and whether SDK offline resilience is required.
- Create, update, and delete feature flags (boolean, string, number, JSON variants)
- Target flags by user ID, user attributes (country, plan, device), percentage rollout
- Gradual rollout: 1% → 5% → 25% → 50% → 100% (with rollback at any point)
- A/B testing integration: assign users to experiment variants deterministically
- Kill switch: instantly disable a feature globally in < 5 seconds
- Flag dependencies: flag B requires flag A to be enabled
- Audit log: who changed what flag, when, and why
- SDK support: server-side (Java, Go, Python) and client-side (JS, iOS, Android)
- Environment separation: dev, staging, prod with independent flag states
- Scheduled flags: auto-enable at a specific time (launch events)
Non-Functional Requirements
Your interviewer will care most about sub-microsecond local evaluation and kill-switch propagation under 5 seconds. Flag checks run on every request — they must never add a network hop on the hot path.
- Ultra-Low Latency: Flag evaluation in < 1µs (local in-memory, no network call)
- High Availability: 99.999%: flag evaluation must never fail (default fallback)
- Consistency: Flag change propagated to all servers within 10 seconds
- Scalability: 10K+ flags, 1B+ evaluations/day
- Resilience: SDK works offline / when flag service is down (cached state)
- Zero Performance Impact: No measurable overhead in the hot path
Capacity Estimations
Run this math before you size relay servers. Flag count and evaluations per day are cheap; SSE connection fan-out from config pushes drives relay cluster sizing.
| Metric | Calculation | Value |
|---|---|---|
| Feature flags (total) | Given | 10,000 |
| Active flags | Given | 2,000 |
| Flag evaluations / sec | Derived from daily volume ÷ 86400 (+ peak factor) | ~12K/sec per instance x 100 instances |
| Flag changes / day | ~50 ÷ 86400 | ~50 |
| SDK instances | Given | 100,000 |
| Flag definition payload | Given | ~50 KB |
| Streaming update bandwidth | Given | ~1.7 MB/sec |
Architecture Diagram
In the room: SDK local evaluation in <1ms — relay pushes config via SSE; never hit the database per flag check on the hot path.
Walk your interviewer through the local-eval vs push-refresh split. SDKs evaluate flags locally from an in-memory config cache, refreshed via SSE push from regional relay servers when flag definitions change in the admin API — I draw evaluation as a pure in-process hash lookup with no network on the hot path.
Component Deep Dives
Next we walk through each box on the diagram. I start with the flag relay because 100K SDKs polling the database directly would kill it — relay multiplexes one Kafka consumer into SSE pushes.
The relay is the scaling bottleneck for config propagation — explain regional deployment and ~10K SSE connections per relay instance.
Flag Relay Service
Why a dedicated relay? Direct DB polling from 100K SDKs = 100K queries/sec → DB dies. Relay multiplexes: 1 Kafka consumer → 100K SSE pushes.
Scaling: 1 relay handles ~10K SSE connections. 10 relays per region.
Regional deployment: us-east, eu-west, ap-south → low latency push.
Deterministic MurmurHash rollout is the core algorithm — users must never flip between on and off when you increase percentage.
Deterministic Percentage Rollout (Key Algorithm)
bucket = murmurhash3(flag_key + ":" + user_id) % 100 rollout_percent = 25 → bucket < 25 → ON Monotonic increase: 25% → 50%: users 0-24 STILL ON, users 25-49 NOW ON Nobody loses access — only gains Why MurmurHash3: fast (~5ns), uniform distribution, deterministic
Multi-Variant Experiments
Flag: "checkout_layout" Variant A: "single_page" (33%), Variant B: "multi_step" (33%), Variant C: "wizard" (34%) Mutual exclusion via experiment layers: Layer 1 (checkout): hash1(user_id) % 100 < 50 Layer 2 (pricing): hash2(user_id) % 100 >= 50 Different hash seed per layer ensures non-correlation
Event Bus Design (Kafka)
Topic: flag-changes Partitions: 16 (low volume ~50 flag edits/day; headroom for bursts) Partition key: flag_key (all variants for one flag stay ordered) Retention: 30 days (audit + rollback replay) Producers: Flag Management API after PostgreSQL commit Consumers: Flag Relay Service (SSE push to 100K SDKs), analytics pipeline, audit log Topic: flag-evaluation-samples Partitions: 64 (partition by flag_key) Sampling: 1% of evaluations for experimentation dashboards Consumers: A/B metrics aggregator (variant exposure counts) Admin edit path: validate rules → PostgreSQL → invalidate Redis → publish flag-changes → 200 Relay pushes full flag snapshot to connected SDKs; disconnected SDKs poll every 30s
API Design
POST /api/flags → Create flag with targeting rules
GET /api/flags → List all flags (paginated)
GET /api/flags/{key} → Flag details with all environments
PUT /api/flags/{key} → Update flag (creates audit entry)
DELETE /api/flags/{key} → Archive flag (soft delete)
POST /api/flags/{key}/toggle → Kill switch enable/disable
GET /api/flags/{key}/audit → Audit log with diffs
POST /api/flags/{key}/schedule → Schedule future enable/disable
# SDK endpoints
GET /api/sdk/flags?env=prod → Full flag definitions
GET /api/sdk/stream?env=prod → SSE stream of changes
POST /api/sdk/evaluate → Server-side evaluation for client SDKsCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
PostgreSQL: Source of Truth
Why PostgreSQL? ACID for metadata, JSONB for flexible rules, reliable for low-write workload (~50 writes/day).
CREATE TABLE feature_flags (
flag_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
flag_key TEXT UNIQUE NOT NULL,
name TEXT NOT NULL,
description TEXT,
flag_type TEXT NOT NULL CHECK (flag_type IN ('boolean','string','number','json')),
created_by UUID, created_at TIMESTAMPTZ DEFAULT NOW(),
archived BOOLEAN DEFAULT FALSE
);
CREATE TABLE flag_environments (
flag_id UUID REFERENCES feature_flags(flag_id),
environment TEXT NOT NULL,
enabled BOOLEAN DEFAULT FALSE,
default_variant JSONB,
fallthrough JSONB,
rules JSONB,
version BIGINT DEFAULT 1,
updated_at TIMESTAMPTZ DEFAULT NOW(),
updated_by UUID,
PRIMARY KEY (flag_id, environment)
);
CREATE TABLE flag_audit_log (
audit_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
flag_id UUID NOT NULL,
flag_key TEXT NOT NULL,
environment TEXT,
action TEXT NOT NULL,
old_value JSONB,
new_value JSONB,
changed_by UUID NOT NULL,
reason TEXT,
changed_at TIMESTAMPTZ DEFAULT NOW()
);Redis Cache
SET flags:production '{...all flags...}' EX 300
SET flag_version:production 42SDK In-Memory (Atomic Pointer Swap)
class FlagStore {
volatile Map<String, FlagDef> flags;
FlagDef get(String key) { return flags.get(key); }
void update(Map<String, FlagDef> n) { this.flags = Map.copyOf(n); }
}Fault Tolerance
| Technique | Application |
|---|---|
| SDK disk cache |
|
| Default values | Flag not found → hardcoded default |
| Streaming + polling | SSE primary, 30s poll backup, disk cache last resort |
| PostgreSQL replicas | Read replicas for API reads |
| Kafka RF=3 | Change events survive broker failure |
| Relay redundancy |
|
SDK Resilience
Startup: disk cache → relay fetch → defaults. Runtime: streaming → polling → cache. Evaluation NEVER makes a network call.
Concurrent Flag Updates
Optimistic locking: UPDATE ... WHERE version = $expected. Conflict → 409, UI shows diff.
Relay Failure
SDKs detect disconnect → reconnect to another relay → send last_seen_version → receive delta. If ALL relays down → poll API. If API down → use cached flags.
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- SDK local eval < 1ms (8 min)
- Config push via SSE from relay (9 min)
- Disk cache survives relay restart (8 min)
- Separate admin plane (flag definitions in PostgreSQL) from data plane (SDK local evaluation) — relay pushes snapshots; the hot path never hits the DB per request.
- Deterministic percentage rollout:
murmurhash3(flag_key + ":" + user_id) % 100— monotonic increase means users never lose access when rollout expands from 25% to 50%. - SDK evaluates flags locally with zero network calls on the hot path — relay/SSE pushes updates; disk cache → relay → defaults is the fallback chain.
- Flag relay multiplexes 100K SDK connections through Kafka consumers — direct DB polling from every SDK instance kills the database.
- Experiment layers with independent hash seeds prevent correlated flag assignments across overlapping experiments.
- Optimistic locking on flag updates (
WHERE version = $expected) prevents silent overwrites during concurrent admin edits. - Common pitfall: random per-request rollout — users flip between variants on every page load, breaking experiments and eroding trust in gradual releases.
Engineering Trade-offs
Feature flags trade evaluation speed against config freshness — local SDK cache vs server-side evaluation.
Full Snapshot Push vs Delta Updates
Full Snapshot (this design, LaunchDarkly approach): Every update: relay sends complete set of all flag definitions (~50 KB) SDK replaces entire in-memory map atomically ✓ Always consistent ✓ Recovery is trivial ✗ Bandwidth: 100K SDKs x 50 KB x 50 updates/day = 250 GB/day Delta Updates: Only send the changed flag definition (~1 KB) ✓ 50x less bandwidth ✗ Ordering matters: missed delta → state diverges ✗ Need sequence numbers and gap detection Recommendation: Full snapshot (simplicity + consistency > bandwidth) 50 KB compressed to ~15 KB with gzip → 75 GB/day → trivial at scale
In-Process SDK vs Remote Evaluation API
In-Process SDK (recommended): Latency: ~100 ns, Availability: 100%, works offline ✗ Flag definitions exposed, SDK must be updated Remote Evaluation API: Latency: 5-20 ms, Availability: depends on service ✓ Secure, no SDK per language ✗ Network failure → flags stop working Production pattern: Server-side: in-process SDK → 0ms latency Client-side: server evaluates on behalf of client
Deterministic Hash vs Server-Assigned Cohort
Deterministic Hash (MurmurHash3): ✓ Stateless, no storage, monotonic rollout ✗ Can't manually override specific users Server-Assigned Cohort: ✓ Full control, stable assignment ✗ Requires DB lookup, storage expensive Recommendation: Deterministic hash for most flags. Server-assigned only for formal A/B experiments.
Consistency Model: Bounded Eventual Consistency
Timeline of a flag change: T+0s: Admin clicks "Enable flag" T+0.1s: API writes to PostgreSQL T+0.2s: API publishes to Kafka T+0.5s: Relay Service consumes from Kafka T+1s: Relay pushes update via SSE T+5s: All healthy SDKs have received update T+10s: Reconnecting SDKs have received update (poll fallback) CAP analysis: AP system by design. Network partition → SDKs serve stale cached flags (availability over consistency) Flags are NOT a source of truth for critical business logic.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.