System Design Problem

Design a Circuit Breaker

Commonly Asked By:NetflixMicrosoftGoogleAWS

Interview Setup

Interview Prompt

Design a circuit breaker library/middleware protecting 500 microservices making 10M inter-service RPC calls/sec, with ~5,000 per-service circuit breakers and configurable failure thresholds.

Clarifying Questions (ask before designing)

QuestionWhy it matters
In-process library or sidecar/service-mesh proxy?
  • Library = zero latency overhead but per-language
  • sidecar = universal but adds hop.
Failure rate or slow-call rate as trigger, or both?A service returning 200 in 30s is 'healthy' by error rate but kills caller latency budget.
Per-endpoint or per-service circuit breakers?
  • 5,000 breakers = per-endpoint
  • one bad endpoint shouldn't open circuit for entire service.
Fallback strategies in scope?
  • Cache stale data, default response, or alternate service
  • each needs different circuit breaker config.

Scope

In scope

  • State machine (closed → open → half-open)
  • Failure rate thresholds
  • Bulkhead pattern
  • Fallback strategies
  • Integration with service mesh
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Full service mesh control plane
  • Application business logic in downstream services
  • Building a custom service registry from scratch when managed options exist

Functional Requirements

Start by confirming whether you are designing an in-process library or a service-mesh sidecar. Ask about fallback strategies, per-endpoint breakers, and manual override requirements.

  • Monitor health of downstream service calls (success/failure rates)
  • Automatically stop sending requests to unhealthy services (circuit OPEN)
  • Periodically probe to check if service has recovered (HALF-OPEN state)
  • Resume traffic when service recovers (circuit CLOSED)
  • Configurable thresholds: failure rate, slow-call rate, minimum call volume
  • Support per-service, per-endpoint, per-client circuit breakers
  • Dashboard showing circuit breaker states across all services
  • Integration with service mesh (Envoy/Istio sidecar) or application library
  • Fallback responses when circuit is open (cached response, default value, degraded mode)
  • Manual override: force-open or force-close circuits

Non-Functional Requirements

Your interviewer will care most about sub-microsecond overhead and fast failure detection, as explored in Circuit Breakers, Retries & Bulkheads. The breaker must add almost zero latency in CLOSED state: it is a local counter check, not a network hop.

  • Ultra-Low Overhead: < 1µs per call decision (in-process check)
  • Fast Detection: Detect failures within 5 to 10 seconds
  • Fast Recovery: Resume traffic within seconds of downstream recovery
  • No SPOF: Circuit breaker itself must not become a single point of failure
  • Consistency: All instances of a service should have similar circuit state (eventual)
  • Observability: Emit metrics for open/close transitions, fallback invocations

Capacity Estimations

Run this math before you size the sliding window. Calls per second per service and failure threshold tell you ring buffer size; probe count in HALF-OPEN drives recovery latency.

MetricCalculationValue
Services in productionGiven500
Inter-service RPC calls / secDerived from daily volume ÷ 86400 (+ peak factor)10M
Circuit breakers (per service x endpoint)Given~5,000
State transitions / min (normal)Given< 10
Metric data points / secDerived from daily volume ÷ 86400 (+ peak factor)50K

Architecture Diagram

Walk your interviewer through the three-state machine first. An in-process circuit breaker library wraps calls to downstream Service B. The state machine transitions between CLOSED (normal operation), OPEN (rejecting requests), and HALF-OPEN (probing recovery). Metrics are emitted to Prometheus for dashboarding; I label each transition and its trigger condition on the whiteboard.

Loading...

Component Deep Dives

In the room: draw the CLOSED → OPEN → HALF-OPEN state machine before any boxes: transitions and thresholds are the whole answer.

Next we walk through each box on the diagram. I start with the state machine because CLOSED → OPEN → HALF-OPEN transitions are the entire interview answer.

State Machine Deep Dive

Walk through each transition trigger (failure rate threshold, wait duration, and probe success criteria) before your interviewer asks about false positives.

CLOSED → OPEN: Triggered when failure_rate > threshold (e.g., 50%) over sliding window AND minimum_calls met (e.g., 20). Sliding window types: count-based (last N calls) or time-based (last T seconds). Failure types counted: exceptions/HTTP 5xx, timeouts, and slow calls (> slow_call_threshold).

OPEN → HALF-OPEN: After wait_duration (e.g., 30 seconds) to give downstream time to recover.

HALF-OPEN → CLOSED: If permitted probe calls succeed (e.g., 5/5). Reset failure metrics, resume normal traffic.

HALF-OPEN → OPEN: If any probe fails. Restart wait_duration timer. Optimization: exponential backoff (30s → 60s → 120s → max 300s).

In-Process Sliding Window (Count-Based)

The ring buffer implementation is worth mentioning: O(1) per call with fixed memory is why in-process breakers add negligible overhead.

Implemented as a ring buffer: O(1) per call, fixed memory. Each slot stores outcome: 0=success, 1=failure, 2=slow. On each call result, update ring buffer head, increment/decrement counters, check threshold, and possibly trip to OPEN.

Distributed Circuit Breaker State

Distributed state sharing is a common follow-up: explain why you keep the breaker in-process and use Redis only for visibility, not decisions.

Redis can be used for cross-instance CB state sharing: HSET cb:payment:/charge state "OPEN". Each instance publishes local metrics every 5 seconds. Trade-off: adds Redis latency to hot path. Recommendation: keep CB in-process, share state for visibility only.

API Design

Configuration API

PUT /api/circuit-breakers/{service}/{endpoint}
{
  "failure_rate_threshold": 50,
  "slow_call_rate_threshold": 80,
  "slow_call_duration_ms": 5000,
  "sliding_window_type": "COUNT",
  "sliding_window_size": 100,
  "minimum_calls": 20,
  "wait_duration_open_ms": 30000,
  "permitted_calls_half_open": 5,
  "fallback_type": "CACHE",
  "manual_override": null
}

GET /api/circuit-breakers/status → All CB states
POST /api/circuit-breakers/{service}/{endpoint}/override
  { "state": "FORCE_OPEN" }

Metrics API

JSON
GET /api/circuit-breakers/metrics?service=payment-service
{
  "circuit_state": "OPEN",
  "failure_rate": 67.5,
  "total_calls": 1500,
  "successful_calls": 487,
  "failed_calls": 1013,
  "not_permitted_calls": 2300,
  "fallback_calls": 2300,
  "state_transitions": [
    {"from": "CLOSED", "to": "OPEN", "at": "2026-03-14T10:05:00Z"}
  ]
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Ring Buffer Implementation (Count-Based)

JAVA
class CircuitBreaker {
    enum State { CLOSED, OPEN, HALF_OPEN }

    State state = CLOSED;
    int[] outcomes;          // ring buffer: 0=success, 1=failure, 2=slow
    int head = 0;
    int totalFailures = 0;
    int totalCalls = 0;
    long openedAt;
    Config config;

    Result execute(Supplier<Result> call, Supplier<Result> fallback) {
        if (state == OPEN) {
            if (System.nanoTime() - openedAt > config.waitDuration) {
                state = HALF_OPEN;
                halfOpenPermits = config.permittedCallsHalfOpen;
            } else {
                metrics.increment("not_permitted");
                return fallback.get();
            }
        }

        if (state == HALF_OPEN && halfOpenPermits <= 0) {
            return fallback.get();
        }

        try {
            long start = System.nanoTime();
            Result r = call.get();
            long duration = System.nanoTime() - start;
            recordSuccess(duration);
            return r;
        } catch (Exception e) {
            recordFailure();
            return fallback.get();
        }
    }
}

Prometheus Metrics

# Circuit breaker state (0=closed, 1=open, 2=half-open)
circuit_breaker_state{service="payment", endpoint="/charge"} 1

# Call outcomes
circuit_breaker_calls_total{service="payment", outcome="success"} 487
circuit_breaker_calls_total{service="payment", outcome="failure"} 1013
circuit_breaker_calls_total{service="payment", outcome="not_permitted"} 2300

# State transitions
circuit_breaker_transitions_total{service="payment", from="closed", to="open"} 3

Fault Tolerance

Cascading Failure Prevention

Without CB: Service C is down → B gets timeouts → B's thread pool exhausted → A's thread pool exhausted → TOTAL SYSTEM DOWN. With CB: C is down → B's CB trips after 50% failures → B returns fallback immediately → B's thread pool stays healthy → A stays healthy. Key insight: CB prevents resource exhaustion by failing fast.

Bulkhead Pattern (Complementary to CB)

Even with CB, a slow service can consume all threads. Bulkhead isolates thread pools per downstream service (e.g., 20 threads for B, 10 for C, 15 for D). If C is slow, only 10 threads blocked. Implementations: Hystrix-style thread pool isolation, semaphore isolation (lighter weight), Envoy max_connections per upstream.

Timeout Strategy

Layered timeouts (defense in depth): connection timeout (1s), read timeout (5s), CB slow-call threshold (3s), retry budget (max 3, total 10s), deadline propagation. Without layered timeouts: client timeout of 30s waits 30s for each failing call. With CB + timeouts: after 20 failures (each 5s timeout) → CB trips → instant fallback.

Distributed CB Coordination

Problem: 10 instances may each see different failure rates. Options: 1) No coordination (simplest, eventually converge). 2) Shared metrics via Redis (consistent but adds dependency). 3) Control plane gossip (eventual consistency, no critical-path dependency). Recommendation: Option 1 for most cases, Option 3 for critical paths.

Retry vs Circuit Breaker Interaction

Bad: retry(3, circuit_breaker(call)): retries waste attempts even when CB is open. Good: circuit_breaker(retry(3, call)): CB wraps retries. When open, no retries attempted. When persistent failures → CB trips, stops all attempts.

Additional Considerations

Real-World Configurations

ServiceFailure ThresholdWait DurationFallback
Low-risk (catalog)70%10scached catalog data
High-risk (payment)30%60squeue for later
Internal microservice50%30sreturn error

Testing Circuit Breakers

Chaos engineering: kill downstream service → verify CB trips, inject latency → verify slow-call detection. Load testing: verify CB doesn't trip under normal load, trips quickly under failure. Integration test: configure low thresholds, send mix of success/failure, assert state transitions.

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Fail-fast against cascading failures (5 min)
    • Three states: CLOSED → OPEN → HALF-OPEN (6 min)
    • Sliding window: count-based or time-based (5 min)
    • Count failures broadly: 5xx, timeouts, slow calls (5 min)
    • Fallback strategies and bulkhead isolation (4 min)
  • Frame as fail-fast protection against cascading failures: a slow downstream must not exhaust the caller's thread pool and take down upstream services.
  • Walk through the three states: CLOSED (normal) → OPEN (reject + fallback after threshold breach) → HALF-OPEN (probe recovery with limited calls).
  • Explain the sliding window: count-based ring buffer or time-based window with minimum_calls before tripping to avoid false positives on cold start.
  • Count failures broadly: exceptions, HTTP 5xx, timeouts, and slow calls exceeding slow_call_duration_ms.
  • Pair with Bulkheads (isolated thread pools per downstream) and layered timeouts; CB alone does not prevent resource exhaustion from slow calls below the trip threshold.
  • Critical ordering: wrap retries inside the circuit breaker (circuit_breaker(retry(call))), not the reverse: retries must not fire when the breaker is open.
  • Pair Resilience4j (~<1μs in-process check) with Istio outlier detection (~1ms sidecar hop): library for semantic failures, mesh for connection-level ejection.

Engineering Trade-offs

Circuit Breaker vs Bulkhead vs Rate Limiter vs Timeout

Circuit breakers trade fast failure against recovery probing, balancing threshold tuning and the half-open probe rate.

PatternProblemSolution
TimeoutDownstream hangs → threads pile upSet max wait time (e.g., 500ms). ALWAYS use this.
RetryTransient failuresRetry with exponential backoff + jitter. Don't retry 4xx.
Circuit BreakerConsistent failures → wasted resourcesTrip breaker → fast fail for 30-60s → probe recovery
BulkheadOne slow dependency exhausts thread poolSeparate thread pools per dependency
Rate LimiterCaller overloads the calleeLimit calls per second to downstream as seen in distributed rate limiting

Half-Open State: The Critical Recovery Mechanism

Without Half-Open: once OPEN, circuit never recovers. With Half-Open: wait duration elapses → allow probe → if succeeds, CLOSED; if fails, OPEN again. Two approaches:

Conservative (1 probe at a time): minimal risk but slow recovery if traffic is low.

Aggressive (sliding window, e.g., 10% of requests): faster recovery, more risk. Resilience4j default: 10 permitted calls in half-open. Best practice: set wait_duration slightly longer than P95 downstream latency.

Service Mesh vs Application Code

Application code (Resilience4j, Hystrix): fine-grained per method, custom fallbacks in app logic, zero latency check. Cons: per-language implementation, code coupling.

Service Mesh (Istio/Envoy): language-agnostic, centralized observability, no code changes. Cons: coarser granularity, network-level only, ~1ms overhead.

Best practice: service mesh for basic CB (~1ms sidecar overhead per hop) + library for complex fallback logic (~<1μs in-process). Istio handles connection-level failures; Resilience4j handles business logic exceptions.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...