Interview Setup
Interview Prompt
Design a circuit breaker library/middleware protecting 500 microservices making 10M inter-service RPC calls/sec, with ~5,000 per-service circuit breakers and configurable failure thresholds.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| In-process library or sidecar/service-mesh proxy? |
|
| Failure rate or slow-call rate as trigger, or both? | A service returning 200 in 30s is 'healthy' by error rate but kills caller latency budget. |
| Per-endpoint or per-service circuit breakers? |
|
| Fallback strategies in scope? |
|
Scope
In scope
- State machine (closed → open → half-open)
- Failure rate thresholds
- Bulkhead pattern
- Fallback strategies
- Integration with service mesh
- Capacity estimation with shown math
Out of scope (state explicitly)
- Full service mesh control plane
- Application business logic in downstream services
- Building a custom service registry from scratch when managed options exist
Functional Requirements
Start by confirming whether you are designing an in-process library or a service-mesh sidecar. Ask about fallback strategies, per-endpoint breakers, and manual override requirements.
- Monitor health of downstream service calls (success/failure rates)
- Automatically stop sending requests to unhealthy services (circuit OPEN)
- Periodically probe to check if service has recovered (HALF-OPEN state)
- Resume traffic when service recovers (circuit CLOSED)
- Configurable thresholds: failure rate, slow-call rate, minimum call volume
- Support per-service, per-endpoint, per-client circuit breakers
- Dashboard showing circuit breaker states across all services
- Integration with service mesh (Envoy/Istio sidecar) or application library
- Fallback responses when circuit is open (cached response, default value, degraded mode)
- Manual override: force-open or force-close circuits
Non-Functional Requirements
Your interviewer will care most about sub-microsecond overhead and fast failure detection, as explored in Circuit Breakers, Retries & Bulkheads. The breaker must add almost zero latency in CLOSED state: it is a local counter check, not a network hop.
- Ultra-Low Overhead: < 1µs per call decision (in-process check)
- Fast Detection: Detect failures within 5 to 10 seconds
- Fast Recovery: Resume traffic within seconds of downstream recovery
- No SPOF: Circuit breaker itself must not become a single point of failure
- Consistency: All instances of a service should have similar circuit state (eventual)
- Observability: Emit metrics for open/close transitions, fallback invocations
Capacity Estimations
Run this math before you size the sliding window. Calls per second per service and failure threshold tell you ring buffer size; probe count in HALF-OPEN drives recovery latency.
| Metric | Calculation | Value |
|---|---|---|
| Services in production | Given | 500 |
| Inter-service RPC calls / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 10M |
| Circuit breakers (per service x endpoint) | Given | ~5,000 |
| State transitions / min (normal) | Given | < 10 |
| Metric data points / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 50K |
Architecture Diagram
Walk your interviewer through the three-state machine first. An in-process circuit breaker library wraps calls to downstream Service B. The state machine transitions between CLOSED (normal operation), OPEN (rejecting requests), and HALF-OPEN (probing recovery). Metrics are emitted to Prometheus for dashboarding; I label each transition and its trigger condition on the whiteboard.
Component Deep Dives
In the room: draw the CLOSED → OPEN → HALF-OPEN state machine before any boxes: transitions and thresholds are the whole answer.
Next we walk through each box on the diagram. I start with the state machine because CLOSED → OPEN → HALF-OPEN transitions are the entire interview answer.
State Machine Deep Dive
Walk through each transition trigger (failure rate threshold, wait duration, and probe success criteria) before your interviewer asks about false positives.
CLOSED → OPEN: Triggered when failure_rate > threshold (e.g., 50%) over sliding window AND minimum_calls met (e.g., 20). Sliding window types: count-based (last N calls) or time-based (last T seconds). Failure types counted: exceptions/HTTP 5xx, timeouts, and slow calls (> slow_call_threshold).
OPEN → HALF-OPEN: After wait_duration (e.g., 30 seconds) to give downstream time to recover.
HALF-OPEN → CLOSED: If permitted probe calls succeed (e.g., 5/5). Reset failure metrics, resume normal traffic.
HALF-OPEN → OPEN: If any probe fails. Restart wait_duration timer. Optimization: exponential backoff (30s → 60s → 120s → max 300s).
In-Process Sliding Window (Count-Based)
The ring buffer implementation is worth mentioning: O(1) per call with fixed memory is why in-process breakers add negligible overhead.
Implemented as a ring buffer: O(1) per call, fixed memory. Each slot stores outcome: 0=success, 1=failure, 2=slow. On each call result, update ring buffer head, increment/decrement counters, check threshold, and possibly trip to OPEN.
Distributed Circuit Breaker State
Distributed state sharing is a common follow-up: explain why you keep the breaker in-process and use Redis only for visibility, not decisions.
Redis can be used for cross-instance CB state sharing: HSET cb:payment:/charge state "OPEN". Each instance publishes local metrics every 5 seconds. Trade-off: adds Redis latency to hot path. Recommendation: keep CB in-process, share state for visibility only.
API Design
Configuration API
PUT /api/circuit-breakers/{service}/{endpoint}
{
"failure_rate_threshold": 50,
"slow_call_rate_threshold": 80,
"slow_call_duration_ms": 5000,
"sliding_window_type": "COUNT",
"sliding_window_size": 100,
"minimum_calls": 20,
"wait_duration_open_ms": 30000,
"permitted_calls_half_open": 5,
"fallback_type": "CACHE",
"manual_override": null
}
GET /api/circuit-breakers/status → All CB states
POST /api/circuit-breakers/{service}/{endpoint}/override
{ "state": "FORCE_OPEN" }Metrics API
GET /api/circuit-breakers/metrics?service=payment-service
{
"circuit_state": "OPEN",
"failure_rate": 67.5,
"total_calls": 1500,
"successful_calls": 487,
"failed_calls": 1013,
"not_permitted_calls": 2300,
"fallback_calls": 2300,
"state_transitions": [
{"from": "CLOSED", "to": "OPEN", "at": "2026-03-14T10:05:00Z"}
]
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
Ring Buffer Implementation (Count-Based)
class CircuitBreaker {
enum State { CLOSED, OPEN, HALF_OPEN }
State state = CLOSED;
int[] outcomes; // ring buffer: 0=success, 1=failure, 2=slow
int head = 0;
int totalFailures = 0;
int totalCalls = 0;
long openedAt;
Config config;
Result execute(Supplier<Result> call, Supplier<Result> fallback) {
if (state == OPEN) {
if (System.nanoTime() - openedAt > config.waitDuration) {
state = HALF_OPEN;
halfOpenPermits = config.permittedCallsHalfOpen;
} else {
metrics.increment("not_permitted");
return fallback.get();
}
}
if (state == HALF_OPEN && halfOpenPermits <= 0) {
return fallback.get();
}
try {
long start = System.nanoTime();
Result r = call.get();
long duration = System.nanoTime() - start;
recordSuccess(duration);
return r;
} catch (Exception e) {
recordFailure();
return fallback.get();
}
}
}Prometheus Metrics
# Circuit breaker state (0=closed, 1=open, 2=half-open)
circuit_breaker_state{service="payment", endpoint="/charge"} 1
# Call outcomes
circuit_breaker_calls_total{service="payment", outcome="success"} 487
circuit_breaker_calls_total{service="payment", outcome="failure"} 1013
circuit_breaker_calls_total{service="payment", outcome="not_permitted"} 2300
# State transitions
circuit_breaker_transitions_total{service="payment", from="closed", to="open"} 3Fault Tolerance
Cascading Failure Prevention
Without CB: Service C is down → B gets timeouts → B's thread pool exhausted → A's thread pool exhausted → TOTAL SYSTEM DOWN. With CB: C is down → B's CB trips after 50% failures → B returns fallback immediately → B's thread pool stays healthy → A stays healthy. Key insight: CB prevents resource exhaustion by failing fast.
Bulkhead Pattern (Complementary to CB)
Even with CB, a slow service can consume all threads. Bulkhead isolates thread pools per downstream service (e.g., 20 threads for B, 10 for C, 15 for D). If C is slow, only 10 threads blocked. Implementations: Hystrix-style thread pool isolation, semaphore isolation (lighter weight), Envoy max_connections per upstream.
Timeout Strategy
Layered timeouts (defense in depth): connection timeout (1s), read timeout (5s), CB slow-call threshold (3s), retry budget (max 3, total 10s), deadline propagation. Without layered timeouts: client timeout of 30s waits 30s for each failing call. With CB + timeouts: after 20 failures (each 5s timeout) → CB trips → instant fallback.
Distributed CB Coordination
Problem: 10 instances may each see different failure rates. Options: 1) No coordination (simplest, eventually converge). 2) Shared metrics via Redis (consistent but adds dependency). 3) Control plane gossip (eventual consistency, no critical-path dependency). Recommendation: Option 1 for most cases, Option 3 for critical paths.
Retry vs Circuit Breaker Interaction
Bad: retry(3, circuit_breaker(call)): retries waste attempts even when CB is open. Good: circuit_breaker(retry(3, call)): CB wraps retries. When open, no retries attempted. When persistent failures → CB trips, stops all attempts.
Additional Considerations
Real-World Configurations
| Service | Failure Threshold | Wait Duration | Fallback |
|---|---|---|---|
| Low-risk (catalog) | 70% | 10s | cached catalog data |
| High-risk (payment) | 30% | 60s | queue for later |
| Internal microservice | 50% | 30s | return error |
Testing Circuit Breakers
Chaos engineering: kill downstream service → verify CB trips, inject latency → verify slow-call detection. Load testing: verify CB doesn't trip under normal load, trips quickly under failure. Integration test: configure low thresholds, send mix of success/failure, assert state transitions.
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Fail-fast against cascading failures (5 min)
- Three states: CLOSED → OPEN → HALF-OPEN (6 min)
- Sliding window: count-based or time-based (5 min)
- Count failures broadly: 5xx, timeouts, slow calls (5 min)
- Fallback strategies and bulkhead isolation (4 min)
- Frame as fail-fast protection against cascading failures: a slow downstream must not exhaust the caller's thread pool and take down upstream services.
- Walk through the three states: CLOSED (normal) → OPEN (reject + fallback after threshold breach) → HALF-OPEN (probe recovery with limited calls).
- Explain the sliding window: count-based ring buffer or time-based window with
minimum_callsbefore tripping to avoid false positives on cold start. - Count failures broadly: exceptions, HTTP 5xx, timeouts, and slow calls exceeding
slow_call_duration_ms. - Pair with Bulkheads (isolated thread pools per downstream) and layered timeouts; CB alone does not prevent resource exhaustion from slow calls below the trip threshold.
- Critical ordering: wrap retries inside the circuit breaker (
circuit_breaker(retry(call))), not the reverse: retries must not fire when the breaker is open. - Pair Resilience4j (~<1μs in-process check) with Istio outlier detection (~1ms sidecar hop): library for semantic failures, mesh for connection-level ejection.
Engineering Trade-offs
Circuit Breaker vs Bulkhead vs Rate Limiter vs Timeout
Circuit breakers trade fast failure against recovery probing, balancing threshold tuning and the half-open probe rate.
| Pattern | Problem | Solution |
|---|---|---|
| Timeout | Downstream hangs → threads pile up | Set max wait time (e.g., 500ms). ALWAYS use this. |
| Retry | Transient failures | Retry with exponential backoff + jitter. Don't retry 4xx. |
| Circuit Breaker | Consistent failures → wasted resources | Trip breaker → fast fail for 30-60s → probe recovery |
| Bulkhead | One slow dependency exhausts thread pool | Separate thread pools per dependency |
| Rate Limiter | Caller overloads the callee | Limit calls per second to downstream as seen in distributed rate limiting |
Half-Open State: The Critical Recovery Mechanism
Without Half-Open: once OPEN, circuit never recovers. With Half-Open: wait duration elapses → allow probe → if succeeds, CLOSED; if fails, OPEN again. Two approaches:
Conservative (1 probe at a time): minimal risk but slow recovery if traffic is low.
Aggressive (sliding window, e.g., 10% of requests): faster recovery, more risk. Resilience4j default: 10 permitted calls in half-open. Best practice: set wait_duration slightly longer than P95 downstream latency.
Service Mesh vs Application Code
Application code (Resilience4j, Hystrix): fine-grained per method, custom fallbacks in app logic, zero latency check. Cons: per-language implementation, code coupling.
Service Mesh (Istio/Envoy): language-agnostic, centralized observability, no code changes. Cons: coarser granularity, network-level only, ~1ms overhead.
Best practice: service mesh for basic CB (~1ms sidecar overhead per hop) + library for complex fallback logic (~<1μs in-process). Istio handles connection-level failures; Resilience4j handles business logic exceptions.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.