Interview Setup
Interview Prompt
Design an API Gateway (Kong/Envoy) handling 5M RPS globally across 2,000 data plane nodes, with dynamic routing, JWT auth, distributed rate limiting, mTLS to upstreams, and WAF at the edge.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Are we designing the data plane or the control plane? | The data plane must be ultra low latency, whereas the control plane is a standard configuration management system. |
| What scale of traffic must it handle? | Determines if we need a highly optimized proxy core (like Nginx or Envoy) or if a Node.js or Go proxy is sufficient. |
| What cross cutting features are required? | Rate limiting, authentication, logging, circuit breaking, and WAF integration have different data and failure dependencies. |
Scope
In scope
- Data plane (Request routing and proxying)
- Control plane (Configuration management)
- Plugin execution model (Auth, Rate Limiting)
- High availability and scaling
Out of scope (state explicitly)
- Detailed implementation of the backend microservices
- Billing systems for monetization
Functional Requirements
Start by asking your interviewer whether you are designing the data plane (request proxy) or the control plane (route configuration), because they have different SLOs. Confirm plugin requirements such as JWT authentication, rate limiting, circuit breaking, and hot reload without dropping active connections.
- Routing: Route incoming HTTP requests to the correct backend microservice based on URL paths and headers.
- Authentication & Authorization: The gateway validates authentication tokens such as JWTs, while backend services enforce resource authorization.
- Rate Limiting: Prevent abuse by limiting the number of requests per user or IP over a time window, applying principles from an API Rate Limiter.
- Load Balancing & Retry: Distribute traffic across healthy backend instances and retry failed idempotent requests.
- Dynamic Configuration: Routes and limits must be updatable without restarting the gateway (zero downtime).
Non-Functional Requirements
Your interviewer will stress-test added latency on every request and the chosen fail open vs fail closed behavior when Redis rate limiting is unavailable. They will also probe control plane vs data plane separation and zero downtime configuration reload through xDS or etcd watches.
- Ultra Low Latency: The gateway sits in the critical path of every request, so total added overhead should remain in the single digit milliseconds, with a p99 target below 5 ms.
- High Throughput: Must support 50M concurrent connections globally, distributed across the gateway fleet.
- Fault Tolerance: The gateway must not crash if backend services fail, relying on the Circuit Breaker pattern.
- Scalability: Scale horizontally by adding gateway nodes behind an L4 Load Balancer.
Capacity Estimations
At 5M RPS globally, each of 2,000 nodes handles about 2,500 RPS. The ~500K Redis operations per second figure assumes 10% synchronous global checks. Strict per request Redis enforcement would approach 5M operations per second, so local bounded state is a deliberate accuracy and latency tradeoff rather than the default for security sensitive limits.
| Metric | Calculation | Value |
|---|---|---|
| Peak requests/sec (global) | Given | 5M RPS |
| Gateway nodes (data plane) | Given | 2,000 |
| RPS per node | 5M ÷ 2000 | ~2,500 RPS/node |
| Avg request size | Given | 2 KB |
| Avg response size | Assumed | 2 KB |
| Bidirectional bandwidth (peak) | 5M x 2KB x 2 (request + response) | ~20 GB/s |
| Concurrent connections | Given | 50M |
| Concurrent connections per node | 50M ÷ 2000 | ~25K/node average |
| Route configuration count | Given | 100K |
| Strict global Redis checks/sec | One atomic decision per request | Up to 5M ops/s across sharded Redis |
| Sampled global Redis checks/sec | 5M checks x 10% reconciliation | ~500K ops/s with bounded overshoot |
At 2,500 RPS/node on an event loop architecture using epoll, each node carries about 25K concurrent keep alive connections under the stated 50M global connection target. The added gateway latency budget is modeled as under 2 ms of proxy overhead, about 1 ms for JWT verification, and about 1 ms for the rate limit decision. This gives a roughly 4 ms internal processing budget before network and upstream time, leaving room under the p99 gateway overhead target. The control plane pushes validated configuration through xDS or etcd watches, allowing 100K routes to reload in under 100 ms without dropping active connections.
Architecture Diagram
At 5M RPS globally, the API gateway introduces overhead on every request, which means added latency must stay in single digit milliseconds. The control plane owns route CRUD and plugin configuration, whereas the data plane manages connection handling, JWT verification, rate limiting, and upstream proxying with hot reload so configuration changes never drop active keep alive sessions.
The WAF and anycast edge absorb raw DDoS attacks before traffic reaches the origin, after which the gateway terminates client TLS, enforces authentication and quotas, and forwards requests to microservices over mTLS with SPIFFE workload identities. Redis holds distributed rate limit state, while optional local counter caches trade strict precision for lower latency when the endpoint policy permits bounded overshoot.
In the room
Separate the control plane (route CRUD, admin API) from the data plane (high throughput proxy) in your first sketch. Clarify that rate limiting can fail open when policy permits it, while JWT authentication fails closed.
Component Deep Dives
1. The Data Plane (Asynchronous I/O)
Request flow crosses four concerns in sequence: terminating TLS, authenticating, rate limiting, and routing to a healthy upstream, with each step implemented as an inline plugin on the event loop proxy core. The sections below connect asynchronous I/O architecture, zero downtime configuration propagation, distributed rate limiting, stateless JWT authentication, circuit breaking, and the edge security layers that sit in front of the gateway.
An event loop architecture with epoll or kqueue handles thousands of concurrent connections per worker thread rather than a thread per connection model, which quickly exhausts CPU through context switching.
To achieve high throughput (10,000+ requests/sec per node), gateways like Nginx (which Kong is based on) or Envoy do not use a thread per connection model (like older Apache or Tomcat servers), which would exhaust CPU and RAM via context switching overhead. Instead, they use an Event Loop architecture (epoll/kqueue) with non-blocking I/O. A single worker thread handles thousands of concurrent connections efficiently, dramatically reducing memory overhead.
2. Dynamic Configuration (Control Plane) ⭐
Dynamic configuration via xDS streams or etcd watches lets gateway nodes receive validated routing snapshots without restarting. The proxy atomically swaps the active configuration in memory, while worker generations can drain existing connections before retiring old state.
If an engineering team deploys a new microservice, the gateway needs to know the new route. If we have 50 gateway nodes, restarting them to update a YAML config file drops active connections and causes an outage. Modern gateways solve this by separating the control plane from the data plane. The control plane stores durable configuration, validates it, and publishes versioned snapshots through xDS or an equivalent distribution channel. Data plane nodes consume those snapshots locally instead of querying PostgreSQL on the request path.
# Versioned configuration snapshot published by the Control Plane
version: 42
route:
name: payment-service-route
path: /api/v1/payments/*
methods:
- GET
- POST
backend:
service: payment-service
port: 8443
plugins:
- rate-limiting
- jwt
# The Control Plane distributes this validated snapshot through xDS.
# Data Plane nodes atomically swap the active snapshot in memory.
# Existing connections continue draining on the previous snapshot.3. Distributed Rate Limiting
Global rate limiting requires shared state across gateway nodes when the quota must be exact. Redis with Lua provides atomic Token Bucket decisions for that strict mode. Optional local token buckets can reduce latency when the policy accepts bounded global overshoot and asynchronous reconciliation.
Rate limiting mitigates application abuse and request floods. Volumetric DDoS is primarily handled at the CDN or anycast WAF edge. For globally enforced limits, the gateway shares Redis state across nodes so an attacker cannot bypass a limit simply by reaching another node.
Optimization (Local Caching vs Accuracy): A synchronous Redis round trip for every request can add 1 to 2 ms of latency. A local token bucket with a preallocated per node budget can reduce that cost, with asynchronous reconciliation to Redis. This trades exact global enforcement for bounded overshoot, so it is appropriate only for policies that explicitly accept approximate global enforcement.
4. Authentication (Stateless JWT)
Stateless JWT validation at the gateway avoids a database round trip on every request by verifying signatures offline using cached JWKS keys and validating the issuer, audience, expiry, not before, and required scopes or roles. Refresh JWKS on an unknown key identifier and allow key overlap during key rotation.
Making a database call to the User Service to validate a session token for every API request creates unnecessary dependency and latency. Instead, the gateway uses stateless JSON Web Tokens (JWT).
Because the JWT is cryptographically signed with asymmetric keys such as RS256, the gateway can verify the token's authenticity offline using cached JWKS keys. It trusts only validated claims such as user_id and roles. The gateway first strips any client supplied identity headers, then injects its own authenticated identity headers over the mTLS protected upstream connection. Backend services still enforce authorization based on the trusted identity and resource policy.
5. Circuit Breaker Pattern
Circuit breakers prevent a failing upstream from exhausting the gateway connection pool and causing cascading failure across the entire API platform.
If the "Order Service" goes down, the gateway might queue up thousands of requests waiting for it to respond, eventually exhausting the gateway's own connection pool and causing a catastrophic cascading failure across the entire API platform.
- The Circuit Breaker tracks failure rates, such as 50% failures occurring in 10 seconds.
- If the threshold is crossed, the circuit Opens. The gateway instantly rejects new requests to the Order Service with
503 Service Unavailablewithout even trying the network call, giving the backend time to recover. - After a timeout, it shifts to Half-Open, allowing a small percentage of test requests through. If they succeed, it Closes the circuit. Failed test requests reopen it.
6. mTLS Upstream & WAF at the Edge
Defense in depth spans three layers. The edge WAF blocks OWASP attacks, bot floods, and volumetric traffic. Gateway plugins enforce authentication, request limits, and circuit breaking. mTLS and workload identity protect the gateway to service path. Rate limiting can fail open where endpoint policy permits it, while authentication failures fail closed.
mTLS (service to service):
- The edge security layer filters public traffic before the gateway.
- The gateway routes to upstream microservices over mTLS using SPIFFE/SPIRE or an equivalent workload PKI.
- The gateway validates the workload identity in the certificate against the expected service identity.
- Backend services accept gateway traffic only over the authenticated channel.
- This limits lateral movement if a pod is compromised.
WAF / edge protection:
Layer 1: Anycast CDN/WAF
OWASP Top 10 rules, SQLi/XSS blocking, geo-blocking, bot scoring, and volumetric DDoS absorption.
If the WAF terminates TLS, it re-encrypts traffic to the gateway. Otherwise the gateway terminates client TLS.
Layer 2: Gateway plugins
IP rate limiting, JWT validation, request size caps, and circuit breaking.
Layer 3: Backend
Resource authorization remains in the backend service. The gateway authenticates the caller but does not replace resource authorization.
Rate limiting policy:
- Strict global limits use an atomic Redis decision.
- Latency sensitive endpoints may use bounded local token buckets with asynchronous reconciliation when the policy permits controlled overshoot.
- Availability oriented endpoints may fail open on Redis outage. Security sensitive endpoints may fail closed.API Design
Typed Gateway Domain Contract
The control plane exposes typed route and plugin configuration objects, while the data plane remains a high throughput proxy interface rather than a business API.
export type RouteId = string;
export type ServiceId = string;
export type HttpMethod = "GET" | "POST" | "PUT" | "PATCH" | "DELETE" | "HEAD" | "OPTIONS";
export type PluginType = "jwt" | "rate_limiting" | "circuit_breaker";
export interface RateLimitPolicy {
limit: number;
windowSeconds: number;
scope: "ip" | "user" | "organization" | "route";
}
export interface RouteConfig {
routeId: RouteId;
pathPrefix: string;
methods: HttpMethod[];
serviceId: ServiceId;
timeoutMs: number;
plugins: PluginType[];
rateLimit?: RateLimitPolicy;
version: number;
}
export interface RouteCreateRequest {
idempotencyKey: string;
route: Omit<RouteConfig, "routeId" | "version">;
}
export interface RouteUpdateRequest {
expectedVersion: number;
route: Omit<RouteConfig, "routeId" | "version">;
}Control Plane Admin API
The Gateway exposes a private administrative interface for route and plugin configuration. It is protected by workload identity or mTLS plus RBAC and is never placed on the public request path.
// Control Plane Admin API: Add a new route
POST /admin/api/routes
Idempotency-Key: route-create-8f3c1b
Content-Type: application/json
{
"name": "payment-service-route",
"paths": ["/api/v1/payments"],
"methods": ["GET", "POST"],
"service": {
"host": "payment.internal",
"port": 8443
},
"plugins": [
{ "name": "rate-limiting", "config": { "limit": 100, "window": "1m" } },
{ "name": "jwt" }
]
}Standard Data Plane Errors
The data plane uses stable HTTP error semantics so clients can distinguish authentication failures, quota rejection, missing routes, upstream failures, and gateway timeouts.
HTTP/1.1 401 Unauthorized
{
"error": "invalid_or_expired_token"
}
HTTP/1.1 403 Forbidden
{
"error": "insufficient_permissions"
}
HTTP/1.1 404 Not Found
{
"error": "route_not_found"
}
HTTP/1.1 409 Conflict
{
"error": "config_version_conflict"
}
HTTP/1.1 429 Too Many Requests
Retry-After: 60
{
"error": "rate_limit_exceeded"
}
HTTP/1.1 502 Bad Gateway
{
"error": "upstream_connect_failed"
}
HTTP/1.1 503 Service Unavailable
{
"error": "upstream_unavailable"
}
HTTP/1.1 504 Gateway Timeout
{
"error": "upstream_timeout"
}Data Model
The Gateway requires two different storage roles. PostgreSQL holds durable control plane configuration, while Redis holds high throughput, volatile rate limiting state and fast distributed coordination data.
CREATE TABLE routes (
id UUID PRIMARY KEY,
name VARCHAR(128) NOT NULL,
path_prefix VARCHAR(512) NOT NULL,
upstream_service_id UUID NOT NULL,
version BIGINT NOT NULL,
updated_at TIMESTAMPTZ NOT NULL
);
CREATE TABLE plugins (
id UUID PRIMARY KEY,
route_id UUID NOT NULL REFERENCES routes(id) ON DELETE CASCADE,
plugin_type VARCHAR(64) NOT NULL,
config JSONB NOT NULL
);# Configuration distribution
# etcd may hold the published snapshot, while xDS is the delivery protocol.
gateway/config/{version} = validated route + plugin snapshot
# Data plane rate limiting state
# Lua atomically consumes tokens, refills according to elapsed time, and refreshes TTL.
rate_limit:{scope}:{client_id}:{route_id} = {
tokens,
last_refill_ms
}
TTL = expire after idle periodFault Tolerance
| Failure Case | System Solution Design |
|---|---|
| Redis Rate Limiter Crash | Fail open only where endpoint policy permits availability to win over strict quota enforcement. Use local bounded fallback state and emit telemetry. Security sensitive endpoints can fail closed. |
| Control Plane Crash | Data Plane nodes retain the last validated routing snapshot in local memory. Existing traffic continues while configuration changes pause until a healthy control plane resumes publication. |
| Stale Configuration Snapshot | Gateways apply monotonically increasing versions, reject out of order snapshots, and canary new versions before global rollout. A previous validated snapshot remains available for rollback. |
| Backend Service Timeout | Gateway enforces strict timeouts and bounded retries. Retry only safe or explicitly idempotent requests before a response is committed, and only against healthy endpoints. |
| Plugin Runtime Failure | Bound plugin CPU, memory, and execution time. A faulty plugin instance is isolated from the proxy event loop, disabled or rolled back, and reported through control plane health signals. |
| WAF or Edge Dependency Failure | Keep multiple edge providers or regional edge capacity where availability requirements justify it. Preserve gateway capacity for traffic that remains admitted after edge protection. |
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless interviewing for a staff or principal role.
- Separate control plane (config CRUD, admin API) from high throughput data plane proxy (5 min)
- Data plane uses event loop with non-blocking I/O rather than thread per connection (6 min)
- Dynamic config via xDS stream or etcd watch (5 min)
- Rate limiting: Redis Lua globally, with optional local bounded cache (5 min)
- JWT validation offline at gateway, injecting X-User-Id downstream (4 min)
- Separate control plane (config CRUD, admin API) from data plane (high throughput proxy), since they operate with different SLOs.
- Data plane uses an event loop with non-blocking I/O rather than thread per connection.
- Dynamic config via xDS stream or etcd watch enables hot reload without dropping connections.
- Rate limiting: Redis Lua globally, with an optional local bounded token bucket where bounded overshoot is acceptable.
- JWT validation offline at gateway, injecting X-User-Id headers for backends.
- WAF at the CDN edge absorbs volumetric and application attacks, while the gateway adds mTLS to upstream microservices.
- Track per route p50, p95, and p99 latency, upstream error rate, circuit breaker state, rate limit decisions, active connections, configuration version skew, and plugin execution time.
- Common pitfall: a synchronous Redis call on every request adds latency, so use exact Redis enforcement where required and local bounded state only where policy permits approximate global enforcement.
Engineering Trade-offs
API Gateway vs Service Mesh (Istio)
An API Gateway handles "north south" traffic (external clients entering the datacenter). A Service Mesh handles "east west" traffic (internal microservices talking to each other). While they share technologies (Envoy is often used for both), they serve different purposes. Gateways focus on Edge security (WAF, OAuth), while Service Meshes focus on internal mTLS, distributed tracing, and complex routing between hundreds of internal pods.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.