System Design Problem

Design an API Gateway (Kong / Envoy)

Commonly Asked By:KongStripeNetflixLyft

Interview Setup

Interview Prompt

Design an API Gateway (Kong/Envoy) handling 5M RPS globally across 2,000 data plane nodes, with dynamic routing, JWT auth, distributed rate limiting, mTLS to upstreams, and WAF at the edge.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Are we designing the data plane or the control plane?The data plane must be ultra low latency, whereas the control plane is a standard configuration management system.
What scale of traffic must it handle?Determines if we need a highly optimized proxy core (like Nginx or Envoy) or if a Node.js or Go proxy is sufficient.
What cross cutting features are required?Rate limiting, authentication, logging, circuit breaking, and WAF integration have different data and failure dependencies.

Scope

In scope

  • Data plane (Request routing and proxying)
  • Control plane (Configuration management)
  • Plugin execution model (Auth, Rate Limiting)
  • High availability and scaling

Out of scope (state explicitly)

  • Detailed implementation of the backend microservices
  • Billing systems for monetization

Functional Requirements

Start by asking your interviewer whether you are designing the data plane (request proxy) or the control plane (route configuration), because they have different SLOs. Confirm plugin requirements such as JWT authentication, rate limiting, circuit breaking, and hot reload without dropping active connections.

  • Routing: Route incoming HTTP requests to the correct backend microservice based on URL paths and headers.
  • Authentication & Authorization: The gateway validates authentication tokens such as JWTs, while backend services enforce resource authorization.
  • Rate Limiting: Prevent abuse by limiting the number of requests per user or IP over a time window, applying principles from an API Rate Limiter.
  • Load Balancing & Retry: Distribute traffic across healthy backend instances and retry failed idempotent requests.
  • Dynamic Configuration: Routes and limits must be updatable without restarting the gateway (zero downtime).

Non-Functional Requirements

Your interviewer will stress-test added latency on every request and the chosen fail open vs fail closed behavior when Redis rate limiting is unavailable. They will also probe control plane vs data plane separation and zero downtime configuration reload through xDS or etcd watches.

  • Ultra Low Latency: The gateway sits in the critical path of every request, so total added overhead should remain in the single digit milliseconds, with a p99 target below 5 ms.
  • High Throughput: Must support 50M concurrent connections globally, distributed across the gateway fleet.
  • Fault Tolerance: The gateway must not crash if backend services fail, relying on the Circuit Breaker pattern.
  • Scalability: Scale horizontally by adding gateway nodes behind an L4 Load Balancer.

Capacity Estimations

At 5M RPS globally, each of 2,000 nodes handles about 2,500 RPS. The ~500K Redis operations per second figure assumes 10% synchronous global checks. Strict per request Redis enforcement would approach 5M operations per second, so local bounded state is a deliberate accuracy and latency tradeoff rather than the default for security sensitive limits.

MetricCalculationValue
Peak requests/sec (global)Given5M RPS
Gateway nodes (data plane)Given2,000
RPS per node5M ÷ 2000~2,500 RPS/node
Avg request sizeGiven2 KB
Avg response sizeAssumed2 KB
Bidirectional bandwidth (peak)5M x 2KB x 2 (request + response)~20 GB/s
Concurrent connectionsGiven50M
Concurrent connections per node50M ÷ 2000~25K/node average
Route configuration countGiven100K
Strict global Redis checks/secOne atomic decision per requestUp to 5M ops/s across sharded Redis
Sampled global Redis checks/sec5M checks x 10% reconciliation~500K ops/s with bounded overshoot

At 2,500 RPS/node on an event loop architecture using epoll, each node carries about 25K concurrent keep alive connections under the stated 50M global connection target. The added gateway latency budget is modeled as under 2 ms of proxy overhead, about 1 ms for JWT verification, and about 1 ms for the rate limit decision. This gives a roughly 4 ms internal processing budget before network and upstream time, leaving room under the p99 gateway overhead target. The control plane pushes validated configuration through xDS or etcd watches, allowing 100K routes to reload in under 100 ms without dropping active connections.

Architecture Diagram

At 5M RPS globally, the API gateway introduces overhead on every request, which means added latency must stay in single digit milliseconds. The control plane owns route CRUD and plugin configuration, whereas the data plane manages connection handling, JWT verification, rate limiting, and upstream proxying with hot reload so configuration changes never drop active keep alive sessions.

The WAF and anycast edge absorb raw DDoS attacks before traffic reaches the origin, after which the gateway terminates client TLS, enforces authentication and quotas, and forwards requests to microservices over mTLS with SPIFFE workload identities. Redis holds distributed rate limit state, while optional local counter caches trade strict precision for lower latency when the endpoint policy permits bounded overshoot.

Loading...

In the room

Separate the control plane (route CRUD, admin API) from the data plane (high throughput proxy) in your first sketch. Clarify that rate limiting can fail open when policy permits it, while JWT authentication fails closed.

Component Deep Dives

1. The Data Plane (Asynchronous I/O)

Request flow crosses four concerns in sequence: terminating TLS, authenticating, rate limiting, and routing to a healthy upstream, with each step implemented as an inline plugin on the event loop proxy core. The sections below connect asynchronous I/O architecture, zero downtime configuration propagation, distributed rate limiting, stateless JWT authentication, circuit breaking, and the edge security layers that sit in front of the gateway.

An event loop architecture with epoll or kqueue handles thousands of concurrent connections per worker thread rather than a thread per connection model, which quickly exhausts CPU through context switching.

To achieve high throughput (10,000+ requests/sec per node), gateways like Nginx (which Kong is based on) or Envoy do not use a thread per connection model (like older Apache or Tomcat servers), which would exhaust CPU and RAM via context switching overhead. Instead, they use an Event Loop architecture (epoll/kqueue) with non-blocking I/O. A single worker thread handles thousands of concurrent connections efficiently, dramatically reducing memory overhead.

2. Dynamic Configuration (Control Plane) ⭐

Dynamic configuration via xDS streams or etcd watches lets gateway nodes receive validated routing snapshots without restarting. The proxy atomically swaps the active configuration in memory, while worker generations can drain existing connections before retiring old state.

If an engineering team deploys a new microservice, the gateway needs to know the new route. If we have 50 gateway nodes, restarting them to update a YAML config file drops active connections and causes an outage. Modern gateways solve this by separating the control plane from the data plane. The control plane stores durable configuration, validates it, and publishes versioned snapshots through xDS or an equivalent distribution channel. Data plane nodes consume those snapshots locally instead of querying PostgreSQL on the request path.

YAML
# Versioned configuration snapshot published by the Control Plane
version: 42
route:
  name: payment-service-route
  path: /api/v1/payments/*
  methods:
    - GET
    - POST
  backend:
    service: payment-service
    port: 8443
  plugins:
    - rate-limiting
    - jwt

# The Control Plane distributes this validated snapshot through xDS.
# Data Plane nodes atomically swap the active snapshot in memory.
# Existing connections continue draining on the previous snapshot.

3. Distributed Rate Limiting

Global rate limiting requires shared state across gateway nodes when the quota must be exact. Redis with Lua provides atomic Token Bucket decisions for that strict mode. Optional local token buckets can reduce latency when the policy accepts bounded global overshoot and asynchronous reconciliation.

Rate limiting mitigates application abuse and request floods. Volumetric DDoS is primarily handled at the CDN or anycast WAF edge. For globally enforced limits, the gateway shares Redis state across nodes so an attacker cannot bypass a limit simply by reaching another node.

Loading...

Optimization (Local Caching vs Accuracy): A synchronous Redis round trip for every request can add 1 to 2 ms of latency. A local token bucket with a preallocated per node budget can reduce that cost, with asynchronous reconciliation to Redis. This trades exact global enforcement for bounded overshoot, so it is appropriate only for policies that explicitly accept approximate global enforcement.

4. Authentication (Stateless JWT)

Stateless JWT validation at the gateway avoids a database round trip on every request by verifying signatures offline using cached JWKS keys and validating the issuer, audience, expiry, not before, and required scopes or roles. Refresh JWKS on an unknown key identifier and allow key overlap during key rotation.

Making a database call to the User Service to validate a session token for every API request creates unnecessary dependency and latency. Instead, the gateway uses stateless JSON Web Tokens (JWT).

Loading...

Because the JWT is cryptographically signed with asymmetric keys such as RS256, the gateway can verify the token's authenticity offline using cached JWKS keys. It trusts only validated claims such as user_id and roles. The gateway first strips any client supplied identity headers, then injects its own authenticated identity headers over the mTLS protected upstream connection. Backend services still enforce authorization based on the trusted identity and resource policy.

5. Circuit Breaker Pattern

Circuit breakers prevent a failing upstream from exhausting the gateway connection pool and causing cascading failure across the entire API platform.

If the "Order Service" goes down, the gateway might queue up thousands of requests waiting for it to respond, eventually exhausting the gateway's own connection pool and causing a catastrophic cascading failure across the entire API platform.

  • The Circuit Breaker tracks failure rates, such as 50% failures occurring in 10 seconds.
  • If the threshold is crossed, the circuit Opens. The gateway instantly rejects new requests to the Order Service with 503 Service Unavailable without even trying the network call, giving the backend time to recover.
  • After a timeout, it shifts to Half-Open, allowing a small percentage of test requests through. If they succeed, it Closes the circuit. Failed test requests reopen it.

6. mTLS Upstream & WAF at the Edge

Defense in depth spans three layers. The edge WAF blocks OWASP attacks, bot floods, and volumetric traffic. Gateway plugins enforce authentication, request limits, and circuit breaking. mTLS and workload identity protect the gateway to service path. Rate limiting can fail open where endpoint policy permits it, while authentication failures fail closed.

mTLS (service to service):
  - The edge security layer filters public traffic before the gateway.
  - The gateway routes to upstream microservices over mTLS using SPIFFE/SPIRE or an equivalent workload PKI.
  - The gateway validates the workload identity in the certificate against the expected service identity.
  - Backend services accept gateway traffic only over the authenticated channel.
  - This limits lateral movement if a pod is compromised.

WAF / edge protection:
  Layer 1: Anycast CDN/WAF
    OWASP Top 10 rules, SQLi/XSS blocking, geo-blocking, bot scoring, and volumetric DDoS absorption.
    If the WAF terminates TLS, it re-encrypts traffic to the gateway. Otherwise the gateway terminates client TLS.
  Layer 2: Gateway plugins
    IP rate limiting, JWT validation, request size caps, and circuit breaking.
  Layer 3: Backend
    Resource authorization remains in the backend service. The gateway authenticates the caller but does not replace resource authorization.

Rate limiting policy:
  - Strict global limits use an atomic Redis decision.
  - Latency sensitive endpoints may use bounded local token buckets with asynchronous reconciliation when the policy permits controlled overshoot.
  - Availability oriented endpoints may fail open on Redis outage. Security sensitive endpoints may fail closed.

API Design

Typed Gateway Domain Contract

The control plane exposes typed route and plugin configuration objects, while the data plane remains a high throughput proxy interface rather than a business API.

TYPESCRIPT
export type RouteId = string;
export type ServiceId = string;
export type HttpMethod = "GET" | "POST" | "PUT" | "PATCH" | "DELETE" | "HEAD" | "OPTIONS";
export type PluginType = "jwt" | "rate_limiting" | "circuit_breaker";

export interface RateLimitPolicy {
  limit: number;
  windowSeconds: number;
  scope: "ip" | "user" | "organization" | "route";
}

export interface RouteConfig {
  routeId: RouteId;
  pathPrefix: string;
  methods: HttpMethod[];
  serviceId: ServiceId;
  timeoutMs: number;
  plugins: PluginType[];
  rateLimit?: RateLimitPolicy;
  version: number;
}

export interface RouteCreateRequest {
  idempotencyKey: string;
  route: Omit<RouteConfig, "routeId" | "version">;
}

export interface RouteUpdateRequest {
  expectedVersion: number;
  route: Omit<RouteConfig, "routeId" | "version">;
}

Control Plane Admin API

The Gateway exposes a private administrative interface for route and plugin configuration. It is protected by workload identity or mTLS plus RBAC and is never placed on the public request path.

HTTP
// Control Plane Admin API: Add a new route
POST /admin/api/routes
Idempotency-Key: route-create-8f3c1b
Content-Type: application/json

{
  "name": "payment-service-route",
  "paths": ["/api/v1/payments"],
  "methods": ["GET", "POST"],
  "service": {
    "host": "payment.internal",
    "port": 8443
  },
  "plugins": [
    { "name": "rate-limiting", "config": { "limit": 100, "window": "1m" } },
    { "name": "jwt" }
  ]
}

Standard Data Plane Errors

The data plane uses stable HTTP error semantics so clients can distinguish authentication failures, quota rejection, missing routes, upstream failures, and gateway timeouts.

HTTP
HTTP/1.1 401 Unauthorized
{
  "error": "invalid_or_expired_token"
}

HTTP/1.1 403 Forbidden
{
  "error": "insufficient_permissions"
}

HTTP/1.1 404 Not Found
{
  "error": "route_not_found"
}

HTTP/1.1 409 Conflict
{
  "error": "config_version_conflict"
}

HTTP/1.1 429 Too Many Requests
Retry-After: 60
{
  "error": "rate_limit_exceeded"
}

HTTP/1.1 502 Bad Gateway
{
  "error": "upstream_connect_failed"
}

HTTP/1.1 503 Service Unavailable
{
  "error": "upstream_unavailable"
}

HTTP/1.1 504 Gateway Timeout
{
  "error": "upstream_timeout"
}

Data Model

The Gateway requires two different storage roles. PostgreSQL holds durable control plane configuration, while Redis holds high throughput, volatile rate limiting state and fast distributed coordination data.

SQL
CREATE TABLE routes (
  id UUID PRIMARY KEY,
  name VARCHAR(128) NOT NULL,
  path_prefix VARCHAR(512) NOT NULL,
  upstream_service_id UUID NOT NULL,
  version BIGINT NOT NULL,
  updated_at TIMESTAMPTZ NOT NULL
);

CREATE TABLE plugins (
  id UUID PRIMARY KEY,
  route_id UUID NOT NULL REFERENCES routes(id) ON DELETE CASCADE,
  plugin_type VARCHAR(64) NOT NULL,
  config JSONB NOT NULL
);
REDIS
# Configuration distribution
# etcd may hold the published snapshot, while xDS is the delivery protocol.
gateway/config/{version} = validated route + plugin snapshot

# Data plane rate limiting state
# Lua atomically consumes tokens, refills according to elapsed time, and refreshes TTL.
rate_limit:{scope}:{client_id}:{route_id} = {
  tokens,
  last_refill_ms
}
TTL = expire after idle period

Fault Tolerance

Failure CaseSystem Solution Design
Redis Rate Limiter CrashFail open only where endpoint policy permits availability to win over strict quota enforcement. Use local bounded fallback state and emit telemetry. Security sensitive endpoints can fail closed.
Control Plane CrashData Plane nodes retain the last validated routing snapshot in local memory. Existing traffic continues while configuration changes pause until a healthy control plane resumes publication.
Stale Configuration SnapshotGateways apply monotonically increasing versions, reject out of order snapshots, and canary new versions before global rollout. A previous validated snapshot remains available for rollback.
Backend Service TimeoutGateway enforces strict timeouts and bounded retries. Retry only safe or explicitly idempotent requests before a response is committed, and only against healthy endpoints.
Plugin Runtime FailureBound plugin CPU, memory, and execution time. A faulty plugin instance is isolated from the proxy event loop, disabled or rolled back, and reported through control plane health signals.
WAF or Edge Dependency FailureKeep multiple edge providers or regional edge capacity where availability requirements justify it. Preserve gateway capacity for traffic that remains admitted after edge protection.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless interviewing for a staff or principal role.

    • Separate control plane (config CRUD, admin API) from high throughput data plane proxy (5 min)
    • Data plane uses event loop with non-blocking I/O rather than thread per connection (6 min)
    • Dynamic config via xDS stream or etcd watch (5 min)
    • Rate limiting: Redis Lua globally, with optional local bounded cache (5 min)
    • JWT validation offline at gateway, injecting X-User-Id downstream (4 min)
  • Separate control plane (config CRUD, admin API) from data plane (high throughput proxy), since they operate with different SLOs.
  • Data plane uses an event loop with non-blocking I/O rather than thread per connection.
  • Dynamic config via xDS stream or etcd watch enables hot reload without dropping connections.
  • Rate limiting: Redis Lua globally, with an optional local bounded token bucket where bounded overshoot is acceptable.
  • JWT validation offline at gateway, injecting X-User-Id headers for backends.
  • WAF at the CDN edge absorbs volumetric and application attacks, while the gateway adds mTLS to upstream microservices.
  • Track per route p50, p95, and p99 latency, upstream error rate, circuit breaker state, rate limit decisions, active connections, configuration version skew, and plugin execution time.
  • Common pitfall: a synchronous Redis call on every request adds latency, so use exact Redis enforcement where required and local bounded state only where policy permits approximate global enforcement.

Engineering Trade-offs

API Gateway vs Service Mesh (Istio)

An API Gateway handles "north south" traffic (external clients entering the datacenter). A Service Mesh handles "east west" traffic (internal microservices talking to each other). While they share technologies (Envoy is often used for both), they serve different purposes. Gateways focus on Edge security (WAF, OAuth), while Service Meshes focus on internal mTLS, distributed tracing, and complex routing between hundreds of internal pods.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...