System Design Problem

Design a Service Discovery System

Commonly Asked By:HashiCorpNetflixAWSGoogle

Interview Setup

Interview Prompt

Design a Consul-like service discovery system for 5,000 services with 100,000 instances, serving 1M lookups/sec (mostly from client cache) with health checking and watch-based updates.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Client-side discovery (app picks instance) or server-side (load balancer proxy)?
  • Client-side scales lookups to clients
  • server-side centralizes routing but is a bottleneck.
DNS-based, API-based, or both?
  • DNS has TTL caching (stale)
  • API with watch gives instant updates but requires SDK.
Health check ownership — central server or local agent?
  • 10K health checks/sec
  • local agent per node avoids central bottleneck.
Multi-datacenter — single global registry or per-DC?
  • Global = cross-DC latency on writes
  • per-DC = eventual cross-DC consistency.

Scope

In scope

  • Client-side vs server-side discovery
  • Health check mechanisms
  • DNS-based vs registry (Consul/etcd)
  • Self-registration
  • Stale entry eviction
  • Capacity estimation with shown math

Out of scope (state explicitly)

  • Full service mesh control plane
  • Application business logic in downstream services
  • Building a custom service registry from scratch when managed options exist

Functional Requirements

Start by confirming client-side vs server-side discovery and health-check depth. Ask about namespaces, watch/subscribe semantics, and DNS vs API-based lookup.

  • Services register themselves on startup (name, address, port, health endpoint, metadata)
  • Services deregister on shutdown (graceful) or get auto-deregistered on failure
  • Clients look up healthy instances of a service by name
  • Health checking: periodic checks (HTTP, TCP, gRPC) to detect unhealthy instances
  • Support for multiple environments/namespaces (prod, staging, per-tenant)
  • Service metadata: version, region, weight, canary flags
  • Watch/subscribe: clients get notified of changes (new instances, removals)
  • DNS-based and API-based discovery
  • Load balancing integration: return instances in weighted/round-robin order

Non-Functional Requirements

Your interviewer will care most about five-nines availability and sub-millisecond lookups. If discovery is down, no service can call another — client-side caching with server push is the standard answer.

  • High Availability: 99.999%: if discovery is down, no service can call another
  • Low Latency: Lookup in < 1ms (client-side caching with server push)
  • Consistency: Eventually consistent with < 5s propagation of changes
  • Scalability: 100K+ service instances, 1M+ lookups/sec
  • Fault Tolerance: Continue operating during network partitions
  • Zero Downtime: Rolling updates, cluster resizing without disruption

Capacity Estimations

Run this math before you size the registry cluster. Service instances and lookup QPS tell you cache hit ratio needs; health-check frequency drives agent CPU on each node.

MetricCalculationValue
ServicesGiven5,000 distinct names
Instances (total)Given100,000
Registrations / secDerived from daily volume ÷ 86400 (+ peak factor)50 (deploys, autoscaling)
Health checks / secDerived from daily volume ÷ 86400 (+ peak factor)10K/sec
Lookups / secDerived from daily volume ÷ 86400 (+ peak factor)1M (with client caching, most served locally)
Watch subscribersGiven50K (one per client instance)
Metadata per instanceGiven~500 bytes

Architecture Diagram

Walk your interviewer through the registry and health-check loop first. A Consul-like service discovery system with a Raft-consistent registry, health checking by local agents, gossip-based anti-entropy, and client-side caching with watch support for instant updates — I draw registration, health probing, and client watch as three separate flows.

Loading...

Component Deep Dives

In the room: client-side vs server-side discovery is the first fork — Consul/Eureka vs load-balancer fronting changes latency and failover.

Next we walk through each box on the diagram. I start with server-side vs client-side discovery because it determines whether clients talk to a load balancer or pick instances directly.

Client-side discovery with Consul or Eureka is the interview default — no extra hop and locality-aware routing.

Server-Side vs Client-Side Discovery

ApproachHowProsCons
Server-Side (ELB/K8s Service)Client → LB → BackendClient is simple (just knows LB address)Extra hop, LB is bottleneck
Client-Side (Consul/Eureka) ⭐Client queries registry → picks instance directlyNo extra hop, locality-aware decisionsPer-language library needed
Service Mesh (Envoy/Istio)Client → local Envoy sidecar → BackendLanguage-agnostic, rich featuresSidecar resource overhead
Kubernetes DNSsvc.namespace.svc.cluster.local → ClusterIPBuilt-in, zero setupLimited health checking, DNS caching

Health check depth is a common trap — shallow HTTP checks every 5s for SD; never use deep dependency checks at high frequency.

Health Check Strategies

Level 0: TCP Check: Is port open? Catches process crash. < 1ms.

Level 1: HTTP Shallow ⭐: GET /healthz → 200. Catches HTTP server failure. < 5ms.

Level 2: HTTP Deep: Checks DB connection, cache, external APIs. Catches dependency failures but risks cascading. 10-100ms.

Level 3: Liveness + Readiness (K8s): /healthz/live (is process stuck?) vs /healthz/ready (can it handle traffic?). Service starting: live=true, ready=false. DB lost: live=true, ready=false. Deadlocked: live=false → K8s KILLS pod.

Recommendation: SD health check = Level 1 (shallow, every 5s). Application readiness = Level 2 (deep, every 30s). K8s liveness = Level 0 or 1. NEVER use deep health checks at high frequency with SD.

Anti-Entropy & Convergence

Consul uses SERF gossip protocol: nodes share state updates peer-to-peer. Anti-entropy sync every 30s: each agent syncs full state with servers. On conflict: latest write wins (LWW with Lamport timestamps). All nodes converge within seconds (< 5s). Health checks are NOT centralized on servers: each local agent health-checks services on its node and gossips status to servers.

Graceful Deployment with Service Discovery

Blue-Green: Deploy v2, register with tag "v2", gradually shift weight v1→v2, deregister v1. Canary: Register 1 canary with tag "canary", route 5% to canary via metadata, monitor, promote or deregister. Connection Draining: Before deregistration, mark as "draining": SD stops sending new requests, waits 30s for in-flight requests to complete, then fully deregisters.

Event Bus Design (Kafka)

Topic: service-registration-events
  Partitions: 32 (partition by service_name — preserves per-service ordering)
  Events: REGISTER, DEREGISTER, HEARTBEAT_MISSED, INSTANCE_UNHEALTHY, INSTANCE_HEALTHY
  Retention: 7 days (audit + post-mortem replay)
  Producers: Service Registry on Raft commit
  Consumers: mesh control-plane sync (Istio/Linkerd), CMDB audit sink, PagerDuty on mass dereg

Topic: service-catalog-changes
  Low volume; schema/version metadata updates
  Consumers: client SDK config refresh, deployment pipeline validators

Registration path: instance POST /register → Raft write → push watch to subscribers (< 100ms)
  Kafka is audit + cross-system fan-out only — live lookups use watch/long-poll, not Kafka

API Design

# Registration
PUT /v1/agent/service/register
{
  "id": "payment-svc-i-abc123",
  "name": "payment-svc",
  "address": "10.0.1.42",
  "port": 8080,
  "tags": ["v2.3.1", "canary"],
  "meta": {"region": "us-east", "weight": "100"},
  "check": {
    "http": "http://10.0.1.42:8080/healthz",
    "interval": "10s",
    "timeout": "3s",
    "deregister_critical_service_after": "60s"
  }
}

# Deregistration
PUT /v1/agent/service/deregister/{service_id}

# Lookup healthy instances
GET /v1/health/service/{service_name}?passing=true&near=_agent&tag=v2.3.1

# Watch for changes (long-poll / blocking query)
GET /v1/health/service/{service_name}?passing=true&index=42&wait=30s
→ Returns immediately if index > 42 (changes occurred)
→ Blocks up to 30s if no changes

# DNS interface
dig payment-svc.service.consul SRV
→ 1 0 8080 i-abc123.node.dc1.consul.
→ 1 0 8080 i-def456.node.dc1.consul.

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

Service Registry (In-Memory + Raft-Replicated)

JSON
{
  "services": {
    "payment-svc": {
      "instances": {
        "payment-svc-i-abc123": {
          "id": "payment-svc-i-abc123",
          "address": "10.0.1.42",
          "port": 8080,
          "tags": ["v2.3.1", "canary"],
          "meta": {"region": "us-east", "weight": "100"},
          "health": "passing",
          "last_heartbeat": "2026-03-14T10:05:00Z",
          "registered_at": "2026-03-14T08:00:00Z",
          "check": {
            "type": "http",
            "endpoint": "http://10.0.1.42:8080/healthz",
            "interval_sec": 10,
            "consecutive_failures": 0
          }
        }
      }
    }
  }
}

Consul vs etcd vs ZooKeeper Comparison

FeatureConsuletcdZooKeeper
ConsensusRaftRaftZAB
Health checks✅ Built-in❌ Must build❌ Must build
DNS interface✅ Built-in❌ External❌ External
Multi-DC✅ WAN gossip❌ Single cluster❌ Single cluster
WatchLong-pollgRPC watchZooKeeper watches
Used byHashiCorp stackKubernetesKafka, HBase

Fault Tolerance

SD Cluster Failure

Defense layers: 1) Raft cluster with 5 nodes (tolerates 2 failures). 2) If majority lost → cluster read-only. 3) Clients use local cache → continue routing to last-known instances. 4) Client-side health checking as backup. 5) On recovery: services re-register. Key principle: Service-to-service calls must NOT depend on SD availability. Client caches must be robust enough to last through SD outages. Graceful degradation: stale routing > no routing.

Race Conditions in Service Discovery

Race 1: Stale Cache During Rolling Deployment: Service B1 deregisters, but A's cache still has B1 (TTL: 30s). A sends request to B1 → connection refused. Mitigation: client-side retry with next instance, circuit breaker, watch-based (not poll-based) cache refresh, and connection draining (B1 sends "draining" status before deregistering).

Race 2: Split-Brain During Network Partition: Minority partition (2/5) becomes read-only. Majority partition (3/5) elects new leader. Clients in minority serve stale but valid data. When partition heals: minority catches up via Raft log replay. Key: clients MUST work with stale data during partition.

Race 3: Zombie Instance: Process hangs but TCP port is open. Health check times out. After 3 consecutive timeouts → deregistered. Window: 30 seconds. Solution: passive health check (if actual request fails, immediately mark unhealthy) + shallow every 5s + deep every 30s.

Additional Considerations

Client-Side Caching: How It Actually Works

Every service maintains a LOCAL cache of SD data. Update strategy (layered): 1) WATCH: blocking query to Consul/etcd (instant notification on change, < 100ms). 2) POLL: every 30s, full refresh as backup. 3) ON-FAILURE: if request to instance fails, immediately refresh cache for that service. 4) DISK: persist cache to disk on exit → load on restart (survives SD outage). Recommended: watch for real-time + TTL=120s as safety net.

Self-Registration vs Third-Party Registration

Self-Registration (Eureka, Consul agent): Service registers itself on startup, sends heartbeats. ✅ Service knows its own state best. ❌ Every service needs registration logic. Third-Party Registration (Kubernetes, Registrator): External component watches for new instances and registers/deregisters on behalf of services. ✅ Services don't need discovery-aware code. ❌ Registrator is another component to manage.

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Discovery patterns: server-side vs client-side vs mesh (5 min)
    • Health check tiers: shallow /healthz every 5s (6 min)
    • Client cache: watch + poll + on-failure refresh (5 min)
    • Connection draining before deregistration (5 min)
    • Raft-backed registry tolerates node failures (4 min)
  • Contrast discovery patterns: server-side (LB/K8s Service), client-side (Consul/Eureka), and service mesh (Envoy sidecar) — pick based on latency vs simplicity trade-off.
  • Health check tiers: shallow HTTP /healthz every 5s for SD; deep dependency checks only at low frequency — never deep-check at SD cadence.
  • Client-side cache with layered refresh: watch (blocking query) + poll backup + on-failure immediate refresh + disk persistence for SD outage survival.
  • Connection draining before deregistration: mark instance as draining → stop new requests → wait 30s for in-flight → fully deregister.
  • Raft-backed registry (Consul/etcd) tolerates node failures; clients must degrade gracefully to stale cache rather than fail entirely.
  • DNS SRV records as a fallback interface — useful for legacy clients that cannot embed a discovery SDK.
  • Common pitfall: deep health checks that verify database connectivity at 5s intervals — one DB blip deregisters all healthy instances and causes a cascading outage.

Engineering Trade-offs

Service discovery trades consistency against availability — CP registry vs AP gossip with stale cache fallback.

DNS-Based vs API-Based Service Discovery

AspectDNS-BasedAPI-Based ⭐
Universality✓ Every language supports DNS natively✗ Requires client library per language
Data richness✗ IP + port only (SRV records help)✓ Full metadata: tags, health, weight
Freshness✗ DNS caching (TTL-dependent)✓ Instant via long-poll/watch
Smart routing✗ Round-robin only✓ Weighted, canary, version-based

Recommendation: Use DNS for simple cases (Kubernetes internal). Use API for microservices with advanced routing needs. Many systems use both: DNS for initial discovery, API for watch/health.

Health Check Depth Spectrum

Level 0 (TCP): < 1ms, catches process crash.

Level 1 ⭐ (HTTP Shallow): < 5ms, catches HTTP failure. Safe for SD at 5s intervals.

Level 2 (HTTP Deep): 10-100ms, catches dependency failures but risks cascading (DB slow → all services marked unhealthy).

Level 3 (Liveness + Readiness): K8s pattern. SD health check = Level 1. Application readiness = Level 2 (every 30s). K8s liveness = Level 0 or 1. NEVER use deep checks at high frequency.

Kubernetes EndpointSlice

In Kubernetes, the legacy Endpoints object scaled poorly (one object per Service, full rewrite on every pod change).EndpointSlice shards endpoints into slices of up to 100 backends each, reducing apiserver churn at 10K+ pod fleets. kube-proxy or CNI dataplane watches EndpointSlices directly; CoreDNS resolves svc.namespace.svc.cluster.local to ClusterIP while the dataplane load-balances to slice endpoints. For headless services or gRPC long-lived connections, clients may watch EndpointSlices via the discovery API for instant endpoint updates without DNS TTL delay.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...