System Design Problem

Design a Content Delivery Network (CDN)

Commonly Asked By:NetflixCloudflareAmazonGoogle

Interview Setup

Interview Prompt

Design a Content Delivery Network (CDN) that caches static and cacheable dynamic content at edge locations worldwide. Origin servers sit in one or few regions, while users worldwide obtain low-latency responses with a high cache hit ratio.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Static only, or dynamic and cacheable API responses too?Static content such as images, JavaScript, and video segments requires long TTLs and immutable URLs, whereas dynamic content requires short TTLs, edge-side includes, or cache keys segmented by Vary headers.
Invalidation latency requirement?Immediate purge in under 30 seconds globally vs eventual expiration via TTL requires a fundamentally different control plane design.
Single tenant or multi-tenant CDN-as-a-service?Multi-tenant architectures require configuration isolation, per-tenant cache namespaces, and fair queuing on origin fetches.
What is peak bandwidth and request rate?A 10 Tbps peak drives PoP count, edge server sizing, and origin shield necessity, while 50M requests per second shapes connection handling architecture.

Scope

In scope

  • PoP architecture and cache hierarchy
  • Geo-routing and DNS or load balancing to nearest edge
  • Origin shield and consistent hashing
  • Cache invalidation (purge API, TTL, surrogate keys)
  • Hit ratio optimization and capacity math

Out of scope (state explicitly)

  • Full DNS provider design (assume GeoDNS exists)
  • DDoS mitigation internals (WAF and scrubbing centers treated as external modules)
  • Video-specific adaptive bitrate streaming (covered in Video Streaming Platform)
  • Origin application design

Functional Requirements

Start by asking your interviewer about static vs dynamic content, cache invalidation requirements, and geographic coverage. Origin shielding and TLS termination are almost always in scope.

  • Cache and serve static content including images, videos, CSS, JavaScript, and web fonts from edge servers closest to the user.
  • Route user requests to the optimal edge server with the lowest network latency.
  • Fetch content from the upstream origin server on a cache miss when content is not cached locally at the edge.
  • Support on-demand cache invalidation and purge operations across all edge locations.
  • Terminate SSL/TLS encryption directly at the edge to minimize handshake latency.
  • Provide real-time telemetry and analytics covering bandwidth usage, cache hit ratios, and regional latency.
  • Support custom caching policies including TTL specifications, cache key customization, and query string filtering.

Non-Functional Requirements

Interviewers evaluate cache hit ratio and edge latency most heavily on this problem. Every design choice, including routing mechanisms, TTL tiers, and origin shield clustering, exists to keep origin load minimal and maintain p99 edge response latency under 50 milliseconds.

  • Low Latency: Serve content in under 50 ms from edge servers compared to 200+ ms from origin.
  • High Availability: Deliver 99.99% availability, because CDN outages degrade user experiences worldwide.
  • Massive Scale: Deliver 10 Tbps of peak aggregate egress bandwidth globally (with architecture headroom scaling to 100+ Tbps for future hyperscale scenarios).
  • Global Coverage: Deploy across 200+ Points of Presence (PoPs) worldwide.
  • Cache Efficiency: Maintain greater than 95% cache hit ratio for popular content.
  • Fault Tolerant: Mask individual edge server and PoP failures transparently from clients.
  • DDoS Resilient: Absorb volumetric attacks and Layer 7 floods directly at the edge.

Capacity Estimations

Run this math before sizing PoP storage, because traffic volume and average object sizes dictate the required edge server count and cache hit ratio targets. At our baseline 50M requests/sec and a 95% edge cache hit ratio, edge misses generate 2.5M requests/sec. Passing through an origin shield with a 90% hit ratio reduces origin load to just 250K requests/sec (0.5% of total traffic), whereas a degraded 90% edge hit ratio without a shield would subject the origin to 5M requests/sec.

MetricCalculationValue
Total PoPsGiven (baseline topology assumption)200+ (50 Tier 1 + 150 Tier 2)
Servers per PoPGiven (assumption documented in value)50-500 (varies by PoP tier)
Total edge serversWeighted average across 200+ PoPs20,000-50,000
Peak bandwidthPrimary baseline egress (headroom for 100+ Tbps hyperscale)10 Tbps
Requests / secFrom Requests / day ÷ 86400 (+ peak factor in value)50M (globally)
Cache storage per PoPGiven (assumption documented in value)100 TB
Total cached content200+ PoPs x 100 TB (scaling up to 30 PB across 300 PoPs at hyperscale)20 PB
Edge misses (95% hit ratio)50M requests/sec x 5% edge miss rate2.5M/sec
Origin requests with shield (90% hit)2.5M misses x 10% shield miss rate (0.5% of total traffic)250K/sec
Degraded origin load (no shield)50M requests/sec x 10% unshielded edge miss rate (stress scenario)5M/sec

Architecture Diagram

In the interview room, clarify cache invalidation requirements early. While TTL-only expiration is architecturally simpler, on-demand active purges necessitate a dedicated asynchronous control plane.

Walk your interviewer through the architecture following the user request path. DNS or Anycast directs users to their nearest edge PoP. On a cache hit, the edge node serves content immediately from local SSD storage or memory. On a cache miss, the PoP proxies through a regional origin shield cluster so that hundreds of PoPs do not overload the origin with duplicate fetches. Purge and warming requests flow asynchronously through a Kafka-backed control plane, guaranteeing that invalidation workflows never stall synchronous edge request handling.

Loading...

Component Deep Dives

Walk through each layer of the architecture sequentially. Begin with client geo-routing, detail local edge caching mechanics, and examine the origin protection tier.

DNS-Based Routing (GeoDNS)

GeoDNS offers an accessible routing mechanism, though you should explain its trade-offs before introducing Anycast.

  • The user's recursive DNS resolver issues a query for cdn.example.com.
  • The CDN's authoritative GeoDNS server maps the client resolver's IP address to a geographic region.
  • The DNS server responds with the IP address of the closest, lowest-latency edge PoP.
  • Pros: Simple to deploy, widely supported across all client devices, and enables granular per-client traffic steering.
  • Cons: DNS resolver caching can delay failover by minutes, and resolver IP locations do not always match the actual user's location (though EDNS Client Subnet helps mitigate this).

Anycast Routing (Alternative and Complementary)

Anycast avoids DNS TTL-based failover delays by allowing multiple PoPs to announce the same IP address through BGP. Cloudflare relies primarily on Anycast, whereas AWS CloudFront combines Anycast with GeoDNS.

  • Multiple global PoPs announce identical IP prefixes to neighboring Autonomous Systems using BGP.
  • Internet routing automatically delivers user packets to the topologically closest PoP along the shortest AS path.
  • Pros: Anycast avoids DNS TTL-based failover delays, natively absorbs volumetric DDoS traffic across multiple PoPs, and automates network-level route withdrawal, though actual failover depends on health-check withdrawal and BGP route convergence (typically 30 to 60 seconds).
  • Cons: Provides less granular per-user traffic steering and remains subject to ISP routing policies and route flappings.
  • Real-world implementation: Cloudflare deploys Anycast globally, whereas AWS CloudFront combines Anycast with GeoDNS-based routing.

Edge Server (Cache Node)

Multi-tiered caching inside edge servers is what keeps read latency within the sub-50 millisecond SLO at global scale.

  • Reverse proxy: NGINX, Envoy, or custom software terminates TLS, inspects headers, and executes caching rules.
  • Multi-tier storage: High-speed RAM stores the hottest objects (64 GB per node), NVMe SSDs retain warm content (2 to 10 TB per node), and high-capacity HDDs retain colder assets (50+ TB).
  • Cache lookup path: The node hashes the cache key, checks memory first, falls back to SSD, checks HDD if configured, and declares a cache miss if absent.
  • Eviction policy: Least Recently Used (LRU) or TwoQ eviction discards cold entries when storage reaches high-water marks.
  • Consistent hashing: Consistent hashing within the PoP establishes a primary cache owner per key to prevent redundant caching across multiple local nodes. For extremely popular (hot) objects, the edge proxy can optionally replicate cached items across a small secondary set of nodes (such as 2 to 3 replicas) to alleviate single-node CPU and network concurrency bottlenecks.

Origin Shield (Mid-Tier Cache)

An origin shield serves as an intermediate regional caching layer that protects upstream origin clusters from concurrent cache miss waves.

  • Problem: Without a shield, 200+ edge PoPs that miss on the same object would simultaneously hit the origin with 200+ redundant requests.
  • Mechanism: Regional shields sit between edge PoPs and origin data centers, aggregating misses across all local PoPs in that continental zone.
  • Impact: Regional shields aggregate traffic across hundreds of edge PoPs, ensuring the origin interfaces with a small set of continental shield aggregates rather than 200+ independent edge sources. Concurrent duplicate fetches are collapsed, and with a 90% shield hit ratio, actual origin traffic drops from 2.5M edge misses/sec down to 250K origin requests/sec (a 90% reduction in origin load).
  • Deployment footprint: Strategically placed across 3 to 5 continental regions including US-East, US-West, Europe, and Asia-Pacific.

Cache Key Configuration

Cache key design directly governs the cache hit ratio. Inefficient key schemas fragment cached objects and turn the CDN into an expensive pass-through proxy. Crucially, personalized and authenticated responses must default to Cache-Control: private, no-store. Raw session tokens, Authorization headers, or user cookies must never be incorporated into cache keys by default, as doing so destroys cache hit ratios and introduces severe cross-user data leak risks. If authenticated responses are explicitly cacheable, the edge should segment cache keys only on coarse, non-sensitive operational dimensions such as user role (X-User-Role) or subscription tier (X-User-Tier).

HTTP
# Default cache key format: scheme + host + path + query_string
GET https://cdn.example.com/images/logo.png?v=2 HTTP/1.1

# Normalized cache key rules:
# 1. Strip irrelevant query strings for static media assets
# 2. Safe authenticated caching: Personalized responses default to Cache-Control: private, no-store.
#    Never include raw session tokens or user cookies in cache keys (destroys hit ratios and risks cross-user data leaks).
#    For explicitly cacheable responses, segment keys only by non-sensitive dimensions (e.g., X-User-Tier, X-User-Role).
# 3. Include Accept-Encoding header (gzip, brotli) to serve matched compression
# 4. Include Accept-Language or device headers when serving localized variants

Cache Control Headers

HTTP Cache-Control and Vary headers instruct edge proxies how to cache, revalidate, and serve stored representations.

HTTP
Cache-Control: public, max-age=86400, s-maxage=604800
# public:    Any downstream shared cache or CDN can store this response
# max-age:   Browser client cache TTL (1 day = 86,400 seconds)
# s-maxage:  Shared CDN edge cache TTL (7 days = 604,800 seconds)

Cache-Control: private, no-store
# private:   Restricted to end-user client, so intermediate CDNs must not cache
# no-store:  Never write to disk or cache memory, and always re-fetch from origin

Vary: Accept-Encoding
# Instructs CDN to maintain separate cache entries for gzip vs brotli

Control Plane: Purge and Warm (Kafka)

Cache invalidation and pre-warming operate as asynchronous control plane workflows so that they never block edge request handling. The Purge API accepts an exact URL, surrogate key, or pattern, publishes the task to the cache-purge topic, and each PoP consumer group deletes matching local cache keys.

YAML
topic_cache_purge:
  partitions: 16 # Partitioned by tenant_id or PoP region
  partition_key: "purge_id"
  retention: "24h" # Audit trail for invalidation jobs
  producers:
    - "Purge API (administrative purge, surrogate-key invalidation, origin webhooks)"
  consumers:
    - "PoP purge workers (each PoP consumer group invalidates matching local cache keys)"
  propagation_sla: "Under 30 seconds globally (p99) for exact URLs, under 5 seconds for regional tag and pattern sweeps (local edge eviction under 1 second once received)"

topic_cache_warm:
  partitions: 8
  producers:
    - "Warm API and origin push hooks for new releases and viral content pre-warming"
  consumers:
    - "Regional warm workers (fetch from origin shield and populate edge SSD or memory)"

topic_edge_access_logs:
  partitions: 256 # Partitioned by PoP_id
  retention: "24h" # Retained for 24 hours then archived to S3 and ClickHouse
  producers:
    - "Edge servers (asynchronously batched cache hit, miss, and latency events)"
  consumers:
    - "Analytics telemetry pipeline (hit ratios, egress bandwidth, and latency percentiles)"

cluster_configuration:
  replication_factor: 3
  min_insync_replicas: 2
  synchronous_path: "User request -> Edge lookup -> Return HTTP 200 on hit, or fetch from origin and populate cache on miss"
  asynchronous_path: "Purge and warming commands execute via Kafka, ensuring telemetry logging never blocks edge delivery"

API Design

While CDNs primarily operate as HTTP caching proxies, the control plane exposes RESTful endpoints for cache invalidation, proactive warming, and operational telemetry.

Control Plane Type Definitions

TypeScript domain signatures defining cache purge requests, warming targets, and regional telemetry metrics:

TYPESCRIPT
// Domain types for CDN control plane and purge management
export type PurgeType = "soft_purge" | "hard_purge";
export type PurgeStatus = "propagating" | "completed" | "failed";
export type WarmStatus = "queued" | "in_progress" | "completed";

export interface PurgeRequest {
  urls?: string[];
  pattern?: string;
  surrogateKeys?: string[];
  purgeType?: PurgeType;
}

export interface PurgeResponse {
  purgeId: string;
  status: PurgeStatus;
  estimatedCompletionSeconds: number; // Operational p99 SLA estimate (exact URL < 30s, regional tag < 5s), not a deterministic guarantee
  initiatedAt: string;
}

export interface WarmRequest {
  urls: string[];
  regions: string[];
  priority?: "normal" | "high";
}

export interface WarmResponse {
  warmId: string;
  status: WarmStatus;
  targetPops: number;
}

export interface RegionalMetrics {
  region: string;
  requests: number;
  hitRatio: number;
  bandwidthGb: number;
  latencyP50Ms: number;
  latencyP99Ms: number;
}

export interface AnalyticsResponse {
  domain: string;
  period: string;
  totalRequests: number; // Example aggregate count across the queried window (e.g. 24h), distinct from the 50M RPS baseline capacity
  cacheHitRatio: number;
  bandwidthGb: number;
  latencyP50Ms: number;
  latencyP99Ms: number;
  byRegion: RegionalMetrics[];
}

export interface CdnControlPlaneClient {
  purgeCache(request: PurgeRequest): Promise<PurgeResponse>;
  warmCache(request: WarmRequest): Promise<WarmResponse>;
  getAnalytics(domain: string, period: string): Promise<AnalyticsResponse>;
}

Cache Purge Endpoint

Dispatches asynchronous purge tasks by exact URL (p99 SLA under 30 seconds globally), glob pattern, or surrogate key tag (target under 5 seconds regionally). The returned estimated_completion communicates an operational p99 SLA estimate rather than a deterministic completion guarantee, while local edge servers evict matching keys in under 1 second once the message arrives at the PoP:

HTTP
POST /api/v1/purge HTTP/1.1
Host: api.cdn.example.com
Authorization: Bearer <tenant_api_token>
Content-Type: application/json

{
  "urls": ["https://cdn.example.com/images/logo.png"],
  "pattern": "https://cdn.example.com/css/*",
  "purge_type": "soft_purge"
}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "purge_id": "purge_7a8b9c0d-1e2f-3a4b-5c6d",
  "status": "propagating",
  "estimated_completion": "30 seconds"
}

Cache Warm Endpoint

Proactively populates regional edge caches ahead of high-traffic software releases or media launches:

HTTP
POST /api/v1/warm HTTP/1.1
Host: api.cdn.example.com
Authorization: Bearer <tenant_api_token>
Content-Type: application/json

{
  "urls": ["https://cdn.example.com/videos/new-release.mp4"],
  "regions": ["us-east", "eu-west", "ap-south"]
}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "warm_id": "warm_4b5c6d7e-8f9a-0b1c-2d3e",
  "status": "queued",
  "target_pops": 85
}

Get Telemetry Analytics Endpoint

Retrieves aggregate request volumes, cache hit ratios, bandwidth consumption, and latency percentiles by region (the 50,000,000 total_requests illustrates a representative multi-hour reporting period aggregate, distinct from our 50M requests/sec instantaneous global capacity baseline):

HTTP
GET /api/v1/analytics?domain=cdn.example.com&period=24h HTTP/1.1
Host: api.cdn.example.com
Authorization: Bearer <tenant_api_token>

HTTP/1.1 200 OK
Content-Type: application/json

{
  "total_requests": 50000000,
  "cache_hit_ratio": 0.93,
  "bandwidth_gb": 15000,
  "latency_p50_ms": 12,
  "latency_p99_ms": 45,
  "by_region": [
    {
      "region": "us-east",
      "requests": 22000000,
      "hit_ratio": 0.95,
      "bandwidth_gb": 6800,
      "latency_p50_ms": 10,
      "latency_p99_ms": 38
    },
    {
      "region": "eu-west",
      "requests": 16000000,
      "hit_ratio": 0.92,
      "bandwidth_gb": 4900,
      "latency_p50_ms": 13,
      "latency_p99_ms": 46
    },
    {
      "region": "ap-south",
      "requests": 12000000,
      "hit_ratio": 0.91,
      "bandwidth_gb": 3300,
      "latency_p50_ms": 15,
      "latency_p99_ms": 52
    }
  ]
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

The CDN data model emphasizes localized in-memory cache indexing, serialized file metadata on SSD, authoritative DNS routing state, and event schemas for asynchronous control operations.

Edge Server Cache Entry

Edge nodes persist object metadata alongside cached file bodies in NVMe SSD storage or RAM:

YAML
cache_key: "https://cdn.example.com/images/logo.png"
metadata:
  content_type: "image/png"
  content_length: 45678
  etag: "\"abc123\""
  last_modified: "2026-03-13T00:00:00Z"
  cache_control: "public, max-age=86400"
  expires_at: "2026-03-14T00:00:00Z"
  created_at: "2026-03-13T00:00:00Z"
  hit_count: 1523
  last_accessed: "2026-03-13T10:30:00Z"
body: "[binary content stored on NVMe SSD or memory]"

DNS Routing Table

GeoDNS authoritative nameservers maintain dynamic routing tables mapping geographical regions to healthy PoP VIPs:

Region    PoP    IP Addresses          Health Status
US-East   NYC    [203.0.113.1, ...]    healthy
US-East   IAD    [203.0.113.5, ...]    healthy
EU-West   LDN    [198.51.100.1, ...]   healthy
AP-South  MUM    [192.0.2.1, ...]      degraded

Purge and Invalidation Event Schemas

Control plane events published to Kafka coordinate invalidation, pre-warming, and telemetry logging across all PoPs:

JSON
{
  "topic": "cache-purge",
  "event_id": "evt_7f8a9b0c-1d2e-3a4b",
  "purge_id": "purge_3c4d5e6f-7a8b-9c0d",
  "tenant_id": "tenant_enterprise_01",
  "pattern": "https://cdn.example.com/css/*",
  "initiated_at": "2026-03-13T10:00:00Z",
  "target_pops": ["all"],
  "purge_type": "soft_purge"
}

Pre-warming jobs follow a similar asynchronous event envelope:

JSON
{
  "topic": "cache-warm",
  "event_id": "evt_1a2b3c4d-5e6f-7a8b",
  "warm_id": "warm_8b9c0d1e-2f3a-4b5c",
  "tenant_id": "tenant_enterprise_01",
  "urls": ["https://cdn.example.com/videos/new-release.mp4"],
  "regions": ["us-east", "eu-west"],
  "priority": "high"
}

Fault Tolerance

A global CDN must gracefully handle origin server outages, cache stampedes during coordinated TTL expirations, regional fiber cuts, and targeted volumetric DDoS attacks without degrading user availability.

ConcernSolution
PoP failureAnycast avoids DNS TTL-based failover delays, but actual failover depends on health-check withdrawal and BGP route convergence (typically 30 to 60 seconds) to reroute traffic to the next closest healthy PoP.
Edge server failureThe load balancer within each PoP routes traffic to healthy servers, and consistent hashing rebalances cache keys.
Origin failureServe stale cached content using stale-while-revalidate and stale-if-error cache directives.
Cache stampedeRequest coalescing sends only one request to origin while all concurrent waiting requests are served from that same response.
DDoS at edgeAbsorb attacks directly at the edge using rate limiting, WAF rules, TCP SYN cookies, and challenge pages.
Cable cut (region offline)Anycast reroutes traffic globally while regional failover redirects traffic to adjacent PoPs.

Cache Stampede and Thundering Herd Mitigation

When a highly popular cached object expires, hundreds or thousands of concurrent incoming requests can trigger simultaneous upstream origin fetches, threatening to take down backend origin servers. The CDN employs three complementary defenses to protect origin infrastructure:

  • Request coalescing (singleflight): When a cache miss occurs for a key, the edge proxy or origin shield locks that key so only the first request dispatches an upstream origin fetch. All subsequent concurrent requests wait on that in-flight network socket and are satisfied by the single upstream response once it returns.
  • Stale-while-revalidate: Edge servers immediately serve the expired cached representation to users with zero added latency while asynchronously dispatching a background thread to fetch the refreshed object from origin.
  • Jittered TTL: The origin or edge proxy adds a pseudo-random jitter of plus or minus 10% to the base TTL value, ensuring that cached representations across 200+ PoPs expire at staggered intervals rather than all at once.

Additional Considerations

Advanced CDN implementations address modern transport protocols (HTTP/3 with QUIC, persistent origin connection reuse, and 103 Early Hints), ingestion strategies, multi-vendor redundancy, edge compute runtimes, TLS termination nuances, and dynamic asset optimization.

Push vs Pull CDN

In a Pull CDN, edge servers do not cache content until a user requests it. On that initial cache miss, the edge node fetches the asset from origin, stores it locally, and serves subsequent requests directly from cache. This pattern is ideal for dynamic websites, rapidly evolving user-generated content, and sites with broad, unpredictable catalogs.

In contrast, a Push CDN requires the origin server or deployment pipeline to upload assets directly to edge locations before any user requests arrive. This approach eliminates cold-miss latencies entirely, making it best suited for predictable, high-profile assets such as major software release binaries, operating system updates, and scheduled video-on-demand releases.

Multi-CDN Strategy

  • Provider diversification: Organizations combine multiple edge providers such as Akamai, CloudFront, and Fastly to prevent vendor lock-in and avoid single-vendor global outages.
  • Regional optimization and cost leverage: Traffic is routed to specific providers based on regional peering advantages, such as one vendor excelling across Latin America and another in East Asia, while enabling competitive commercial rate negotiation.
  • Real-time synthetic switching: A client-side SDK or external DNS traffic director continuously samples latency and packet loss to steer traffic dynamically to the fastest available CDN network for each user.

Edge Computing

  • Runtime execution at the edge: Serverless execution environments like Cloudflare Workers and AWS Lambda@Edge execute lightweight JavaScript or WebAssembly logic directly within PoP facilities.
  • Primary use cases: Edge runtimes evaluate JWT authentication tokens, rewrite request headers, run split A/B tests, and perform dynamic geo-targeted content personalization.
  • Origin round-trip reduction: Handling computation at edge servers eliminates origin round-trips for request verification, reducing response latency from hundreds of milliseconds to under 10 milliseconds.

TLS at Edge

  • Edge TLS termination: Cryptographic handshakes terminate at the nearest edge server, slashing connection establishment times for distant international users from multiple origin round-trips down to a single local RTT.
  • Automated certificate management: The edge control plane provisions and renews TLS certificates automatically through ACME protocols like Let's Encrypt while allowing enterprise customers to upload custom certificates.
  • Tiered upstream connections: Edge servers maintain persistent, pre-warmed HTTPS connection pools back to the origin, amortizing TLS handshake and TCP slow-start overhead across thousands of user requests.

Image Optimization at Edge

  • Automatic format negotiation: The edge proxy inspects client Accept headers to convert legacy JPEG or PNG images into modern WebP or AVIF formats on the fly.
  • Responsive resizing: Images are resized and cropped dynamically based on client viewport dimensions specified in query parameters or Client Hints headers.
  • Adaptive quality compression: Dynamic compression algorithms adjust image bitrate based on client network conditions, reducing total bandwidth consumption by 30% to 50% without visible quality loss.

Monitoring and Telemetry

  • Real-time PoP telemetry: Operations dashboards track aggregate requests per second, egress bandwidth, cache hit ratio, and HTTP status code distributions across every PoP.
  • Automated threshold alerts: Automated alerts trigger when global cache hit ratio drops below 80%, origin 5xx error rates exceed 5%, or p99 edge response latency exceeds 100 milliseconds.
  • Synthetic health probes: Edge clusters dispatch synthetic HTTP health checks to origin endpoints every 5 to 10 seconds to detect regional network routing degradations immediately.

Origin Shield Worked Capacity Calculation

Consider a representative sub-scale cohort handling 10 million requests per second (which scales linearly to our 50M requests/sec baseline). With a 95% edge cache hit ratio, 500,000 cache misses per second leave the edge tier and reach the regional origin shield (scaling to 2.5M/sec at full 50M scale). With a 90% shield hit ratio, upstream origin servers receive only 50,000 requests per second (scaling to 250K requests/sec at 50M baseline), which represents a highly manageable 0.5% origin load. Without an origin shield layer, the full miss volume (500K/sec cohort, 2.5M/sec baseline) would hit the origin directly, overwhelming a standard object storage bucket or web cluster. Consistent hashing maps identical URLs to the exact same shield node, ensuring that subsequent misses for the same URL route to that responsible node where they either hit the already cached object or join the existing in-flight singleflight request if the first fetch has not completed, collapsing concurrent in-flight misses into a single origin fetch.

Related Problems and Core Concepts

Media distribution and adaptive bitrate streaming connect directly to Video Streaming Platform. Reverse proxying, connection pooling, and SSL termination patterns align with Proxy and Reverse Proxy Patterns. To deepen your architectural foundation, review CDN and Edge Delivery, Consistent Hashing, Caching Patterns and Invalidation, Load Balancing Algorithms, and System Design Interview Patterns.

Interview Walkthrough

  • 25-minute cut

    Skip arch50 and arch75 depth unless interviewing for a staff-level role.

    • Functional and non-functional requirements with traffic profiling (3 min)
    • GeoDNS vs Anycast routing trade-offs (6 min)
    • Cache key design, Vary headers, and TTL strategies (7 min)
    • Origin shield architecture for cache miss collapse (5 min)
    • Active purge API vs TTL-only invalidation (4 min)
  • Start with the read-heavy traffic profile to establish why origin offload is the primary engineering objective, referencing System Design Interview Patterns to structure the cache hierarchy discussion.
  • Explain DNS-based GeoDNS vs Anycast routing to the nearest PoP, contrasting how DNS TTL caching delays failover during origin outages against BGP route convergence.
  • Walk through cache key design encompassing URLs, Vary headers, and query string stripping policies, including negative caching rules for HTTP 404 responses.
  • Detail cache invalidation strategies across TTL expiration, active purge APIs with pub/sub fan-out, and content-addressed immutable URLs for static assets.
  • Explain how an origin shield leverages Consistent Hashing and singleflight request coalescing to collapse thundering herds during cache miss storms.
  • Address common interview pitfalls such as caching personalized user HTML at the edge, clarifying how to use Vary headers, edge-side includes, or dynamic edge compute to prevent cross-tenant data leaks.

Engineering Trade-offs

Your interviewer will evaluate your judgment on routing trade-offs and invalidation mechanics. Contrast push vs pull caching paradigms and defend your architectural recommendations for global asset distribution.

Detailed Cache Miss Flow from Edge to Shield to Origin

A cold cache miss incurs approximately 106 milliseconds of total latency across the tiered lookup path. This includes 0 ms for established DNS and TLS, 6 ms for local edge L1 memory, L2 SSD, and L3 storage checks, 0.1 ms for consistent hash node computation, 15 ms to reach the regional origin shield, 80 ms for origin round-trip retrieval, and 5 ms for local SSD backfill. In contrast, subsequent requests result in immediate cache hits served in approximately 5 ms from RAM or 6 ms from local SSD. This creates a 20x latency difference between a hit and a miss, illustrating why maintaining a global cache hit ratio above 95% is the single most critical CDN performance metric.

TLS 1.3 Handshake at Edge (Why Edge Termination Matters)

Without a CDN, establishing a TLS handshake directly with an origin located 200 milliseconds away requires two round trips, totaling 400 ms before the client receives the first byte. Terminating TLS at an edge PoP with a 5 ms round-trip time completes the standard TLS 1.3 handshake in just 10 ms for cached assets, dropping to 5 ms when utilizing TLS 1.3 zero round-trip time (0-RTT) connection resumption. Reducing connection setup from 400 ms down to 5 ms yields an 80x acceleration in transport security negotiation. In addition, enabling Online Certificate Status Protocol (OCSP) stapling at the edge relieves clients from performing separate OCSP revocation lookups, eliminating yet another network round trip.

Consistent Hashing Within a PoP

In a PoP housing 100 edge servers, naive independent caching would cause identical URLs to be duplicated across all 100 nodes, wasting up to 100x the available storage capacity. Implementing Consistent Hashing maps each asset URL to a specific server using a virtual-node hash ring. In the baseline design, each key resolves to exactly one primary cache owner within the facility. For viral or ultra-hot objects, the cluster can optionally replicate the key across a small secondary set of nodes (such as 2 or 3 replicas) to fan out read concurrency without forfeiting storage efficiency. If an individual node fails, only approximately 1/N of the cached keys require re-caching on adjacent ring nodes rather than triggering a full PoP invalidation.

Loading...

Cache Invalidation Propagation Flow

When an administrator or automated webhook triggers an invalidation, the Purge Service validates the request and publishes an event to the cache-purge Kafka topic. Purge consumer daemons in each global PoP ingest the message from their respective partition and fan out the deletion command across all local edge servers, which purge matching cache keys from memory and disk. Global purge propagation achieves a p99 under 30 seconds for exact URLs and under 5 seconds for regional tag or pattern sweeps. Once an invalidation message arrives at a local PoP consumer daemon, local edge server eviction executes in under 1 second across memory and disk indexes. PoPs optimize pattern matching using a prefix-tree trie index over cached URLs to execute wildcard purges in O(prefix_len) time, while surrogate tags enable instantaneous O(1) tag-indexed invalidations.

Cache Warming vs On-Demand Pull

Under default on-demand pull caching, the first incoming user request triggers a cold miss that fetches content from origin. While this avoids wasting cache capacity on unrequested files, the initial user in each PoP experiences higher latency. Conversely, proactive cache warming pushes content to edge servers ahead of time, ensuring zero cold-start latency for end users at the cost of consuming valuable edge SSD space if the content fails to achieve expected viewership. Production systems favor a hybrid architecture that relies on on-demand pull caching as the general default while selectively pre-warming high-impact assets like major software updates and anticipated video releases.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...