System Design Problem

Design a URL Shortener (TinyURL / Bit.ly)

Commonly Asked By:GoogleMetaAmazonMicrosoftTwitterUber

Interview Setup

Interview Prompt

Design a URL shortening service like TinyURL or Bit.ly. Users submit long URLs and receive short links. Short links redirect to the original URL. Support optional custom aliases and basic click analytics.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What is the expected read-to-write ratio?A 100:1 read-to-write ratio is typical, which means we must design a cache-first read path alongside a write-optimized database.
Do we need per-click analytics or just redirects?This decides between HTTP 301 and 302 redirects. If real-time analytics are required, we need an asynchronous processing pipeline with Kafka, Flink, and ClickHouse.
Should the same long URL always map to the same short URL?This dictates whether we use content-based hashing for global deduplication or a Key Generation Service for collision-free key assignments.
Are custom aliases globally unique or scoped per user?This determines how we handle collision detection, such as using lightweight database transactions for global aliases or namespacing aliases per user.

Scope

In scope

  • Create, redirect, and delete URLs
  • Custom aliases with uniqueness enforcement
  • Click analytics aggregation
  • Capacity estimation and read-path optimization

Out of scope (state explicitly)

  • Link preview and Open Graph scraping
  • User authentication system design (assume API keys exist)
  • Billing and tier management

Functional Requirements

Start by clarifying the scope with your interviewer. For a URL shortener, the core flow is shortening and redirecting links. You should also confirm whether custom aliases, expiration times, and click analytics are required.

  • Given a long URL, generate a unique, short URL (for example, https://short.ly/xK9b2).
  • Given a short URL, redirect the user to the original long URL (HTTP 301 or 302).
  • Users can optionally provide a custom alias for their short URL.
  • Short URLs have a configurable expiration with a default of 5 years.
  • Analytics track total clicks, geographic distribution, referrers, and device types.
  • Users can delete their own shortened links.
  • Duplicate long URLs submitted by the same user return the existing short link.

Non-Functional Requirements

Redirect latency and system availability are the primary evaluation metrics for this problem. The heavy read-to-write ratio justifies aggressive caching and an isolated analytics pipeline.

  • High Availability (99.99%): Redirects must always remain accessible.
  • Low Latency: Redirect responses should complete in under 10 milliseconds at the 99th percentile.
  • Read-Heavy Workload: Expected read-to-write ratio is approximately 100:1.
  • Scalability: Support billions of stored URLs and peak traffic above 100,000 reads per second.
  • Durability: Once stored, a URL record must persist reliably until its expiration date.
  • Uniqueness: Two distinct destination URLs must never generate the same short code.
  • Consistency Model: Eventual consistency is acceptable for click analytics, whereas URL creation requires strong consistency.

Capacity Estimations

Calculate capacity figures before finalizing your storage architecture. The read-to-write ratio and total URL volume define the cache requirements, while the short-code length determines total namespace capacity.

MetricCalculationValue
New URLs / monthGiven100M
Redirects / month100:1 read:write10B
Writes / sec100M / (30 x 86400)~38 writes/s
Reads / sec10B / (30 x 86400)~3,800 reads/s (peak 5x: ~19K)
Record sizeshort_code + long_url + metadata~500 bytes
Storage / year100M x 12 x 500B~600 GB
5-year storage~3 TB
Cache size (20% hot)0.2 x daily_reads x 500B = 0.2 x 333M x 500B~33 GB

Short Code Length

We use Base62 encoding consisting of lowercase letters, uppercase letters, and numbers (a-z, A-Z, 0-9).

  • 6 characters yield 626 ≈ 56.8 billion unique combinations, which is sufficient for initial needs.
  • 7 characters yield 627 ≈ 3.5 trillion unique combinations, providing extensive future headroom.
  • Choose 7 characters to provide extensive long-term capacity without collision pressure.

Architecture Diagram

Interview tip: Ask about HTTP 301 vs 302 redirects early because it changes whether every click reaches your application servers.

Walk your interviewer through the system by tracing read and write traffic separately. URL creation and redirection have different latency constraints and scaling requirements. On the write path, the URL Write Service retrieves short codes from the Key Generation Service and persists records in Cassandra. On the read path, where most traffic occurs, requests pass from the CDN to Redis, reaching Cassandra only on cache misses. Click events are streamed asynchronously through Kafka, allowing the Analytics API to serve pre-aggregated metrics from ClickHouse without impacting redirect performance.

Loading...

Component Deep Dives

Walk through each component in the architecture systematically. Start at the edge and follow the read path inward, then examine the write workflow and asynchronous analytics pipeline.

API Gateway

An API gateway serves as the single entry point to centralize authentication, rate limiting, and request routing before traffic reaches backend services.

  • Purpose: Enforces rate limiting per API key using token buckets, validates JWT authentication, terminates SSL connections, and routes requests.
  • Routing: Directs /api/v1/urls to the URL Write Service and /{short_code} to the URL Read Service.
  • Fault Tolerance: Deployed as multiple stateless instances behind a Layer 4 load balancer for seamless horizontal scaling.

Key Generation Service (KGS)

Pre-generating unique keys offline eliminates runtime collision checks on the write path, keeping URL creation fast and predictable. At initial or modest traffic volume, a simpler key-generation strategy such as a database sequence or auto-increment counter may be sufficient. This batched offline KGS design is a deliberate scalability choice when independent allocation, high availability, and future scale justify the operational complexity.

  • Why KGS: Pre-generating keys avoids expensive real-time collision checks and database lookups when creating short links.
  • How It Works:
    1. An offline process generates all possible 7-character Base62 keys and stores them in a key database containing two tables: unused_keys and used_keys.
    2. Each application server requests a batch of keys (for example, 1,000 keys) from KGS.
    3. KGS atomically moves the keys from unused_keys to used_keys and assigns them to the requesting server.
    4. The application server allocates keys from its local in-memory batch, requiring zero database queries per user request.
    5. If an application server crashes, its unused in-memory keys are discarded. This waste is acceptable because the 7-character keyspace contains over 3.5 trillion values.
  • Coordination: ZooKeeper or etcd coordinates non-overlapping key ranges across active KGS instances.
  • Fault Tolerance: Multiple KGS replicas run with automated leader election. If a leader fails, a follower takes over with a new key range.

URL Write Service

The write service manages URL creation, custom alias validation, and database persistence.

  1. Receive the long destination URL and an optional custom alias.
  2. If a custom alias is requested, check the Bloom filter. If it indicates the alias is definitely absent, proceed directly to an atomic conditional insert (such as Cassandra INSERT IF NOT EXISTS) as the authoritative uniqueness check. If the filter indicates the alias might exist, verify against or attempt the conditional insert, rejecting with a 409 Conflict if already taken. On a successful insert, update the Bloom filter.
  3. If no custom alias is provided, pop the next short code from the local KGS in-memory batch.
  4. Write the record containing short_code, long_url, user_id, and expires_at to Cassandra.
  5. Populate the Redis cache with the newly created mapping.
  6. Return the generated short URL to the user.

URL Read Service (Redirect Service)

The redirect service is the performance-critical path of the system, designed to return HTTP redirects within 10 milliseconds.

  1. Receive an incoming GET /{short_code} request.
  2. Check the Redis cache. If the key exists, return the redirect immediately.
  3. If a cache miss occurs, query Cassandra for the long URL, populate the Redis cache, and return the redirect.
  4. Asynchronously publish a click event to Kafka for analytics processing.

HTTP 301 vs 302 Redirects:

  • HTTP 301 (Permanent Redirect): The browser caches the redirect URL, reducing server load on repeat visits, but preventing the server from recording individual click analytics.
  • HTTP 302 (Temporary Redirect): Every click routes through our servers, allowing comprehensive real-time click tracking.
  • Recommendation: Use 302 redirects when analytics are required, and offer 301 redirects as an option for high-volume static links.

Redis Cluster (Cache Layer)

Because read traffic heavily outweighs writes, an in-memory caching tier absorbs the majority of redirect requests ahead of Cassandra.

  • Technology Choice: Redis provides in-memory sub-millisecond lookups and native TTL support for automatic cache expiration.
  • Caching Strategy: Uses the cache-aside pattern with an LRU (Least Recently Used) eviction policy.
  • Key Structure: key=url:{short_code}, value=long_url, with a 24-hour TTL and random jitter.
  • Capacity: Approximately 33 GB of cache holds 20% of daily hot URLs comfortably on a small cluster.
  • High Availability: Redis Cluster runs with 6 nodes (3 primaries and 3 replicas) with automated failover promotion.

Cassandra (URL Store)

Cassandra acts as the primary distributed database for persistent URL mappings, offering horizontal scalability and low-latency point lookups.

  • Why Cassandra:
    • Single-key lookups using short_code as the partition key execute in O(1) time.
    • High write throughput powered by an LSM-tree storage engine.
    • Built-in row-level TTL automatically purges expired URLs during background compaction.
    • Native multi-datacenter replication without requiring third-party tooling.
    • No complex joins or multi-table transactions are necessary.
  • Replication Strategy: Replication factor of 3, with QUORUM consistency for writes and ONE for reads to deliver low read latency and high durability.
  • Compaction: Employs Leveled Compaction Strategy (LCS) to optimize for read-heavy workloads.

Kafka (Click Event Stream)

Streaming click events through Kafka decouples tracking analytics from the user redirect path, protecting latency and absorbing traffic spikes.

  • Purpose: Decouples click event ingestion from downstream stream processing and absorbs sudden traffic bursts.
  • Partitioning: The click-events topic is partitioned by short_code, ensuring that all events for a given URL land in the same partition for ordered processing.
  • Replication: Replication factor of 3 with a minimum of 2 in-sync replicas to prevent data loss.
  • Retention: Configured for 7-day retention to allow downstream stream processors sufficient time for recovery.

Analytics API Service

The Analytics API operates independently of the redirect path to handle analytical queries under a separate service level agreement.

  • Isolation: Separating analytics from redirects ensures analytical dashboard queries do not degrade redirect response times.
  • Query Path: Serves GET /api/v1/urls/{short_code}/analytics by querying pre-aggregated rollups in ClickHouse instead of scanning raw event logs.
  • Data Flow: The URL Read Service sends raw click events to Kafka, Apache Flink computes windowed aggregates into ClickHouse, and this service reads the aggregated views.

Apache Flink (Stream Processing)

Apache Flink aggregates raw click events in near real time across time windows, geographies, and device categories before storing them in ClickHouse.

  • Processing Model: Computes sliding and tumbling window aggregations (such as clicks per minute, hour, and day) directly from the Kafka stream.
  • Aggregation Dimensions: Groups events by short code, time window, geographic region, and device type before writing batches to ClickHouse.
  • Fault Tolerance: Leverages distributed checkpointing to provide consistent state recovery, and when paired with replayable Kafka sources and idempotent ClickHouse sinks, supports end-to-end exactly-once processing semantics for windowed aggregates.

ClickHouse (Analytics Store)

ClickHouse is a columnar database optimized for fast analytical aggregations over large datasets, powering user dashboards and metrics reports.

  • Why ClickHouse: Columnar storage and vectorized execution make it fast for aggregate queries such as sums, counts, and group-by filters over billions of rows.
  • Use Cases: Powers queries such as top URLs by clicks today, breakdown by country, and referrer trends.

API Design

Define clear REST endpoints for URL creation, redirection, deletion, and analytics retrieval before detailing the underlying data models.

Create Short URL

HTTP
POST /api/v1/urls
Authorization: Bearer <token>
Content-Type: application/json

{
  "long_url": "https://example.com/some/very/long/path?query=param",
  "custom_alias": "my-brand",
  "expires_at": "2031-03-13T00:00:00Z"
}

Response: 201 Created
{
  "short_url": "https://short.ly/xK9b2",
  "long_url": "https://example.com/some/very/long/path?query=param",
  "short_code": "xK9b2",
  "expires_at": "2031-03-13T00:00:00Z",
  "created_at": "2026-03-13T10:00:00Z"
}

Redirect

HTTP
GET /{short_code}

Response: 302 Found
Location: https://example.com/some/very/long/path?query=param

Delete URL

HTTP
DELETE /api/v1/urls/{short_code}
Authorization: Bearer <token>

Response: 204 No Content

Get Analytics

HTTP
GET /api/v1/urls/{short_code}/analytics?period=7d
Authorization: Bearer <token>

Response: 200 OK
{
  "short_code": "xK9b2",
  "total_clicks": 152437,
  "clicks_by_day": [
    {"date": "2026-03-12", "count": 1234}
  ],
  "top_countries": [
    {"country": "US", "count": 50234}
  ],
  "top_referrers": [
    {"referrer": "twitter.com", "count": 30211}
  ]
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

The data model reflects key access patterns: constant-time lookups by short code for redirects, indexed queries by user ID for dashboards, and columnar rollups for click analytics.

Cassandra: URL Table

Primary store for URL mappings. Partition key is short_code for O(1) lookups. Built-in TTL auto-deletes expired rows.

SQL
CREATE TABLE url_mappings (
    short_code  TEXT,          -- Partition key
    long_url    TEXT,
    user_id     UUID,
    created_at  TIMESTAMP,
    expires_at  TIMESTAMP,
    is_custom   BOOLEAN,
    PRIMARY KEY (short_code)
) WITH default_time_to_live = 157680000;  -- 5 years in seconds

Cassandra: User URLs Table

Enables listing all URLs created by a user for dashboard views, clustered by creation time in descending order.

SQL
CREATE TABLE user_urls (
    user_id     UUID,
    created_at  TIMESTAMP,
    short_code  TEXT,
    long_url    TEXT,
    PRIMARY KEY (user_id, created_at)
) WITH CLUSTERING ORDER BY (created_at DESC);

Cassandra: Click Events Table

Note: In the recommended production architecture, raw click events stream through Kafka and Apache Flink directly into ClickHouse for analytics queries. The schema below represents a direct Cassandra storage alternative.

SQL
CREATE TABLE click_events (
    short_code  TEXT,
    timestamp   TIMESTAMP,
    user_agent  TEXT,
    referer     TEXT,
    country     TEXT,
    device      TEXT,
    PRIMARY KEY (short_code, timestamp)
) WITH CLUSTERING ORDER BY (timestamp DESC);

Redis Cache Structure

Key:    url:xK9b2
Value:  "https://example.com/some/very/long/path?query=param"
TTL:    86400 (24 hours)

Kafka Topic: click-events

JSON
{
  "event_id": "uuid-v4",
  "short_code": "xK9b2",
  "timestamp": "2026-03-13T10:30:00Z",
  "ip": "203.0.113.42",
  "user_agent": "Mozilla/5.0...",
  "referrer": "https://twitter.com/post/123",
  "country": "US",
  "city": "San Francisco",
  "device_type": "mobile",
  "os": "iOS"
}

ClickHouse: Analytics Table

Columnar analytical table for aggregated click metrics, using SummingMergeTree for efficient count rollups.

SQL
CREATE TABLE url_analytics (
    short_code  String,
    event_date  Date,
    hour        UInt8,
    country     LowCardinality(String),
    referrer    String,
    device_type LowCardinality(String),
    click_count UInt64
) ENGINE = SummingMergeTree()
ORDER BY (short_code, event_date, hour, country);

Key Generation Service: Key Store (PostgreSQL)

Tracks pre-generated short code keys, their assignment status, and server allocations.

SQL
CREATE TABLE keys (
    key_value   CHAR(7) PRIMARY KEY,
    status      ENUM('unused', 'assigned', 'used'),
    assigned_to VARCHAR(64),  -- server instance ID
    assigned_at TIMESTAMP
);

Fault Tolerance

Address critical failure scenarios to demonstrate how the system maintains high availability during infrastructure outages and traffic surges.

General Resilience Techniques

TechniqueApplication
ReplicationCassandra with RF=3, Kafka with RF=3, and Redis Cluster with 3 primaries and 3 replicas
Health ChecksLoad balancers continuously monitor service instance health and deregister unhealthy nodes
Circuit BreakersApplied between services (such as Write Service to KGS) to prevent cascading failures
Retry with BackoffExponential backoff with jitter applied to transient database and network timeouts
IdempotencySubmitting the same long URL and user token returns the existing short link without creating duplicates
Graceful DegradationIf the analytics pipeline becomes unavailable, user redirects continue operating without interruption

Problem-Specific Scenarios

1. KGS Failure and Key Pool Depletion

  • Deploy multiple KGS instances with pre-assigned, non-overlapping key ranges.
  • Each application server pre-fetches a batch of 1,000 keys, allowing it to continue serving writes during brief KGS outages.
  • If an application server crashes, its in-memory keys are discarded safely because the keyspace is vast.
  • Automated alerts trigger when the unused key pool drops below 10% capacity.

2. Cache Stampede (Thundering Herd)

  • When a hot URL cache entry expires, thousands of concurrent requests can overwhelm the database.
  • Solution: Apply request coalescing using the singleflight pattern so only a single thread queries Cassandra while concurrent requests wait for the result.
  • Add random jitter to cache TTL values (such as 24 hours ± 2 hours) to avoid synchronized cache expirations across keys.

3. Viral URL Traffic Spikes

  • A viral link can generate millions of redirects per second and overload a single Redis instance.
  • Solution: Add an in-process L1 cache (such as Caffeine or Guava) with a 60-second TTL on each application server, and replicate the hot key across Redis read replicas.

4. Custom Alias Concurrency Conflicts

  • Two users might simultaneously attempt to register the exact same custom alias.
  • Solution: Use Cassandra Lightweight Transactions (LWT) with INSERT IF NOT EXISTS on the custom alias partition key so that exactly one write succeeds and the second returns a 409 Conflict.

5. Datacenter Outages

  • Cassandra multi-datacenter replication ensures read requests can be served from surviving regions.
  • GeoDNS automatically routes traffic away from the degraded datacenter to healthy regions.

Additional Considerations

These additional production considerations demonstrate operational maturity and practical system awareness.

URL Normalization

Normalizing destination URLs before persistence avoids allocating multiple short keys to identical target addresses:

  • Lowercase the scheme and hostname (for example, convert HTTP://Example.COM to http://example.com).
  • Remove default protocol ports (:80 for HTTP and :443 for HTTPS).
  • Sort query parameters alphabetically.
  • Remove trailing slashes where appropriate.
  • Decode unnecessary percent-encoded characters.

Abuse Prevention

  • Rate Limiting: Implement token bucket rate limiting per API key (such as 100 URL creations per hour on the free tier).
  • Malicious URL Filtering: Integrate with the Google Safe Browsing API to detect and block malicious or phishing URLs.
  • CAPTCHA Verification: Require CAPTCHA challenges for anonymous or unauthenticated URL creation.
  • Spam Detection: Validate destination domains against known spam and threat intelligence databases.

Bloom Filter for Duplicate Detection

  • Use a distributed Bloom filter to optimize custom alias creation by quickly identifying definitely absent aliases before touching the database.
  • Because a Bloom filter has no false negatives, a negative result guarantees the alias is currently absent from the filter, allowing the system to skip preliminary existence queries and proceed directly to the conditional insert.
  • A positive result indicates the alias might already exist due to potential false positives, prompting the system to verify against or conditionally write to the authoritative store.
  • The authoritative datastore (such as Cassandra with Lightweight Transactions) always enforces the final uniqueness decision. On successful creation, the new alias is added to the Bloom filter.
  • With an illustrative 1% false-positive rate, approximately 1% of truly absent aliases may be reported as possibly present and take the slower check path, while 99% of absent aliases correctly test negative and avoid preliminary database existence reads. A filter size of ~1 GB comfortably supports 1 billion entries at this false-positive rate.

Monitoring and Alerting

  • Redirect Latency: Track p50, p95, and p99 latency metrics against the 10-millisecond target.
  • Cache Hit Ratio: Alert when the cache hit ratio falls below 80%.
  • KGS Key Pool Health: Monitor remaining unused keys and alert when capacity drops below threshold levels.
  • Kafka Consumer Lag: Track consumer lag on the click-events topic to ensure the analytics pipeline remains healthy.
  • Error Rates: Monitor 4xx and 5xx response codes across all API endpoints.

GDPR and Data Privacy

  • Allow users to permanently delete shortened links along with all associated analytics data.
  • Anonymize IP addresses in analytics storage after 30 days.
  • Provide data export capabilities for user records and analytics history.

Related Problems and Concepts

Short-link key generation and read-heavy caching strategies share patterns with Pastebin, which uses a similar database and object storage split at lower scale. Rate limiting on the URL creation endpoint applies concepts covered in API Rate Limiter. You can also review foundational caching strategies in Caching Patterns and Invalidation and back-of-the-envelope estimation guidelines in Back-of-the-Envelope Estimation.

Interview Walkthrough

  • 25-Minute Interview Strategy

    Focus on the core read-path caching and key generation decisions unless the interviewer explicitly requests multi-datacenter depth.

    • Functional and Non-Functional Requirements (3 min)
    • Capacity Calculations (5 min)
    • Read Path and Multi-Layer Cache Architecture (8 min)
    • Key Generation Service vs Hashing (5 min)
    • Asynchronous Analytics Pipeline (4 min)
  • Highlight the 100:1 read-to-write ratio early. Redirect performance is the core bottleneck, so design the read path before detailing URL creation.
  • Discuss HTTP 301 vs 302 redirect trade-offs upfront, noting that 302 preserves click analytics while 301 reduces server traffic via browser and CDN caching.
  • Compare KGS against URL hashing. Hashing offers automatic deduplication but requires collision retry handling, whereas KGS pre-allocates keys so write operations avoid runtime collision overhead.
  • Detail the four caching layers (CDN, local memory, Redis cluster, and Cassandra) and connect the 1.9-millisecond average latency back to the 10-millisecond p99 SLO.
  • Explain how a Bloom filter pre-check quickly identifies definitely absent custom aliases so the system can skip preliminary existence reads before Cassandra executes an authoritative conditional insert.
  • Stream click analytics through Kafka to ensure background processing does not block the user redirect response.
  • Use capacity calculations for 7-character Base62 encoding (approximately 3.5 trillion keys) to demonstrate that the design scales for decades.

Engineering Trade-offs

Interviewers evaluate how you navigate architectural forks and justify trade-offs. The sections below analyze the primary design decisions for this system.

Key Generation Strategy: KGS vs Hashing

Selecting the method for producing short codes impacts collision handling, write latency, and storage deduplication.

ApproachImplementationProsCons
MD5 or SHA-256 Hash with Base62Hash the destination URL and take the first 7 charactersDeterministic and provides automatic global deduplication of identical URLsCollision risks from the birthday paradox require runtime collision detection and retry loops
Centralized CounterIncrement a centralized database counter and Base62 encode the valueSimple to implement and produces unique sequential valuesSingle point of failure, predictable URL sequences, and a single-writer bottleneck at scale
Key Generation Service (KGS) ⭐Pre-generate keys offline in batches and dispense to application serversZero runtime collisions, non-sequential codes, O(1) allocation time, and no single point of failure with multiple instancesUnused in-memory keys are lost if a server restarts, and requires dedicated KGS infrastructure
UUID with Base62Generate a UUID v4 and encode to Base62Decentralized and simple generation128-bit UUID produces 22 Base62 characters, which is too long for a short URL
Snowflake ID with Base62Generate a 64-bit time-ordered unique ID and Base62 encodeUnique, decentralized, and sortable by time64-bit ID requires 11 Base62 characters and exposes the creation timestamp

Why KGS is chosen for high-scale URL shorteners: The primary requirement is a compact 7-character short URL. KGS delivers exactly 7 characters with zero collision checks at runtime. At initial or modest scale, a simpler approach such as a database auto-increment counter or single-node key sequence may be sufficient. The distributed KGS architecture with offline pre-generation and leader election is a deliberate scalability choice that provides collision-free, non-sequential allocation and high write availability as traffic scales.

When to choose hashing instead: Choose hashing when global content deduplication is a mandatory requirement where every submission of the exact same long URL must map to the same short code without maintaining a secondary lookup index.

Database Selection: Cassandra vs Alternatives

The database tier must support high write throughput, massive dataset growth, and low-latency key-value lookups.

DatabaseStrengths for This Use CaseWeaknesses
Cassandra ⭐Masterless architecture with no single point of failure, linear horizontal scaling, tunable consistency, built-in row TTL for automatic expiration, and native multi-datacenter replicationDoes not support ACID transactions or multi-table joins, and requires operational expertise
DynamoDBFully managed by AWS, automatic scaling, single-digit millisecond latency, and built-in TTL supportVendor lock-in, higher cost profile at extreme traffic volumes, and throughput provisioning limits
MySQLFamiliar ecosystem, strong ACID guarantees, and mature toolingSingle-writer primary creates write bottlenecks at scale, sharding requires manual maintenance, and lacks built-in TTL support
PostgreSQLRich feature set, robust JSON support, and strong relational integrityShares the same horizontal scaling constraints as MySQL for massive write-heavy key-value workloads

Deciding Factors:

  1. The primary access pattern is pure key-value point lookup: no multi-table joins or relational constraints are required.
  2. Cassandra uses an LSM-tree storage engine that handles write bursts with consistent low latency.
  3. Cassandra provides native row-level TTL, automatically purging expired records during background compaction without application cleanup jobs.
  4. Native multi-datacenter replication allows read requests to be served locally from the nearest region.
  5. DynamoDB is a viable alternative if the infrastructure is AWS-native and managed operations are preferred, though Cassandra remains cost-effective at multi-billion record scale.

HTTP Redirect Status Codes: 301 vs 302

The choice of HTTP status code determines how browsers and intermediate caches handle redirects.

Status CodeBehaviorProsCons
301 (Permanent Redirect)Browsers and CDNs cache the target URL indefinitely or according to cache headersMinimizes server load and provides the fastest redirect on repeat visitsPrevents per-click tracking on cached visits, and changing destination URLs requires cache invalidation
302 (Temporary Redirect) ⭐Browsers do not cache the redirect and query the shortener on every clickEnables accurate real-time click tracking, geolocation capture, and dynamic destination updatesIncurs a network hop to the redirect service on every user click
307 (Temporary Redirect, Preserves Method)Similar to 302, but specifies that the HTTP request method remains unchanged according to RFC specificationsPreserves POST or PUT methods across redirectsOnly relevant for non-GET requests, which are rarely used in public URL shortening

Recommendation: Default to HTTP 302 redirects because analytics tracking and link flexibility are critical for most URL shortening platforms. Provide 301 redirects as an option for enterprise high-volume marketing campaigns where redirect speed takes precedence.

Read Path Optimization and Latency Budget

At peak loads exceeding 19,000 reads per second, multi-layer caching ensures the system meets its sub-10-millisecond latency target.

  • Layer 1: CDN Edge (CloudFront or Fastly): For this illustrative design, assume that for 301 redirects the CDN caches the Location header at edge locations close to users, achieving roughly a 30% hit ratio with ~5ms response times.
  • Layer 2: In-Process Application Cache (Caffeine): Each application server caches the top 10,000 most active URLs in memory with a 60-second TTL, absorbing an assumed 20% to 30% of remaining requests with ~0.01ms access times.
  • Layer 3: Distributed Redis Cluster: Caches active URLs with a 24-hour TTL and jitter, handling roughly 90% of requests that reach the backend with ~0.5ms network latency.
  • Layer 4: Cassandra Primary Storage: Serves cold URLs that miss the cache layers (approximately 5% of total traffic) with ~2ms to 5ms disk reads under LOCAL_ONE consistency.
LayerTechnologyHit RatioLatencyTTL
L1CDN Edge (CloudFront or Fastly)~30%~5msCDN cache headers
L2In-Process Cache (Caffeine)~20-30%~0.01ms60 seconds
L3Redis Cluster~90% of remaining~0.5ms24 hours with jitter
L4Cassandra~5% (cold URLs)~2-5ms5 years (persistent)

Effective Latency Calculation: Under these illustrative cache-hit and component latency assumptions: 0.30 x 5ms + 0.25 x 0.01ms + 0.40 x 0.5ms + 0.05 x 3ms = ~1.85 milliseconds weighted average latency, comfortably within the 10-millisecond p99 latency target.

Encoding Schemes: Base62 vs Base58 and Base64

The character set chosen for short codes affects URL readability, length, and transmission safety.

EncodingCharacter SetURL SafeCharacter Ambiguity
Base62a-z, A-Z, 0-9YesSimilar glyphs such as 0/O and 1/l can cause confusion if manually transcribed
Base58Base62 minus 0, O, l, and IYesEliminates visually ambiguous characters, ideal for manual entry
Base64Base62 plus + and /No (requires percent-encoding for URLs)Contains non-alphanumeric characters that need escaping

Recommendation: Use Base62 for standard web shorteners where links are clicked digitally, as it maximizes combinations per character. Use Base58 if short codes are printed or frequently entered by hand.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...