Interview Setup
Interview Prompt
Design a URL shortening service like TinyURL or Bit.ly. Users submit long URLs and receive short links. Short links redirect to the original URL. Support optional custom aliases and basic click analytics.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What is the expected read-to-write ratio? | A 100:1 read-to-write ratio is typical, which means we must design a cache-first read path alongside a write-optimized database. |
| Do we need per-click analytics or just redirects? | This decides between HTTP 301 and 302 redirects. If real-time analytics are required, we need an asynchronous processing pipeline with Kafka, Flink, and ClickHouse. |
| Should the same long URL always map to the same short URL? | This dictates whether we use content-based hashing for global deduplication or a Key Generation Service for collision-free key assignments. |
| Are custom aliases globally unique or scoped per user? | This determines how we handle collision detection, such as using lightweight database transactions for global aliases or namespacing aliases per user. |
Scope
In scope
- Create, redirect, and delete URLs
- Custom aliases with uniqueness enforcement
- Click analytics aggregation
- Capacity estimation and read-path optimization
Out of scope (state explicitly)
- Link preview and Open Graph scraping
- User authentication system design (assume API keys exist)
- Billing and tier management
Functional Requirements
Start by clarifying the scope with your interviewer. For a URL shortener, the core flow is shortening and redirecting links. You should also confirm whether custom aliases, expiration times, and click analytics are required.
- Given a long URL, generate a unique, short URL (for example,
https://short.ly/xK9b2). - Given a short URL, redirect the user to the original long URL (HTTP 301 or 302).
- Users can optionally provide a custom alias for their short URL.
- Short URLs have a configurable expiration with a default of 5 years.
- Analytics track total clicks, geographic distribution, referrers, and device types.
- Users can delete their own shortened links.
- Duplicate long URLs submitted by the same user return the existing short link.
Non-Functional Requirements
Redirect latency and system availability are the primary evaluation metrics for this problem. The heavy read-to-write ratio justifies aggressive caching and an isolated analytics pipeline.
- High Availability (99.99%): Redirects must always remain accessible.
- Low Latency: Redirect responses should complete in under 10 milliseconds at the 99th percentile.
- Read-Heavy Workload: Expected read-to-write ratio is approximately 100:1.
- Scalability: Support billions of stored URLs and peak traffic above 100,000 reads per second.
- Durability: Once stored, a URL record must persist reliably until its expiration date.
- Uniqueness: Two distinct destination URLs must never generate the same short code.
- Consistency Model: Eventual consistency is acceptable for click analytics, whereas URL creation requires strong consistency.
Capacity Estimations
Calculate capacity figures before finalizing your storage architecture. The read-to-write ratio and total URL volume define the cache requirements, while the short-code length determines total namespace capacity.
| Metric | Calculation | Value |
|---|---|---|
| New URLs / month | Given | 100M |
| Redirects / month | 100:1 read:write | 10B |
| Writes / sec | 100M / (30 x 86400) | ~38 writes/s |
| Reads / sec | 10B / (30 x 86400) | ~3,800 reads/s (peak 5x: ~19K) |
| Record size | short_code + long_url + metadata | ~500 bytes |
| Storage / year | 100M x 12 x 500B | ~600 GB |
| 5-year storage | ~3 TB | |
| Cache size (20% hot) | 0.2 x daily_reads x 500B = 0.2 x 333M x 500B | ~33 GB |
Short Code Length
We use Base62 encoding consisting of lowercase letters, uppercase letters, and numbers (a-z, A-Z, 0-9).
- 6 characters yield 626 ≈ 56.8 billion unique combinations, which is sufficient for initial needs.
- 7 characters yield 627 ≈ 3.5 trillion unique combinations, providing extensive future headroom.
- Choose 7 characters to provide extensive long-term capacity without collision pressure.
Architecture Diagram
Interview tip: Ask about HTTP 301 vs 302 redirects early because it changes whether every click reaches your application servers.
Walk your interviewer through the system by tracing read and write traffic separately. URL creation and redirection have different latency constraints and scaling requirements. On the write path, the URL Write Service retrieves short codes from the Key Generation Service and persists records in Cassandra. On the read path, where most traffic occurs, requests pass from the CDN to Redis, reaching Cassandra only on cache misses. Click events are streamed asynchronously through Kafka, allowing the Analytics API to serve pre-aggregated metrics from ClickHouse without impacting redirect performance.
Component Deep Dives
Walk through each component in the architecture systematically. Start at the edge and follow the read path inward, then examine the write workflow and asynchronous analytics pipeline.
API Gateway
An API gateway serves as the single entry point to centralize authentication, rate limiting, and request routing before traffic reaches backend services.
- Purpose: Enforces rate limiting per API key using token buckets, validates JWT authentication, terminates SSL connections, and routes requests.
- Routing: Directs
/api/v1/urlsto the URL Write Service and/{short_code}to the URL Read Service. - Fault Tolerance: Deployed as multiple stateless instances behind a Layer 4 load balancer for seamless horizontal scaling.
Key Generation Service (KGS)
Pre-generating unique keys offline eliminates runtime collision checks on the write path, keeping URL creation fast and predictable. At initial or modest traffic volume, a simpler key-generation strategy such as a database sequence or auto-increment counter may be sufficient. This batched offline KGS design is a deliberate scalability choice when independent allocation, high availability, and future scale justify the operational complexity.
- Why KGS: Pre-generating keys avoids expensive real-time collision checks and database lookups when creating short links.
- How It Works:
- An offline process generates all possible 7-character Base62 keys and stores them in a key database containing two tables:
unused_keysandused_keys. - Each application server requests a batch of keys (for example, 1,000 keys) from KGS.
- KGS atomically moves the keys from
unused_keystoused_keysand assigns them to the requesting server. - The application server allocates keys from its local in-memory batch, requiring zero database queries per user request.
- If an application server crashes, its unused in-memory keys are discarded. This waste is acceptable because the 7-character keyspace contains over 3.5 trillion values.
- An offline process generates all possible 7-character Base62 keys and stores them in a key database containing two tables:
- Coordination: ZooKeeper or etcd coordinates non-overlapping key ranges across active KGS instances.
- Fault Tolerance: Multiple KGS replicas run with automated leader election. If a leader fails, a follower takes over with a new key range.
URL Write Service
The write service manages URL creation, custom alias validation, and database persistence.
- Receive the long destination URL and an optional custom alias.
- If a custom alias is requested, check the Bloom filter. If it indicates the alias is definitely absent, proceed directly to an atomic conditional insert (such as Cassandra INSERT IF NOT EXISTS) as the authoritative uniqueness check. If the filter indicates the alias might exist, verify against or attempt the conditional insert, rejecting with a 409 Conflict if already taken. On a successful insert, update the Bloom filter.
- If no custom alias is provided, pop the next short code from the local KGS in-memory batch.
- Write the record containing
short_code,long_url,user_id, andexpires_atto Cassandra. - Populate the Redis cache with the newly created mapping.
- Return the generated short URL to the user.
URL Read Service (Redirect Service)
The redirect service is the performance-critical path of the system, designed to return HTTP redirects within 10 milliseconds.
- Receive an incoming
GET /{short_code}request. - Check the Redis cache. If the key exists, return the redirect immediately.
- If a cache miss occurs, query Cassandra for the long URL, populate the Redis cache, and return the redirect.
- Asynchronously publish a click event to Kafka for analytics processing.
HTTP 301 vs 302 Redirects:
- HTTP 301 (Permanent Redirect): The browser caches the redirect URL, reducing server load on repeat visits, but preventing the server from recording individual click analytics.
- HTTP 302 (Temporary Redirect): Every click routes through our servers, allowing comprehensive real-time click tracking.
- Recommendation: Use 302 redirects when analytics are required, and offer 301 redirects as an option for high-volume static links.
Redis Cluster (Cache Layer)
Because read traffic heavily outweighs writes, an in-memory caching tier absorbs the majority of redirect requests ahead of Cassandra.
- Technology Choice: Redis provides in-memory sub-millisecond lookups and native TTL support for automatic cache expiration.
- Caching Strategy: Uses the cache-aside pattern with an LRU (Least Recently Used) eviction policy.
- Key Structure:
key=url:{short_code},value=long_url, with a 24-hour TTL and random jitter. - Capacity: Approximately 33 GB of cache holds 20% of daily hot URLs comfortably on a small cluster.
- High Availability: Redis Cluster runs with 6 nodes (3 primaries and 3 replicas) with automated failover promotion.
Cassandra (URL Store)
Cassandra acts as the primary distributed database for persistent URL mappings, offering horizontal scalability and low-latency point lookups.
- Why Cassandra:
- Single-key lookups using
short_codeas the partition key execute in O(1) time. - High write throughput powered by an LSM-tree storage engine.
- Built-in row-level TTL automatically purges expired URLs during background compaction.
- Native multi-datacenter replication without requiring third-party tooling.
- No complex joins or multi-table transactions are necessary.
- Single-key lookups using
- Replication Strategy: Replication factor of 3, with QUORUM consistency for writes and ONE for reads to deliver low read latency and high durability.
- Compaction: Employs Leveled Compaction Strategy (LCS) to optimize for read-heavy workloads.
Kafka (Click Event Stream)
Streaming click events through Kafka decouples tracking analytics from the user redirect path, protecting latency and absorbing traffic spikes.
- Purpose: Decouples click event ingestion from downstream stream processing and absorbs sudden traffic bursts.
- Partitioning: The
click-eventstopic is partitioned byshort_code, ensuring that all events for a given URL land in the same partition for ordered processing. - Replication: Replication factor of 3 with a minimum of 2 in-sync replicas to prevent data loss.
- Retention: Configured for 7-day retention to allow downstream stream processors sufficient time for recovery.
Analytics API Service
The Analytics API operates independently of the redirect path to handle analytical queries under a separate service level agreement.
- Isolation: Separating analytics from redirects ensures analytical dashboard queries do not degrade redirect response times.
- Query Path: Serves
GET /api/v1/urls/{short_code}/analyticsby querying pre-aggregated rollups in ClickHouse instead of scanning raw event logs. - Data Flow: The URL Read Service sends raw click events to Kafka, Apache Flink computes windowed aggregates into ClickHouse, and this service reads the aggregated views.
Apache Flink (Stream Processing)
Apache Flink aggregates raw click events in near real time across time windows, geographies, and device categories before storing them in ClickHouse.
- Processing Model: Computes sliding and tumbling window aggregations (such as clicks per minute, hour, and day) directly from the Kafka stream.
- Aggregation Dimensions: Groups events by short code, time window, geographic region, and device type before writing batches to ClickHouse.
- Fault Tolerance: Leverages distributed checkpointing to provide consistent state recovery, and when paired with replayable Kafka sources and idempotent ClickHouse sinks, supports end-to-end exactly-once processing semantics for windowed aggregates.
ClickHouse (Analytics Store)
ClickHouse is a columnar database optimized for fast analytical aggregations over large datasets, powering user dashboards and metrics reports.
- Why ClickHouse: Columnar storage and vectorized execution make it fast for aggregate queries such as sums, counts, and group-by filters over billions of rows.
- Use Cases: Powers queries such as top URLs by clicks today, breakdown by country, and referrer trends.
API Design
Define clear REST endpoints for URL creation, redirection, deletion, and analytics retrieval before detailing the underlying data models.
Create Short URL
POST /api/v1/urls
Authorization: Bearer <token>
Content-Type: application/json
{
"long_url": "https://example.com/some/very/long/path?query=param",
"custom_alias": "my-brand",
"expires_at": "2031-03-13T00:00:00Z"
}
Response: 201 Created
{
"short_url": "https://short.ly/xK9b2",
"long_url": "https://example.com/some/very/long/path?query=param",
"short_code": "xK9b2",
"expires_at": "2031-03-13T00:00:00Z",
"created_at": "2026-03-13T10:00:00Z"
}Redirect
GET /{short_code}
Response: 302 Found
Location: https://example.com/some/very/long/path?query=paramDelete URL
DELETE /api/v1/urls/{short_code}
Authorization: Bearer <token>
Response: 204 No ContentGet Analytics
GET /api/v1/urls/{short_code}/analytics?period=7d
Authorization: Bearer <token>
Response: 200 OK
{
"short_code": "xK9b2",
"total_clicks": 152437,
"clicks_by_day": [
{"date": "2026-03-12", "count": 1234}
],
"top_countries": [
{"country": "US", "count": 50234}
],
"top_referrers": [
{"referrer": "twitter.com", "count": 30211}
]
}Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
The data model reflects key access patterns: constant-time lookups by short code for redirects, indexed queries by user ID for dashboards, and columnar rollups for click analytics.
Cassandra: URL Table
Primary store for URL mappings. Partition key is short_code for O(1) lookups. Built-in TTL auto-deletes expired rows.
CREATE TABLE url_mappings (
short_code TEXT, -- Partition key
long_url TEXT,
user_id UUID,
created_at TIMESTAMP,
expires_at TIMESTAMP,
is_custom BOOLEAN,
PRIMARY KEY (short_code)
) WITH default_time_to_live = 157680000; -- 5 years in secondsCassandra: User URLs Table
Enables listing all URLs created by a user for dashboard views, clustered by creation time in descending order.
CREATE TABLE user_urls (
user_id UUID,
created_at TIMESTAMP,
short_code TEXT,
long_url TEXT,
PRIMARY KEY (user_id, created_at)
) WITH CLUSTERING ORDER BY (created_at DESC);Cassandra: Click Events Table
Note: In the recommended production architecture, raw click events stream through Kafka and Apache Flink directly into ClickHouse for analytics queries. The schema below represents a direct Cassandra storage alternative.
CREATE TABLE click_events (
short_code TEXT,
timestamp TIMESTAMP,
user_agent TEXT,
referer TEXT,
country TEXT,
device TEXT,
PRIMARY KEY (short_code, timestamp)
) WITH CLUSTERING ORDER BY (timestamp DESC);Redis Cache Structure
Key: url:xK9b2 Value: "https://example.com/some/very/long/path?query=param" TTL: 86400 (24 hours)
Kafka Topic: click-events
{
"event_id": "uuid-v4",
"short_code": "xK9b2",
"timestamp": "2026-03-13T10:30:00Z",
"ip": "203.0.113.42",
"user_agent": "Mozilla/5.0...",
"referrer": "https://twitter.com/post/123",
"country": "US",
"city": "San Francisco",
"device_type": "mobile",
"os": "iOS"
}ClickHouse: Analytics Table
Columnar analytical table for aggregated click metrics, using SummingMergeTree for efficient count rollups.
CREATE TABLE url_analytics (
short_code String,
event_date Date,
hour UInt8,
country LowCardinality(String),
referrer String,
device_type LowCardinality(String),
click_count UInt64
) ENGINE = SummingMergeTree()
ORDER BY (short_code, event_date, hour, country);Key Generation Service: Key Store (PostgreSQL)
Tracks pre-generated short code keys, their assignment status, and server allocations.
CREATE TABLE keys (
key_value CHAR(7) PRIMARY KEY,
status ENUM('unused', 'assigned', 'used'),
assigned_to VARCHAR(64), -- server instance ID
assigned_at TIMESTAMP
);Fault Tolerance
Address critical failure scenarios to demonstrate how the system maintains high availability during infrastructure outages and traffic surges.
General Resilience Techniques
| Technique | Application |
|---|---|
| Replication | Cassandra with RF=3, Kafka with RF=3, and Redis Cluster with 3 primaries and 3 replicas |
| Health Checks | Load balancers continuously monitor service instance health and deregister unhealthy nodes |
| Circuit Breakers | Applied between services (such as Write Service to KGS) to prevent cascading failures |
| Retry with Backoff | Exponential backoff with jitter applied to transient database and network timeouts |
| Idempotency | Submitting the same long URL and user token returns the existing short link without creating duplicates |
| Graceful Degradation | If the analytics pipeline becomes unavailable, user redirects continue operating without interruption |
Problem-Specific Scenarios
1. KGS Failure and Key Pool Depletion
- Deploy multiple KGS instances with pre-assigned, non-overlapping key ranges.
- Each application server pre-fetches a batch of 1,000 keys, allowing it to continue serving writes during brief KGS outages.
- If an application server crashes, its in-memory keys are discarded safely because the keyspace is vast.
- Automated alerts trigger when the unused key pool drops below 10% capacity.
2. Cache Stampede (Thundering Herd)
- When a hot URL cache entry expires, thousands of concurrent requests can overwhelm the database.
- Solution: Apply request coalescing using the singleflight pattern so only a single thread queries Cassandra while concurrent requests wait for the result.
- Add random jitter to cache TTL values (such as 24 hours ± 2 hours) to avoid synchronized cache expirations across keys.
3. Viral URL Traffic Spikes
- A viral link can generate millions of redirects per second and overload a single Redis instance.
- Solution: Add an in-process L1 cache (such as Caffeine or Guava) with a 60-second TTL on each application server, and replicate the hot key across Redis read replicas.
4. Custom Alias Concurrency Conflicts
- Two users might simultaneously attempt to register the exact same custom alias.
- Solution: Use Cassandra Lightweight Transactions (LWT) with
INSERT IF NOT EXISTSon the custom alias partition key so that exactly one write succeeds and the second returns a 409 Conflict.
5. Datacenter Outages
- Cassandra multi-datacenter replication ensures read requests can be served from surviving regions.
- GeoDNS automatically routes traffic away from the degraded datacenter to healthy regions.
Additional Considerations
These additional production considerations demonstrate operational maturity and practical system awareness.
URL Normalization
Normalizing destination URLs before persistence avoids allocating multiple short keys to identical target addresses:
- Lowercase the scheme and hostname (for example, convert
HTTP://Example.COMtohttp://example.com). - Remove default protocol ports (
:80for HTTP and:443for HTTPS). - Sort query parameters alphabetically.
- Remove trailing slashes where appropriate.
- Decode unnecessary percent-encoded characters.
Abuse Prevention
- Rate Limiting: Implement token bucket rate limiting per API key (such as 100 URL creations per hour on the free tier).
- Malicious URL Filtering: Integrate with the Google Safe Browsing API to detect and block malicious or phishing URLs.
- CAPTCHA Verification: Require CAPTCHA challenges for anonymous or unauthenticated URL creation.
- Spam Detection: Validate destination domains against known spam and threat intelligence databases.
Bloom Filter for Duplicate Detection
- Use a distributed Bloom filter to optimize custom alias creation by quickly identifying definitely absent aliases before touching the database.
- Because a Bloom filter has no false negatives, a negative result guarantees the alias is currently absent from the filter, allowing the system to skip preliminary existence queries and proceed directly to the conditional insert.
- A positive result indicates the alias might already exist due to potential false positives, prompting the system to verify against or conditionally write to the authoritative store.
- The authoritative datastore (such as Cassandra with Lightweight Transactions) always enforces the final uniqueness decision. On successful creation, the new alias is added to the Bloom filter.
- With an illustrative 1% false-positive rate, approximately 1% of truly absent aliases may be reported as possibly present and take the slower check path, while 99% of absent aliases correctly test negative and avoid preliminary database existence reads. A filter size of ~1 GB comfortably supports 1 billion entries at this false-positive rate.
Monitoring and Alerting
- Redirect Latency: Track p50, p95, and p99 latency metrics against the 10-millisecond target.
- Cache Hit Ratio: Alert when the cache hit ratio falls below 80%.
- KGS Key Pool Health: Monitor remaining unused keys and alert when capacity drops below threshold levels.
- Kafka Consumer Lag: Track consumer lag on the
click-eventstopic to ensure the analytics pipeline remains healthy. - Error Rates: Monitor 4xx and 5xx response codes across all API endpoints.
GDPR and Data Privacy
- Allow users to permanently delete shortened links along with all associated analytics data.
- Anonymize IP addresses in analytics storage after 30 days.
- Provide data export capabilities for user records and analytics history.
Related Problems and Concepts
Short-link key generation and read-heavy caching strategies share patterns with Pastebin, which uses a similar database and object storage split at lower scale. Rate limiting on the URL creation endpoint applies concepts covered in API Rate Limiter. You can also review foundational caching strategies in Caching Patterns and Invalidation and back-of-the-envelope estimation guidelines in Back-of-the-Envelope Estimation.
Interview Walkthrough
- 25-Minute Interview Strategy
Focus on the core read-path caching and key generation decisions unless the interviewer explicitly requests multi-datacenter depth.
- Functional and Non-Functional Requirements (3 min)
- Capacity Calculations (5 min)
- Read Path and Multi-Layer Cache Architecture (8 min)
- Key Generation Service vs Hashing (5 min)
- Asynchronous Analytics Pipeline (4 min)
- Highlight the 100:1 read-to-write ratio early. Redirect performance is the core bottleneck, so design the read path before detailing URL creation.
- Discuss HTTP 301 vs 302 redirect trade-offs upfront, noting that 302 preserves click analytics while 301 reduces server traffic via browser and CDN caching.
- Compare KGS against URL hashing. Hashing offers automatic deduplication but requires collision retry handling, whereas KGS pre-allocates keys so write operations avoid runtime collision overhead.
- Detail the four caching layers (CDN, local memory, Redis cluster, and Cassandra) and connect the 1.9-millisecond average latency back to the 10-millisecond p99 SLO.
- Explain how a Bloom filter pre-check quickly identifies definitely absent custom aliases so the system can skip preliminary existence reads before Cassandra executes an authoritative conditional insert.
- Stream click analytics through Kafka to ensure background processing does not block the user redirect response.
- Use capacity calculations for 7-character Base62 encoding (approximately 3.5 trillion keys) to demonstrate that the design scales for decades.
Engineering Trade-offs
Interviewers evaluate how you navigate architectural forks and justify trade-offs. The sections below analyze the primary design decisions for this system.
Key Generation Strategy: KGS vs Hashing
Selecting the method for producing short codes impacts collision handling, write latency, and storage deduplication.
| Approach | Implementation | Pros | Cons |
|---|---|---|---|
| MD5 or SHA-256 Hash with Base62 | Hash the destination URL and take the first 7 characters | Deterministic and provides automatic global deduplication of identical URLs | Collision risks from the birthday paradox require runtime collision detection and retry loops |
| Centralized Counter | Increment a centralized database counter and Base62 encode the value | Simple to implement and produces unique sequential values | Single point of failure, predictable URL sequences, and a single-writer bottleneck at scale |
| Key Generation Service (KGS) ⭐ | Pre-generate keys offline in batches and dispense to application servers | Zero runtime collisions, non-sequential codes, O(1) allocation time, and no single point of failure with multiple instances | Unused in-memory keys are lost if a server restarts, and requires dedicated KGS infrastructure |
| UUID with Base62 | Generate a UUID v4 and encode to Base62 | Decentralized and simple generation | 128-bit UUID produces 22 Base62 characters, which is too long for a short URL |
| Snowflake ID with Base62 | Generate a 64-bit time-ordered unique ID and Base62 encode | Unique, decentralized, and sortable by time | 64-bit ID requires 11 Base62 characters and exposes the creation timestamp |
Why KGS is chosen for high-scale URL shorteners: The primary requirement is a compact 7-character short URL. KGS delivers exactly 7 characters with zero collision checks at runtime. At initial or modest scale, a simpler approach such as a database auto-increment counter or single-node key sequence may be sufficient. The distributed KGS architecture with offline pre-generation and leader election is a deliberate scalability choice that provides collision-free, non-sequential allocation and high write availability as traffic scales.
When to choose hashing instead: Choose hashing when global content deduplication is a mandatory requirement where every submission of the exact same long URL must map to the same short code without maintaining a secondary lookup index.
Database Selection: Cassandra vs Alternatives
The database tier must support high write throughput, massive dataset growth, and low-latency key-value lookups.
| Database | Strengths for This Use Case | Weaknesses |
|---|---|---|
| Cassandra ⭐ | Masterless architecture with no single point of failure, linear horizontal scaling, tunable consistency, built-in row TTL for automatic expiration, and native multi-datacenter replication | Does not support ACID transactions or multi-table joins, and requires operational expertise |
| DynamoDB | Fully managed by AWS, automatic scaling, single-digit millisecond latency, and built-in TTL support | Vendor lock-in, higher cost profile at extreme traffic volumes, and throughput provisioning limits |
| MySQL | Familiar ecosystem, strong ACID guarantees, and mature tooling | Single-writer primary creates write bottlenecks at scale, sharding requires manual maintenance, and lacks built-in TTL support |
| PostgreSQL | Rich feature set, robust JSON support, and strong relational integrity | Shares the same horizontal scaling constraints as MySQL for massive write-heavy key-value workloads |
Deciding Factors:
- The primary access pattern is pure key-value point lookup: no multi-table joins or relational constraints are required.
- Cassandra uses an LSM-tree storage engine that handles write bursts with consistent low latency.
- Cassandra provides native row-level TTL, automatically purging expired records during background compaction without application cleanup jobs.
- Native multi-datacenter replication allows read requests to be served locally from the nearest region.
- DynamoDB is a viable alternative if the infrastructure is AWS-native and managed operations are preferred, though Cassandra remains cost-effective at multi-billion record scale.
HTTP Redirect Status Codes: 301 vs 302
The choice of HTTP status code determines how browsers and intermediate caches handle redirects.
| Status Code | Behavior | Pros | Cons |
|---|---|---|---|
| 301 (Permanent Redirect) | Browsers and CDNs cache the target URL indefinitely or according to cache headers | Minimizes server load and provides the fastest redirect on repeat visits | Prevents per-click tracking on cached visits, and changing destination URLs requires cache invalidation |
| 302 (Temporary Redirect) ⭐ | Browsers do not cache the redirect and query the shortener on every click | Enables accurate real-time click tracking, geolocation capture, and dynamic destination updates | Incurs a network hop to the redirect service on every user click |
| 307 (Temporary Redirect, Preserves Method) | Similar to 302, but specifies that the HTTP request method remains unchanged according to RFC specifications | Preserves POST or PUT methods across redirects | Only relevant for non-GET requests, which are rarely used in public URL shortening |
Recommendation: Default to HTTP 302 redirects because analytics tracking and link flexibility are critical for most URL shortening platforms. Provide 301 redirects as an option for enterprise high-volume marketing campaigns where redirect speed takes precedence.
Read Path Optimization and Latency Budget
At peak loads exceeding 19,000 reads per second, multi-layer caching ensures the system meets its sub-10-millisecond latency target.
- Layer 1: CDN Edge (CloudFront or Fastly): For this illustrative design, assume that for 301 redirects the CDN caches the Location header at edge locations close to users, achieving roughly a 30% hit ratio with ~5ms response times.
- Layer 2: In-Process Application Cache (Caffeine): Each application server caches the top 10,000 most active URLs in memory with a 60-second TTL, absorbing an assumed 20% to 30% of remaining requests with ~0.01ms access times.
- Layer 3: Distributed Redis Cluster: Caches active URLs with a 24-hour TTL and jitter, handling roughly 90% of requests that reach the backend with ~0.5ms network latency.
- Layer 4: Cassandra Primary Storage: Serves cold URLs that miss the cache layers (approximately 5% of total traffic) with ~2ms to 5ms disk reads under LOCAL_ONE consistency.
| Layer | Technology | Hit Ratio | Latency | TTL |
|---|---|---|---|---|
| L1 | CDN Edge (CloudFront or Fastly) | ~30% | ~5ms | CDN cache headers |
| L2 | In-Process Cache (Caffeine) | ~20-30% | ~0.01ms | 60 seconds |
| L3 | Redis Cluster | ~90% of remaining | ~0.5ms | 24 hours with jitter |
| L4 | Cassandra | ~5% (cold URLs) | ~2-5ms | 5 years (persistent) |
Effective Latency Calculation: Under these illustrative cache-hit and component latency assumptions: 0.30 x 5ms + 0.25 x 0.01ms + 0.40 x 0.5ms + 0.05 x 3ms = ~1.85 milliseconds weighted average latency, comfortably within the 10-millisecond p99 latency target.
Encoding Schemes: Base62 vs Base58 and Base64
The character set chosen for short codes affects URL readability, length, and transmission safety.
| Encoding | Character Set | URL Safe | Character Ambiguity |
|---|---|---|---|
| Base62 | a-z, A-Z, 0-9 | Yes | Similar glyphs such as 0/O and 1/l can cause confusion if manually transcribed |
| Base58 | Base62 minus 0, O, l, and I | Yes | Eliminates visually ambiguous characters, ideal for manual entry |
| Base64 | Base62 plus + and / | No (requires percent-encoding for URLs) | Contains non-alphanumeric characters that need escaping |
Recommendation: Use Base62 for standard web shorteners where links are clicked digitally, as it maximizes combinations per character. Use Base58 if short codes are printed or frequently entered by hand.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.