System Design Problem

Design Pastebin

Commonly Asked By:GoogleDropboxAmazonMicrosoft

Interview Setup

Interview Prompt

Design a Pastebin style service where users paste text or code snippets and receive a short shareable URL

Clarifying Questions (ask before designing)

QuestionWhy it matters
Should the same content always map to the same URL, or does every paste get a unique key?Deduplication saves storage but alters privacy semantics. Pastebin assigns a fresh key per paste even when payloads are identical to prevent cross account metadata leaks.
What expiration options do we support, and what happens to expired pastes?Expiration policies dictate background cleanup workers, Redis TTL alignment, CDN cache-control headers, and S3 lifecycle rules.
Are private, unlisted, and burn after reading pastes in scope?Private pastes are owner only, require authenticated authorization, and must bypass public edge CDN caching. Unlisted pastes use secret URL access without authentication. Burn after reading requires atomic distributed locking and immediate destructive purging after the initial retrieval.
What scale are we designing for?The capacity model assumes 1M new pastes per day, roughly 12 writes per second, 300 peak reads per second, and 10 GB of new text storage daily. Scale assumptions dictate component choices.

Scope

In scope

  • Create, read, and delete paste APIs
  • Key generation service delivering collision free 8 character Base62 keys
  • Hybrid storage separating MySQL metadata from S3 content blobs
  • Multi tier read caching utilizing edge CDN and Redis clusters
  • Configurable expiration with background garbage collection
  • Private owner only and unlisted pastes with IP rate limiting
  • Asynchronous analytics pipelines via Kafka off the critical path

Out of scope (state explicitly)

  • Client desktop and mobile applications
  • Real time collaborative editing
  • Full machine learning content moderation pipelines beyond asynchronous scanning queues

Functional Requirements

Start by asking your interviewer about paste creation, raw text viewing, and expiration defaults. This problem is simpler than a full URL shortener, so confirm functional scope before introducing complex components.

  • Create paste: Submit text content and receive a unique short URL
  • Read paste: Retrieve and view the paste using its shareable URL
  • Expiration: Support configurable duration including 10 minutes, 1 hour, 1 day, 1 week, or permanent retention
  • Syntax highlighting: Support automatic or user selected programming languages
  • Unlisted pastes: Restrict discovery to clients possessing the secret URL
  • Private pastes: Restrict access to the authenticated owner. Private cache entries are scoped to the owner and are never publicly edge cached.
  • User accounts: Allow authenticated users to view, manage, and delete their saved pastes
  • Raw view: Provide plain text retrieval without HTML formatting or styling
  • Paste size limit: Enforce a strict payload ceiling of 10 MB per paste

Non-Functional Requirements

Interviewers focus primarily on read latency for popular pastes in this design. The 5:1 read to write ratio justifies a dedicated caching layer in front of object storage, following the caching patterns demonstrated in URL Shortener at an accessible operational scale.

  • High Availability: Deliver 99.99% uptime for the read path because paste viewing represents the core user experience
  • Low Latency: Ensure paste retrieval returns in under 100 ms at p99
  • Read Heavy: Support a read to write ratio of approximately 5:1
  • Scalability: Handle millions of stored pastes and scale beyond 10,000 read requests per second during peak viral traffic
  • Durability: Target durable committed paste content until its scheduled expiration, with S3 durability and replicated MySQL metadata
  • Unique URLs: Guarantee zero collisions across generated paste identifiers

Capacity Estimations

New pastes per day and average paste payload sizes determine the primary storage and cache requirements. These values provide the baseline numbers for sizing.

MetricCalculationValue
New pastes / dayGiven1M
Reads / dayGiven5M
Writes / sec1M ÷ 86400~12 average
Reads / sec5M ÷ 86400~58 average, peak 300
Avg paste sizeGiven10 KB
Storage / day1M x 10 KB10 GB raw before compression
Storage / year10 GB x 3653.65 TB
Metadata per pasteGiven500 bytes

Architecture Diagram

In an interview setting, this system represents a read heavy key value architecture where MySQL metadata, S3 blob storage, and a Redis cache tier satisfy requirements.

Walk your interviewer through the system by traffic type. On the create path, the service generates a short key, stores metadata in MySQL, and writes content to S3. On the read path, where the vast majority of traffic lands, requests query CloudFront CDN, then Redis cache, and fall back to MySQL and S3 only on a cache miss. Time to live settings and background workers clean up expired pastes without requiring blocking full table scans.

Loading...

Component Deep Dives

CDN (CloudFront / Fastly)

Key generation, object storage, and read caching paths are straightforward. This keeps the design lean and cost effective.

Paste reads are globally distributed and content is immutable once created. The CDN caches rendered HTML and raw text pages at edge locations using the paste URL path as the cache key, applying a time to live of min(expiry_remaining, 24 hours) for expiring pastes and 7 days for permanent pastes.

API Gateway

The API gateway centralizes authentication, rate limiting, and request routing before traffic reaches internal services. It enforces token bucket rate limiting by client IP, allowing 10 pastes per hour for anonymous users and 100 pastes per hour for authenticated accounts. Anonymous creation requests require a CAPTCHA challenge during burst traffic, while the gateway handles SSL termination and uniform request validation.

Key Generation Service (KGS)

At Pastebin scale, random key generation often suffices, whereas a dedicated Key Generation Service becomes necessary when write throughput increases, as explored in URL Shortener.

The Key Generation Service pre allocates 8 character Base62 identifiers in batches from a theoretical space of 218 trillion keys, persisting them in a dedicated key pool database partitioned into unused_keys and used_keys tables. During paste creation, application workers atomically pop a preallocated batch of keys from memory. This approach eliminates runtime collision checks, generates non sequential identifiers that make enumeration impractical, and maintains submillisecond key allocation times while local batches remain available.

Write Service

The write path must remain durable and idempotent. The Write Service receives the paste payload, expiration policy, syntax language, visibility, and idempotency key, acquires a pregenerated unique key from the local KGS buffer, and uploads the compressed body to S3 under s3://pastebin-content/{key_prefix}/{key}. Once the S3 upload succeeds, it commits the metadata record and a paste created row to the MySQL outbox in the same transaction. The API returns 201 Created after the database commit. If the database transaction fails after the S3 upload, the object is treated as an orphan candidate and cleaned by the lifecycle worker or retry path. For authenticated users, the idempotency key is scoped to the user so retries return the original result without creating a second paste. The server derives idempotency_scope from the authenticated user or anonymous request principal. Anonymous requests that need retry safety therefore use a stable request principal together with the idempotency key.

Read Service

The read path is the latency critical workflow that demands multi tier caching. The Read Service resolves visibility before serving a cached document. Public and unlisted requests use the standard Redis key. Private requests require owner authorization and use an auth aware Redis key. On a cache miss, the service queries MySQL to confirm that the paste exists, verifies authorization when the paste is private, and checks that it has not expired. Burn after reading pastes bypass all caches and use an atomic MySQL consume claim so only one request can obtain the right to read the content. For a standard paste, the service fetches compressed content from S3 and asynchronously warms Redis with a time to live of min(paste expiry, 1 hour) before returning the content with syntax highlighting.

Safe Rendering and Raw Content

Rendered paste pages must HTML escape user supplied content and use a strict Content Security Policy so pasted code cannot become stored cross site scripting. Raw responses should use text/plain with MIME sniffing disabled. Syntax highlighting must operate on escaped text and never execute the pasted content.

Content Storage: S3 vs Database

Large pastes are stored in S3, while compact metadata stays in MySQL for fast indexed key lookups. Paste bodies reach up to 10 MB, making them ill suited for relational database storage where large text blobs bloat buffer pools, degrade index traversal, and complicate backup schedules. Amazon S3 provides 11 nines durability and cost effective object storage, allowing MySQL to retain only compact 500 byte metadata records that index cleanly. Compressing text payloads with zstd prior to S3 upload yields 3x to 10x storage reduction and significantly curtails bandwidth expenses.

Cleanup Worker

Background workers pull expired records and report completion without blocking user traffic. A background cleanup worker runs on an hourly schedule to query expired records from MySQL in manageable batches. For each expired batch, the worker marks records expired when needed, deletes the underlying objects from S3, purges eligible metadata rows from MySQL, evicts matching keys from Redis, and writes a paste-deleted event to the MySQL outbox. The Outbox Relay then publishes that event to Kafka for downstream analytics and cache invalidation. The delete path immediately invalidates Redis entries after the MySQL transaction and publishes the deletion through the outbox. The Cache Invalidation Worker retries Redis invalidation when needed and purges public CDN objects. Origin metadata and expiration checks are authoritative immediately, while public CDN invalidation is asynchronous, so a bounded stale edge cache window can exist until the purge completes. A burn after reading paste uses the same deletion pipeline after its atomic consume claim, and if physical deletion fails after the claim, the consumed marker prevents another read while cleanup retries the storage deletion.

Analytics Worker

Analytics runs on a separate SLA from the hot path. The Analytics Worker consumes paste-created and paste-viewed events from Kafka and executes periodic micro batch inserts into ClickHouse. The View Count Batch Updater applies idempotent aggregate deltas to MySQL and updates or invalidates affected Redis entries after successful commits so cached view counts converge. This keeps view counts and trending metrics eventually consistent without imposing latency penalties or database write contention on the primary read path.

Event Bus Design (Kafka)

The event bus decouples producers from consumers and buffers traffic spikes across downstream subsystems. Database state changes use a transactional outbox so Kafka publication does not depend on a second non atomic write.

YAML
topics:
  paste-created:
    partitions: 16
    partition_key: paste_key
    retention_days: 7
    producers:
      - Outbox Relay (publishes after the MySQL metadata and outbox row commit)
    consumers:
      - Analytics Worker (batch writes to ClickHouse)
      - Content Moderation Scanner (asynchronous security scan)
    dlq: paste-created-dlq

  paste-deleted:
    partitions: 8
    partition_key: paste_key
    retention_days: 7
    producers:
      - Delete Service (writes the delete event to the MySQL outbox)
      - Cleanup Worker (writes the delete event to the MySQL outbox)
    consumers:
      - Analytics Worker
      - Cache Invalidation Worker (CDN and Redis)

  paste-viewed:
    partitions: 32
    partition_key: paste_key
    retention_days: 3
    producers:
      - Read Service (buffers every counted view in memory. Optional sampling is reserved for analytics telemetry only)
    consumers:
      - View Count Batch Updater (idempotent aggregate updates and Redis refresh/invalidation)

replication:
  replication_factor: 3
  min_insync_replicas: 2

API Design

Service Interface (TypeScript)

Create, read, and delete endpoints form the core API surface. Strong typing and explicit expiration controls keep the contract clear.

Typed contract defining paste lifecycle operations, payload schemas, and client responses:

TYPESCRIPT
export type PasteKey = string;
export type AuthToken = string;
export type ISODateTime = string;
export type ExpiryOption = "10m" | "1h" | "1d" | "1w" | "never";
export type PasteVisibility = "public" | "unlisted" | "private";

export interface CreatePasteRequest {
  content: string;
  title?: string;
  language?: string;
  expiry?: ExpiryOption;
  visibility?: PasteVisibility;
  burn_after_reading?: boolean;
  idempotency_key?: string;
}

export interface CreatePasteResponse {
  key: PasteKey;
  url: string;
  raw_url: string;
  expires_at: ISODateTime | null;
}

export interface PasteMetadata {
  key: PasteKey;
  title: string | null;
  language: string;
  content_size: number;
  s3_path: string;
  visibility: PasteVisibility;
  burn_after_reading: boolean;
  view_count: number;
  expires_at: ISODateTime | null;
  created_at: ISODateTime;
}

export type Paste = PasteMetadata & { content: string };

export interface PastebinService {
  createPaste(request: CreatePasteRequest, authToken?: AuthToken): Promise<CreatePasteResponse>;
  getPaste(key: PasteKey, authToken?: AuthToken): Promise<Paste>;
  deletePaste(key: PasteKey, authToken: AuthToken): Promise<void>;
}

Create Paste

HTTP POST endpoint for uploading text payloads with optional expiration and privacy settings. Creating a private paste requires the authenticated owner's identity.

HTTP
POST /api/v1/pastes
Content-Type: application/json
Idempotency-Key: 01JPASTEEXAMPLE000000000000

{
  "content": "print('Hello, World!')",
  "language": "python",
  "expiry": "1d",
  "visibility": "public",
  "burn_after_reading": false,
  "title": "My first paste"
}

Response: 201 Created
Content-Type: application/json

{
  "key": "abc12345",
  "url": "https://pastebin.com/abc12345",
  "raw_url": "https://pastebin.com/raw/abc12345",
  "expires_at": "2026-03-14T10:00:00Z"
}

Read Paste

HTTP GET endpoint for retrieving paste content, metadata, and view metrics by key:

HTTP
GET /api/v1/pastes/{key}

Response: 200 OK
Content-Type: application/json

{
  "key": "abc12345",
  "content": "print('Hello, World!')",
  "language": "python",
  "title": "My first paste",
  "created_at": "2026-03-13T10:00:00Z",
  "expires_at": "2026-03-14T10:00:00Z",
  "burn_after_reading": false,
  "view_count": 42
}

Delete Paste

HTTP DELETE endpoint requiring bearer token authentication to logically delete a paste and begin asynchronous storage and cache cleanup:

HTTP
DELETE /api/v1/pastes/{key}
Authorization: Bearer <token>

Response: 204 No Content

Common Error Responses

Standardized error payloads returned across validation, authorization, and rate limiting boundaries:

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff

Data Model

MySQL: Paste Metadata

Structured metadata resides in MySQL for transactional lookups. Raw text payloads are persisted in S3 object storage and served through a multi layer Redis cache.

Relational schema storing indexed metadata, user associations, and expiration timestamps:

SQL
CREATE TABLE pastes (
    paste_key       CHAR(8) PRIMARY KEY,
    user_id         UUID,
    title           VARCHAR(256),
    language        VARCHAR(32) DEFAULT 'text',
    content_size    INT,
    s3_path         VARCHAR(256),
    sha256_checksum CHAR(64) NOT NULL,
    visibility      ENUM('public', 'unlisted', 'private') NOT NULL DEFAULT 'public',
    burn_after_reading BOOLEAN NOT NULL DEFAULT FALSE,
    view_count      BIGINT UNSIGNED NOT NULL DEFAULT 0,
    idempotency_key VARCHAR(64),
    idempotency_scope VARCHAR(128),
    expires_at      TIMESTAMP NULL,
    consumed_at     TIMESTAMP NULL,
    deleted_at      TIMESTAMP NULL,
    created_at      TIMESTAMP NOT NULL,
    UNIQUE KEY uniq_idempotency (idempotency_scope, idempotency_key),
    INDEX idx_user (user_id, created_at DESC),
    INDEX idx_expiry (expires_at)
);

S3: Paste Content

Object storage key prefixing and storage class configuration for compressed paste bodies:

YAML
bucket: pastebin-content
key_pattern: "pastes/{first_2_chars_of_key}/{paste_key}"
storage_class: STANDARD
compression: zstd
metadata:
  content_type: text/plain
  max_size_bytes: 10485760
  sha256_checksum: e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855

checksum_policy: "store SHA-256 in MySQL metadata and object metadata; validate on read"

Redis Cache

In memory cache aside hash structure storing active paste payloads and TTL bounds:

REDIS
key_pattern: "paste:{key}"
private_key_pattern: "private-paste:{user_id}:{key}"
data_structure: HASH
ttl_strategy: "min(paste_expiry_remaining, 3600 seconds)"
fields:
  content: "print('Hello, World!')"
  language: python
  title: "My first paste"
  visibility: public
  expires_at: "2026-03-14T10:00:00Z"
  view_count: 42

KGS: Key Pool

Relational schema supporting atomic key preallocation and assignment tracking. Idempotency scope is derived server side from the authenticated user or the anonymous request principal so retries remain isolated between callers.

At approximately 12 writes per second, a Key Generation Service is optional because cryptographically random 8 character Base62 keys with a Bloom filter precheck and MySQL UNIQUE constraint easily suffice. The service is included here because it mirrors URL Shortener and eliminates collision retries under higher write traffic.

SQL
CREATE TABLE unused_keys (
    key_value   CHAR(8) PRIMARY KEY
);

CREATE TABLE used_keys (
    key_value   CHAR(8) PRIMARY KEY,
    used_at     TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);

Transactional Outbox

Reliable event publication that closes the database to Kafka dual write gap.

SQL
CREATE TABLE paste_event_outbox (
    event_id        CHAR(36) PRIMARY KEY,
    paste_key       CHAR(8) NOT NULL,
    event_type      VARCHAR(32) NOT NULL,
    payload         JSON NOT NULL,
    created_at      TIMESTAMP NOT NULL,
    published_at    TIMESTAMP NULL,
    INDEX idx_publish (published_at, created_at)
);

Outbox flow:
  1. Commit paste metadata and the outbox row in one MySQL transaction.
  2. Outbox Relay publishes the event to Kafka with event_id.
  3. Relay retries until Kafka acknowledges publication.
  4. Consumers deduplicate by event_id so retries do not duplicate effects.

Kafka Topics

Topic partition strategies and consumer group bindings for asynchronous event fan out:

YAML
topics:
  paste-created:
    partitions: 16
    partition_key: paste_key
    retention_days: 7
    producers:
      - Outbox Relay (publishes after the MySQL metadata and outbox row commit)
    consumers:
      - Analytics Worker (batch writes to ClickHouse)
      - Content Moderation Scanner (asynchronous security scan)
    dlq: paste-created-dlq

  paste-deleted:
    partitions: 8
    partition_key: paste_key
    retention_days: 7
    producers:
      - Delete Service (writes the delete event to the MySQL outbox)
      - Cleanup Worker (writes the delete event to the MySQL outbox)
    consumers:
      - Analytics Worker
      - Cache Invalidation Worker (CDN and Redis)

  paste-viewed:
    partitions: 32
    partition_key: paste_key
    retention_days: 3
    producers:
      - Read Service (buffers every counted view in memory. Optional sampling is reserved for analytics telemetry only)
    consumers:
      - View Count Batch Updater (idempotent aggregate updates and Redis refresh/invalidation)

replication:
  replication_factor: 3
  min_insync_replicas: 2

Fault Tolerance

ConcernSolution
S3 unavailableUse cached reads where available, retry writes, and fail over reads to the secondary S3 region
MySQL failurePromote the standby replica to primary and route reads to remaining replicas
KGS failureApplication servers serve from local in memory buffers containing 1000 keys each
Expired paste servedEnforce expires_at on origin reads and align Redis TTL with expiration
Content corruptionCryptographic SHA 256 checksum stored in metadata and validated on read
Duplicate keyKGS preallocation uniqueness backed by database primary key constraints
Lost Kafka eventCommit metadata and the outbox row atomically, then relay to Kafka with retries and event ID deduplication
Stale CDN or Redis cache after deleteMark the paste deleted and write the deletion event in the MySQL transaction, then invalidate Redis immediately and publish through the outbox so cache invalidation and public CDN purge can be retried. Origin metadata remains authoritative, and a bounded stale CDN edge window can exist until invalidation completes

Additional Considerations

Related Problems

Operational safeguards for viral paste spikes, background garbage collection, and defensive security measures. These areas connect directly to the related systems and patterns below.

Key generation and read heavy caching overlap with URL Shortener. Object storage patterns match Dropbox or Google Drive at a simpler scale. Foundational caching patterns connect to Caching Patterns and Invalidation.

Abuse Prevention

Defensive controls combine IP based token bucket rate limiting with automated CAPTCHA challenges on burst creation attempts. Incoming paste content undergoes asynchronous machine learning scans to detect phishing URLs, malware payloads, and credentials, while size constraints restrict anonymous pastes to 512 KB and authenticated uploads to a maximum of 10 MB.

Analytics

View counts are eventually consistent. The Read Service buffers view events for Kafka, and the batch updater applies aggregate deltas to MySQL in batches to minimize write pressure during viral traffic surges. Consumers deduplicate repeated event IDs before applying the deltas. A Read Service or process failure can lose events that remain only in memory, so the displayed count can temporarily undercount views until reconciliation. The displayed count is therefore an operational metric rather than an authoritative ledger of every read. These event streams also feed operational dashboards that visualize trending pastes, geographic readership patterns, and programming language popularity distributions.

Burn After Reading

Burn after reading pastes enforce single view consumption by claiming the paste exactly once before allowing it to be served. The Read Service bypasses CDN and Redis, fetches the content from S3, and then executes a conditional MySQL transaction such as UPDATE pastes SET consumed_at = NOW(), deleted_at = NOW() WHERE paste_key = ? AND consumed_at IS NULL AND deleted_at IS NULL while inserting the corresponding paste-deleted outbox event. Only the request that successfully changes one row may return the fetched content. A losing concurrent request receives a not found or already consumed response. Physical S3 deletion and cache invalidation happen asynchronously after the transaction commits. If the process fails after the consume claim, the consumed and deleted markers still prevent another read, and cleanup retries the physical deletion. A Redis lock or Lua script may be used as an optimization for contention, but MySQL provides the authoritative single consume decision.

Interview Walkthrough

  • 25 minute cut

    Skip multi region and deep scale mechanics unless interviewing for staff roles.

    • Functional requirements, nonfunctional requirements, and the 5:1 read ratio (3 min)
    • Separating MySQL metadata from S3 blob content (6 min)
    • Key generation strategies at this operational scale (5 min)
    • Four tier read path traversing CDN, Redis, MySQL, and S3 (7 min)
    • Time to live expiration without expensive table sweep jobs (4 min)
  • Compare the design to a URL shortener while emphasizing whether content addressable keys or random identifiers should be used, clarifying whether identical content maps to the same URL.
  • Explain the read heavy ratio where reads outnumber writes by 5:1, and layer caching patterns on hot pastes across edge CDN nodes and Redis memory clusters.
  • Detail expiration time to live mechanisms and asynchronous garbage collection of expired content from object storage and the relational metadata store.
  • Discuss syntax highlighting as an asynchronous processing step rather than blocking the synchronous creation response.
  • Describe content moderation workflows for malware and spam detection through an asynchronous queue. Flagged public pastes can be quarantined or removed after creation, and moderation can run before any optional search or discovery indexing.
  • Avoid the common pitfall of storing paste bodies directly in the relational database, because large text belongs in S3 blob storage with only a lightweight metadata row in MySQL.

Engineering Trade-offs

Storage Architecture: MySQL + S3 vs Pure NoSQL

Key architectural decisions center on blob storage partitioning and garbage collection cadence for expired content.

Evaluating relational metadata pairing against single store NoSQL architectures:

Option 1: MySQL (metadata) + S3 (content) [Selected]
  - Paste content reaches up to 10 MB, whereas relational database rows are not optimized for large text blobs
  - S3 is designed for 11 nines of durability, cost effective storage, and native CDN caching compatibility
  - MySQL provides ACID transactions for metadata, indexed key lookups, and clear lifecycle queries

Option 2: Cassandra for both metadata and content
  - Advantage: Eliminates single points of failure with native peer to peer replication
  - Drawback: 2 GB maximum cell ceiling. TTL is available on column writes, but generated tombstones and compaction can degrade read performance
  - Drawback: Does not manage the lifecycle of external S3 objects

Option 3: DynamoDB + S3
  - Advantage: Fully managed infrastructure with automated capacity scaling
  - Drawback: Cloud vendor lock-in and higher request and indexing costs at high read volume

Decision: MySQL plus S3 delivers the optimal balance of transactional metadata integrity and cost effective blob storage.

Expiry Implementation

Comparing scheduled database triggers, NoSQL cell TTLs, and asynchronous lifecycle workers:

Option 1: MySQL Scheduled Events
  - Advantage: Operates without external infrastructure and runs within database transactions
  - Drawback: Fails to purge S3 objects and degrades performance during burst expirations

Option 2: Cassandra TTL
  - Advantage: Requires zero application tier cleanup logic
  - Drawback: Generates tombstones that cause read performance degradation and does not clean S3 objects

Option 3: Background Worker + S3 Lifecycle Policy [Selected]
  - Background worker runs hourly to delete matching database rows, S3 objects, and Redis cache entries
  - Amazon S3 Lifecycle rules target an orphan cleanup prefix or tag and serve as a safety net after six months
  - Advantage: Transactional cleanup handles high volume deletions in manageable batches
  - Trade-off: Expired database rows and orphaned objects can remain until cleanup runs, but read paths enforce expiration immediately

Read Path Optimization

Multi tier caching efficiency across edge CDN, in memory Redis, and origin storage tiers:

Standard read path:
Client -> CDN (edge hit ~40%) -> Redis (hit ~85%) -> MySQL metadata -> S3 content

Cache distribution policies:
  Public pastes: CloudFront CDN caches rendered HTML and raw text pages
  Unlisted pastes: Bypass public edge CDN and cache in Redis because access is controlled by the secret URL
  Private pastes: Require authentication, bypass public edge CDN, and use an auth aware Redis key
  Burn after reading pastes: Strictly bypass all caches to enforce single view semantics

Effective multi tier hit rates:
  CDN edge tier: Serves approximately 40% of total read traffic
  Redis memory cache: Absorbs approximately 85% of remaining requests
  MySQL and S3 origin: Handles approximately 9% of requests
  Net result: Under 10% of total read requests reach MySQL and S3 origin storage

Deduplication: Same Content, Same URL?

Pastebin deliberately avoids content based deduplication because the storage savings for lightweight text snippets are negligible and do not justify the added architectural complexity. Furthermore, assigning unique cryptographic identifiers to identical content preserves user privacy expectations and prevents cross account enumeration leaks that arise when deduplication is enforced across tenants.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...