System Design Problem

Design a Flash Sale System

Commonly Asked By:AlibabaAmazonFlipkartShopify

Interview Setup

Interview Prompt

Design a flash sale system where 1M users compete for 1,000 items in seconds, ensuring strict FIFO ordering, zero inventory oversell, bot protection, and post-reservation payment authorization.

Clarifying Questions (ask before designing)

QuestionWhy it matters
Allow all 500K RPS to hit Redis directly?Redis can handle the raw rate, but application servers and load balancers cannot, making a virtual queue essential.
Reserve stock before or after payment?Reserve stock first with a 10-minute payment window. Release stock with an increment when payment expires or fails.
Per-user purchase limit policy?Lua scripts enforce a maximum of 2 units per user ID and unique device fingerprint.
Expected bot traffic ratio?Anticipate 30% to 50% automated bot traffic at T=0, mitigated by CAPTCHA, TLS fingerprinting, and account age validation.

Scope

In scope

  • Virtual waiting room
  • Redis atomic stock decrement
  • Admission JWT tokens
  • Anti-bot security layers
  • Payment reservation lifecycle
  • Real-time sold-out broadcasts

Out of scope (state explicitly)

  • Full payment gateway processing
  • Product catalog management
  • Warehouse inventory management

Functional Requirements

Confirm the traffic spike profile with your interviewer before proposing components. Inquire about concurrent users at T=0, purchase limits per buyer, and whether a virtual waiting room or anti-bot defense falls within scope.

  • Scheduled sale lifecycle: Administrators schedule flash promotions with precise start and end times, assigned SKU inventory, and steep promotional discounts.
  • Synchronized countdown timer: Display a synchronized countdown on client interfaces, revealing product items and purchase buttons at the exact start timestamp.
  • Atomic stock reservation: Decrement inventory atomically and reserve units for a dedicated payment window, as contrasted with standard cart flows in the Shopping Cart System.
  • Virtual waiting room: Enqueue surplus traffic into a structured queue when concurrent volume exceeds cluster processing capacity.
  • Purchase limits per customer: Enforce a strict cap of 1 to 2 units per user and verified device fingerprint to block scalpers.
  • Real-time stock indicators: Stream remaining available inventory counts to active users and broadcast immediate sell-out notifications.
  • Strict admission fairness: Provide first-come, first-served admission where refreshing browsers provides zero competitive advantage.
  • Anti-bot defenses: Filter out automated scrapers and headless checkout bots from sniping available inventory before real users can purchase.

Non-Functional Requirements

Interviewers prioritize zero inventory overselling and resilient stability under sudden million-user traffic bursts. Establish the critical invariant upfront: if 1,000 units are allocated, exactly 1,000 successful orders can be confirmed.

  • Extreme Throughput: Sustain 1M+ concurrent active connections and 500K purchase attempts per second at sale kickoff.
  • Low Latency: Return atomic purchase decisions in under 100 ms p99.
  • Strong Consistency: Maintain strict inventory accuracy with zero overselling, respecting the principles in CAP Theorem and Consistency Models.
  • Graceful Degradation: Protect downstream payment and order management services from cascade failures via admission gating.
  • Guaranteed Fairness: Enforce verifiable first-in, first-out (FIFO) queue ordering using synchronized timestamp rankings.
  • Idempotency: Ensure rapid double-clicks or client network retries never generate duplicate inventory reservations.

Capacity Estimations

Calculate capacity constraints before defending Redis as the primary reservation hot path. Peak purchase requests and concurrent user volume determine whether a virtual queue is mandatory above the atomic decrement layer.

MetricCalculationValue
Concurrent users at sale startPeak assumption1M+
Purchase attempts / sec (T=0)Peak assumption500K
Items for saleGiven1,000 - 10,000 units
Time to sell outGiven5-30 seconds
Page load requests / secPeak refresh-storm assumption2M (pre-sale refresh storm)
Bot traffic ratioGiven30-50% of requests

Architecture Diagram

In the room: draw the virtual queue before the Redis Lua purchase path, because admission control separates browse traffic from the atomic buy lane.

Trace the admission funnel from edge distribution to order completion. Pre-rendered static landing pages live on global CDN caches, ensuring millions of pre-sale refreshes bypass origin servers entirely. Waiting customers enter a virtual queue, receive short-lived admission tokens upon dequeue, and only then access the Redis Lua purchase path, keeping browse traffic strictly decoupled from the atomic buy lane.

Loading...

Component Deep Dives

The Critical Purchase Path: Redis Lua Script

This is the core implementation detail interviewers probe most aggressively. The entire reservation decision must execute atomically in Redis in under 1 ms, eliminating the race conditions inherent in separate read, decrement, and write commands. The script also makes retries idempotent and checks the per user and per device limits in the same atomic path. For broader patterns, explore Redis Patterns.

LUA
-- Keys:
-- 1: flash_stock:{sale_id}:{sku_id}
-- 2: user_limit:{sale_id}:{user_id}
-- 3: idempotency:{sale_id}:{idempotency_key}
-- 4: reservation:{reservation_token}
-- 5: device:{sale_id}:{fingerprint}
-- Args: user_id, sku_id, quantity, reservation_token, idempotency_key, device_fingerprint

-- Step 0: Return the original result for a retried request
local previous = redis.call('GET', KEYS[3])
if previous then
  return {2, previous}
end

-- Step 1: Check per-user purchase limit
local user_purchased = redis.call('GET', KEYS[2])
local quantity = tonumber(ARGV[3])
if user_purchased and tonumber(user_purchased) + quantity > 2 then
  return {0, 'LIMIT_EXCEEDED'}
end

-- Step 2: Check per-device purchase ownership
local device_user = redis.call('GET', KEYS[5])
if device_user and device_user ~= ARGV[1] then
  return {0, 'DEVICE_LIMIT_EXCEEDED'}
end

-- Step 3: Atomic stock decrement
local remaining = redis.call('DECRBY', KEYS[1], quantity)
if remaining < 0 then
  redis.call('INCRBY', KEYS[1], quantity)  -- undo
  return {0, 'SOLD_OUT'}
end

-- Step 4: Record purchase limits and reservation
redis.call('INCRBY', KEYS[2], quantity)
redis.call('EXPIRE', KEYS[2], 86400)
redis.call('SET', KEYS[5], ARGV[1], 'EX', 86400)
redis.call('SET', KEYS[4], cjson.encode({
  user_id = ARGV[1],
  sku_id = ARGV[2],
  qty = quantity
}), 'EX', 600)

local result = cjson.encode({
  reservation_token = ARGV[4],
  remaining_stock = remaining
})
redis.call('SET', KEYS[3], result, 'EX', 86400)

return {1, result}

Why execute a Lua script rather than individual commands? Redis runs Lua scripts atomically without interleaving other commands. Separate read, decrement, and check calls expose a window where concurrent requests can oversell. A single Redis node can comfortably serialize 100K+ Lua executions per second, and sharding via Sharding and Partitioning scales this further across independent sales.

Virtual Queue: Handling 1M Concurrent Users

Even if Redis sustains high decrement throughput, the upstream API gateway and load balancer tier can collapse under a 500K request per second burst. A virtual waiting room regulates inbound pressure by admitting customers at a stable, controlled drain rate that can be tuned to inventory and downstream capacity.

T-5 min: Users "Enter Queue" early
  Position assigned: INCR queue_position:{sale_id} --> position 347,231
  User shown: "Your position: 347,231. Estimated wait: ~5 minutes"

T=0: Sale starts. Queue processes users FIFO.
  Gate rate: 10,000 users admitted per second (tunable)
  
  Admitted users:
  1. Receive short-lived JWT token (valid 60 seconds)
  2. Token authorizes call to purchase API
  3. Purchase API validates token --> runs Lua script on Redis
  
  Users not yet admitted:
  - See "Please wait..." with live position via WebSocket/SSE
  - Position updates every 5 seconds
  
  Stock gone:
  - Broadcast SOLD_OUT to ALL remaining queue members immediately
  - Don't make users wait if nothing left to buy

Queue implementation:
  Redis sorted set: ZADD queue:{sale_id} {timestamp} {user_id}
  Processing: ZPOPMIN queue:{sale_id} 10000  (pop 10K per second)

Anti-Bot Measures

Scalpers and automated purchase bots represent a primary operational hazard. Defend against unauthorized syndicates by deploying layered inspection filters from the network edge to the checkout service.

Layer 1: CDN/WAF (Cloudflare, AWS WAF)
  - Rate limit per IP: max 10 req/sec
  - Known bot signatures blocked
  - JavaScript challenge (bots can't execute JS)
  - TLS fingerprinting (JA3 hash) flag non-browser clients

Layer 2: Queue Entry Validation
  - CAPTCHA at queue entry (invisible reCAPTCHA)
  - Device fingerprint (canvas hash, WebGL, screen resolution)
  - Account age check: accounts < 24 hours old are blocked

Layer 3: Purchase Validation
  - One purchase per user_id (Redis user_limit)
  - One purchase per device_fingerprint
  - One purchase per payment method

Layer 4: Post-Purchase Fraud Detection
  - Multiple orders to same shipping address from different accounts are cancelled
  - Reseller pattern detection flags suspicious activity

API Design

API payloads use explicit domain types so request and response contracts remain stable as the implementation evolves.

TYPESCRIPT
type EnterQueueRequest = {
  captcha_token: string;
  device_fingerprint: string;
};

type EnterQueueResponse = {
  queue_position: number;
  estimated_wait_seconds: number;
  queue_token: string;
};

type PurchaseRequest = {
  sku_id: string;
  quantity: number;
};

type PurchaseResponse = {
  status: "reserved" | "sold_out" | "limit_exceeded" | "device_limit_exceeded";
  reservation_token?: string;
  payment_deadline?: string;
  remaining_stock?: number;
};

type SaleStatusItem = {
  sku_id: string;
  name: string;
  flash_price: number;
  original_price: number;
  total_stock: number;
  remaining: number;
};

type SaleStatusResponse = {
  sale_id: string;
  status: "scheduled" | "active" | "ended" | "cancelled";
  items: SaleStatusItem[];
};

Enter Queue

Clients invoke this endpoint to receive a queue placement and initial wait estimate. The request maps to EnterQueueRequest and the response maps to EnterQueueResponse.

HTTP
POST /api/v1/flash-sale/{sale_id}/enter-queue
Content-Type: application/json

{
  "captcha_token": "recaptcha-response-token",
  "device_fingerprint": "fp-hash-abc"
}

Response: 200 OK
Content-Type: application/json

{
  "queue_position": 12345,
  "estimated_wait_seconds": 120,
  "queue_token": "qt-uuid"
}

Purchase (After Admitted)

Admitted clients supply their cryptographic admission token along with an idempotency key to claim inventory. The request and response map to PurchaseRequest and PurchaseResponse.

HTTP
POST /api/v1/flash-sale/{sale_id}/purchase
Content-Type: application/json
Idempotency-Key: "purchase-user123-sale456"
Authorization: Bearer {admission_jwt}

{
  "sku_id": "SKU-FLASH-1",
  "quantity": 1
}

Response: 200 OK
Content-Type: application/json

{
  "status": "reserved",
  "reservation_token": "res-uuid",
  "payment_deadline": "2026-03-14T11:10:00Z",
  "remaining_stock": 423
}

OR { "status": "sold_out" }
OR { "status": "limit_exceeded" }
OR { "status": "device_limit_exceeded" }

Get Sale Status

Clients can poll this endpoint or receive equivalent updates over WebSockets to display live stock remaining and event status. The response maps to SaleStatusResponse.

HTTP
GET /api/v1/flash-sale/{sale_id}/status
Response: 200 OK
Content-Type: application/json

{
  "sale_id": "sale-456",
  "status": "active",
  "items": [
    {"sku_id": "SKU-FLASH-1", "name": "iPhone 16", "flash_price": 499.00,
     "original_price": 999.00, "total_stock": 1000, "remaining": 423}
  ]
}

Common Error Responses

Structured error representations for rejected admissions, rate limits, and sold-out states.

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
402 Payment Required: account balance or payment method has insufficient funds
502 Bad Gateway: payment gateway provider timeout, poll transaction status endpoint

Data Model

Redis: Flash Sale State

In memory keys maintain stock counters, user purchase limits, reservation holds, admission state, device ownership, and idempotency records for high speed evaluation.

REDIS
flash_stock:{sale_id}:{sku_id}      --> INT (atomic DECR)
user_limit:{sale_id}:{user_id}      --> INT (max 2), TTL 86400
reservation:{token}                 --> JSON { user_id, sku_id, qty }, TTL 600
queue_position:{sale_id}            --> INT (INCR for each entrant)
queue:{sale_id}                     --> Sorted Set { user_id: monotonic position }
admission:{sale_id}:{user_id}       --> "admitted", TTL 60
device:{sale_id}:{fingerprint}      --> user_id, TTL 86400
idempotency:{sale_id}:{key}         --> JSON reservation result, TTL 86400

PostgreSQL: Durable Records

Relational tables provide durable audit trails and order histories populated asynchronously after successful Redis reservation holds.

SQL
CREATE TABLE flash_sales (
    sale_id         UUID PRIMARY KEY,
    name            VARCHAR(255),
    start_time      TIMESTAMPTZ NOT NULL,
    end_time        TIMESTAMPTZ NOT NULL,
    status          ENUM('scheduled','active','ended','cancelled'),
    created_at      TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE flash_sale_items (
    sale_id         UUID NOT NULL,
    sku_id          VARCHAR(50) NOT NULL,
    flash_price     DECIMAL(10,2) NOT NULL,
    original_price  DECIMAL(10,2) NOT NULL,
    total_stock     INT NOT NULL,
    sold_count      INT DEFAULT 0,
    PRIMARY KEY (sale_id, sku_id)
);

CREATE TABLE flash_sale_orders (
    order_id        UUID PRIMARY KEY,
    sale_id         UUID NOT NULL,
    user_id         UUID NOT NULL,
    sku_id          VARCHAR(50) NOT NULL,
    quantity        INT NOT NULL,
    price           DECIMAL(10,2),
    status          ENUM('reserved','paid','cancelled','expired'),
    reservation_token VARCHAR(64),
    created_at      TIMESTAMPTZ DEFAULT NOW(),
    INDEX idx_sale_user (sale_id, user_id)
);

Event Bus Design (Kafka)

Asynchronous event topics decouple the low latency purchase reservation lane from payment processing and the downstream Order Management System.

Topic: flash-order-events
  Partitions: 64
  Partition key: sale_id (preserves per-sale reservation ordering)
  Retention: 3 days (allows replaying failed order creation workflows)

Producer: Flash Sale Service after Redis Lua reservation success
  Event: { reservation_id, sale_id, user_id, sku_id, qty, admission_jwt_jti, timestamp }

Consumer groups:
  1. order-creator: idempotent INSERT into PostgreSQL flash_sale_orders
  2. payment: processes authorization charges within the 10-minute reservation window
  3. notification: dispatches confirmation email and push alerts on payment success

On payment failure or timeout: INCR stock back in Redis as a compensating action.
Sync path: Redis Lua reservation completes in under 50 ms, while order creation proceeds asynchronously via Kafka.
DLQ: flash-order-events-dlq, triggering alerts when consumer lag exceeds 30 seconds during active sales.

Fault Tolerance

ConcernSolution
Redis crash mid-saleRedis Cluster with WAIT 1 for replica acknowledgment, safety buffer (load 990/1000), and post-sale reconciliation.
OversellingLua script executes atomically with remaining stock verification and immediate undo via INCRBY. Failover reconciliation covers replica lag.
Payment timeoutReservation TTL set to 10 minutes. Expired reservations trigger a compensating stock increment.
Double purchaseIdempotency key enforcement, user_limit validation, and device validation execute in the atomic Lua path.
1M page loadsCDN pre-rendered static landing pages with zero origin load until the user clicks Buy.
Bot snipingMulti-layered defense with CAPTCHA, device fingerprinting, rate limiting, and minimum account age checks.
Queue fairnessRedis Sorted Set uses the monotonic queue position as its score to enforce deterministic FIFO ordering.

Specific: Redis Data Loss Mid-Sale

If a primary Redis node crashes, an asynchronous replica may miss the last 1 to 2 seconds of decrements, potentially allowing up to 10 to 20 units to be oversold.

Mitigation strategies include:

  1. WAIT command: Calling WAIT 1 0 immediately following the Lua script waits for at least 1 replica to acknowledge the write before the client is confirmed. The design assumes approximately 1 ms of additional latency for this acknowledgment path.
  2. Safety buffer: For a 1,000-unit promotional allocation, load only 990 units into Redis. Reserve the remaining 10 units as an internal buffer to absorb replication discrepancies.
  3. Post sale reconciliation: Tally confirmed orders in PostgreSQL against physical warehouse stock in the Inventory Management System. If confirmed orders exceed available stock, cancel the latest timestamped orders and issue immediate customer compensation.

Additional Considerations

Interview Walkthrough

  • 25-minute pacing strategy

    Prioritize core admission control and Lua atomicity before deep-diving into multi-region edge caches.

    • Contention on finite inventory: protect downstream services from a 1000x spike (5 min)
    • CDN static pre-rendered sale page before queue (6 min)
    • Virtual queue gates purchase requests with admission token (5 min)
    • Redis Lua atomic inventory decrement on buy (5 min)
    • Circuit breaker on payment and order services (4 min)
  • Frame the entire problem as a contention pattern on finite inventory where the primary goal is protecting downstream services from a 1000x traffic spike at T=0.
  • Lead with CDN and Edge Delivery: serve a static pre-rendered sale page so 1M page loads never touch origin servers.
  • Gate purchase requests through a virtual queue that admission-controls at a calculated drain rate using Back-of-the-Envelope Estimation.
  • Decrement inventory atomically in Redis via a Lua script, avoiding PostgreSQL queries for stock checks during the sale window.
  • Apply Circuit Breaker and Retries and Bulkheads on the checkout path so a slow payment provider does not cascade into total failure.
  • Plan for oversell risk: asynchronous replication lag can lose 1 to 2 seconds of decrements, so employ a safety buffer and post-sale reconciliation to resolve any excess reservations.
  • Reject early clicks with server-side start_time validation regardless of client clock skew.
  • Common pitfall: letting all 1M users hit the inventory API simultaneously without queue gating or edge caching.

Engineering Trade-offs

Throughput vs Fairness in Flash Sales

Flash sales balance fairness against throughput. A virtual waiting room absorbs traffic spikes while preventing origin saturation and reducing inventory contention.

Pre-Rendered Static Pages: Surviving the Traffic Spike

Serving static landing pages from CDN edge nodes absorbs millions of requests with zero origin load. Client-side JavaScript manages the countdown and reveals the Buy button at T=0. Only admitted purchase API calls reach backend clusters, preventing catastrophic server outages.

Countdown Precision and Clock Drift

The client fetches server time once on initial load, computes local clock drift, and adjusts the local countdown accordingly. The backend API validates against authoritative server clocks, rejecting any attempt submitted before sale.start_time.

Virtual Queue vs Direct Connection

No Queue:
  500K req/sec all hit API + Redis directly.
  Redis handles it. BUT: API servers, load balancers, network with
  500K concurrent connections --> likely crash.
  Result: errors, unfair (network latency determines winners).

Virtual Queue:
  All users enqueued. Admitted at a controlled rate, up to 10K/sec.
  The effective rate is tuned to stock availability and downstream capacity.
  API servers handle the admitted rate rather than the full 500K req/sec burst.
  Result: stable FIFO admission and controlled downstream load.
  Trade-off: users wait in the queue, and the sale can sell out before
  every queued user is admitted.

Why Not Rely Solely on Auto-Scaling?

Cloud auto-scaling typically requires 2 to 5 minutes to launch and initialize compute instances. Flash sales surge from 0 to 500K requests per second in less than 1 second, and available stock often depletes within 30 seconds. Pre-provisioning infrastructure 24 hours prior, combined with edge caching and queue admission, is essential.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...