Interview Setup
Interview Prompt
Design a flash sale system where 1M users compete for 1,000 items in seconds, ensuring strict FIFO ordering, zero inventory oversell, bot protection, and post-reservation payment authorization.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Allow all 500K RPS to hit Redis directly? | Redis can handle the raw rate, but application servers and load balancers cannot, making a virtual queue essential. |
| Reserve stock before or after payment? | Reserve stock first with a 10-minute payment window. Release stock with an increment when payment expires or fails. |
| Per-user purchase limit policy? | Lua scripts enforce a maximum of 2 units per user ID and unique device fingerprint. |
| Expected bot traffic ratio? | Anticipate 30% to 50% automated bot traffic at T=0, mitigated by CAPTCHA, TLS fingerprinting, and account age validation. |
Scope
In scope
- Virtual waiting room
- Redis atomic stock decrement
- Admission JWT tokens
- Anti-bot security layers
- Payment reservation lifecycle
- Real-time sold-out broadcasts
Out of scope (state explicitly)
- Full payment gateway processing
- Product catalog management
- Warehouse inventory management
Functional Requirements
Confirm the traffic spike profile with your interviewer before proposing components. Inquire about concurrent users at T=0, purchase limits per buyer, and whether a virtual waiting room or anti-bot defense falls within scope.
- Scheduled sale lifecycle: Administrators schedule flash promotions with precise start and end times, assigned SKU inventory, and steep promotional discounts.
- Synchronized countdown timer: Display a synchronized countdown on client interfaces, revealing product items and purchase buttons at the exact start timestamp.
- Atomic stock reservation: Decrement inventory atomically and reserve units for a dedicated payment window, as contrasted with standard cart flows in the Shopping Cart System.
- Virtual waiting room: Enqueue surplus traffic into a structured queue when concurrent volume exceeds cluster processing capacity.
- Purchase limits per customer: Enforce a strict cap of 1 to 2 units per user and verified device fingerprint to block scalpers.
- Real-time stock indicators: Stream remaining available inventory counts to active users and broadcast immediate sell-out notifications.
- Strict admission fairness: Provide first-come, first-served admission where refreshing browsers provides zero competitive advantage.
- Anti-bot defenses: Filter out automated scrapers and headless checkout bots from sniping available inventory before real users can purchase.
Non-Functional Requirements
Interviewers prioritize zero inventory overselling and resilient stability under sudden million-user traffic bursts. Establish the critical invariant upfront: if 1,000 units are allocated, exactly 1,000 successful orders can be confirmed.
- Extreme Throughput: Sustain 1M+ concurrent active connections and 500K purchase attempts per second at sale kickoff.
- Low Latency: Return atomic purchase decisions in under 100 ms p99.
- Strong Consistency: Maintain strict inventory accuracy with zero overselling, respecting the principles in CAP Theorem and Consistency Models.
- Graceful Degradation: Protect downstream payment and order management services from cascade failures via admission gating.
- Guaranteed Fairness: Enforce verifiable first-in, first-out (FIFO) queue ordering using synchronized timestamp rankings.
- Idempotency: Ensure rapid double-clicks or client network retries never generate duplicate inventory reservations.
Capacity Estimations
Calculate capacity constraints before defending Redis as the primary reservation hot path. Peak purchase requests and concurrent user volume determine whether a virtual queue is mandatory above the atomic decrement layer.
| Metric | Calculation | Value |
|---|---|---|
| Concurrent users at sale start | Peak assumption | 1M+ |
| Purchase attempts / sec (T=0) | Peak assumption | 500K |
| Items for sale | Given | 1,000 - 10,000 units |
| Time to sell out | Given | 5-30 seconds |
| Page load requests / sec | Peak refresh-storm assumption | 2M (pre-sale refresh storm) |
| Bot traffic ratio | Given | 30-50% of requests |
Architecture Diagram
In the room: draw the virtual queue before the Redis Lua purchase path, because admission control separates browse traffic from the atomic buy lane.
Trace the admission funnel from edge distribution to order completion. Pre-rendered static landing pages live on global CDN caches, ensuring millions of pre-sale refreshes bypass origin servers entirely. Waiting customers enter a virtual queue, receive short-lived admission tokens upon dequeue, and only then access the Redis Lua purchase path, keeping browse traffic strictly decoupled from the atomic buy lane.
Component Deep Dives
The Critical Purchase Path: Redis Lua Script
This is the core implementation detail interviewers probe most aggressively. The entire reservation decision must execute atomically in Redis in under 1 ms, eliminating the race conditions inherent in separate read, decrement, and write commands. The script also makes retries idempotent and checks the per user and per device limits in the same atomic path. For broader patterns, explore Redis Patterns.
-- Keys:
-- 1: flash_stock:{sale_id}:{sku_id}
-- 2: user_limit:{sale_id}:{user_id}
-- 3: idempotency:{sale_id}:{idempotency_key}
-- 4: reservation:{reservation_token}
-- 5: device:{sale_id}:{fingerprint}
-- Args: user_id, sku_id, quantity, reservation_token, idempotency_key, device_fingerprint
-- Step 0: Return the original result for a retried request
local previous = redis.call('GET', KEYS[3])
if previous then
return {2, previous}
end
-- Step 1: Check per-user purchase limit
local user_purchased = redis.call('GET', KEYS[2])
local quantity = tonumber(ARGV[3])
if user_purchased and tonumber(user_purchased) + quantity > 2 then
return {0, 'LIMIT_EXCEEDED'}
end
-- Step 2: Check per-device purchase ownership
local device_user = redis.call('GET', KEYS[5])
if device_user and device_user ~= ARGV[1] then
return {0, 'DEVICE_LIMIT_EXCEEDED'}
end
-- Step 3: Atomic stock decrement
local remaining = redis.call('DECRBY', KEYS[1], quantity)
if remaining < 0 then
redis.call('INCRBY', KEYS[1], quantity) -- undo
return {0, 'SOLD_OUT'}
end
-- Step 4: Record purchase limits and reservation
redis.call('INCRBY', KEYS[2], quantity)
redis.call('EXPIRE', KEYS[2], 86400)
redis.call('SET', KEYS[5], ARGV[1], 'EX', 86400)
redis.call('SET', KEYS[4], cjson.encode({
user_id = ARGV[1],
sku_id = ARGV[2],
qty = quantity
}), 'EX', 600)
local result = cjson.encode({
reservation_token = ARGV[4],
remaining_stock = remaining
})
redis.call('SET', KEYS[3], result, 'EX', 86400)
return {1, result}Why execute a Lua script rather than individual commands? Redis runs Lua scripts atomically without interleaving other commands. Separate read, decrement, and check calls expose a window where concurrent requests can oversell. A single Redis node can comfortably serialize 100K+ Lua executions per second, and sharding via Sharding and Partitioning scales this further across independent sales.
Virtual Queue: Handling 1M Concurrent Users
Even if Redis sustains high decrement throughput, the upstream API gateway and load balancer tier can collapse under a 500K request per second burst. A virtual waiting room regulates inbound pressure by admitting customers at a stable, controlled drain rate that can be tuned to inventory and downstream capacity.
T-5 min: Users "Enter Queue" early
Position assigned: INCR queue_position:{sale_id} --> position 347,231
User shown: "Your position: 347,231. Estimated wait: ~5 minutes"
T=0: Sale starts. Queue processes users FIFO.
Gate rate: 10,000 users admitted per second (tunable)
Admitted users:
1. Receive short-lived JWT token (valid 60 seconds)
2. Token authorizes call to purchase API
3. Purchase API validates token --> runs Lua script on Redis
Users not yet admitted:
- See "Please wait..." with live position via WebSocket/SSE
- Position updates every 5 seconds
Stock gone:
- Broadcast SOLD_OUT to ALL remaining queue members immediately
- Don't make users wait if nothing left to buy
Queue implementation:
Redis sorted set: ZADD queue:{sale_id} {timestamp} {user_id}
Processing: ZPOPMIN queue:{sale_id} 10000 (pop 10K per second)Anti-Bot Measures
Scalpers and automated purchase bots represent a primary operational hazard. Defend against unauthorized syndicates by deploying layered inspection filters from the network edge to the checkout service.
Layer 1: CDN/WAF (Cloudflare, AWS WAF) - Rate limit per IP: max 10 req/sec - Known bot signatures blocked - JavaScript challenge (bots can't execute JS) - TLS fingerprinting (JA3 hash) flag non-browser clients Layer 2: Queue Entry Validation - CAPTCHA at queue entry (invisible reCAPTCHA) - Device fingerprint (canvas hash, WebGL, screen resolution) - Account age check: accounts < 24 hours old are blocked Layer 3: Purchase Validation - One purchase per user_id (Redis user_limit) - One purchase per device_fingerprint - One purchase per payment method Layer 4: Post-Purchase Fraud Detection - Multiple orders to same shipping address from different accounts are cancelled - Reseller pattern detection flags suspicious activity
API Design
API payloads use explicit domain types so request and response contracts remain stable as the implementation evolves.
type EnterQueueRequest = {
captcha_token: string;
device_fingerprint: string;
};
type EnterQueueResponse = {
queue_position: number;
estimated_wait_seconds: number;
queue_token: string;
};
type PurchaseRequest = {
sku_id: string;
quantity: number;
};
type PurchaseResponse = {
status: "reserved" | "sold_out" | "limit_exceeded" | "device_limit_exceeded";
reservation_token?: string;
payment_deadline?: string;
remaining_stock?: number;
};
type SaleStatusItem = {
sku_id: string;
name: string;
flash_price: number;
original_price: number;
total_stock: number;
remaining: number;
};
type SaleStatusResponse = {
sale_id: string;
status: "scheduled" | "active" | "ended" | "cancelled";
items: SaleStatusItem[];
};Enter Queue
Clients invoke this endpoint to receive a queue placement and initial wait estimate. The request maps to EnterQueueRequest and the response maps to EnterQueueResponse.
POST /api/v1/flash-sale/{sale_id}/enter-queue
Content-Type: application/json
{
"captcha_token": "recaptcha-response-token",
"device_fingerprint": "fp-hash-abc"
}
Response: 200 OK
Content-Type: application/json
{
"queue_position": 12345,
"estimated_wait_seconds": 120,
"queue_token": "qt-uuid"
}Purchase (After Admitted)
Admitted clients supply their cryptographic admission token along with an idempotency key to claim inventory. The request and response map to PurchaseRequest and PurchaseResponse.
POST /api/v1/flash-sale/{sale_id}/purchase
Content-Type: application/json
Idempotency-Key: "purchase-user123-sale456"
Authorization: Bearer {admission_jwt}
{
"sku_id": "SKU-FLASH-1",
"quantity": 1
}
Response: 200 OK
Content-Type: application/json
{
"status": "reserved",
"reservation_token": "res-uuid",
"payment_deadline": "2026-03-14T11:10:00Z",
"remaining_stock": 423
}
OR { "status": "sold_out" }
OR { "status": "limit_exceeded" }
OR { "status": "device_limit_exceeded" }Get Sale Status
Clients can poll this endpoint or receive equivalent updates over WebSockets to display live stock remaining and event status. The response maps to SaleStatusResponse.
GET /api/v1/flash-sale/{sale_id}/status
Response: 200 OK
Content-Type: application/json
{
"sale_id": "sale-456",
"status": "active",
"items": [
{"sku_id": "SKU-FLASH-1", "name": "iPhone 16", "flash_price": 499.00,
"original_price": 999.00, "total_stock": 1000, "remaining": 423}
]
}Common Error Responses
Structured error representations for rejected admissions, rate limits, and sold-out states.
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 402 Payment Required: account balance or payment method has insufficient funds 502 Bad Gateway: payment gateway provider timeout, poll transaction status endpoint
Data Model
Redis: Flash Sale State
In memory keys maintain stock counters, user purchase limits, reservation holds, admission state, device ownership, and idempotency records for high speed evaluation.
flash_stock:{sale_id}:{sku_id} --> INT (atomic DECR)
user_limit:{sale_id}:{user_id} --> INT (max 2), TTL 86400
reservation:{token} --> JSON { user_id, sku_id, qty }, TTL 600
queue_position:{sale_id} --> INT (INCR for each entrant)
queue:{sale_id} --> Sorted Set { user_id: monotonic position }
admission:{sale_id}:{user_id} --> "admitted", TTL 60
device:{sale_id}:{fingerprint} --> user_id, TTL 86400
idempotency:{sale_id}:{key} --> JSON reservation result, TTL 86400PostgreSQL: Durable Records
Relational tables provide durable audit trails and order histories populated asynchronously after successful Redis reservation holds.
CREATE TABLE flash_sales (
sale_id UUID PRIMARY KEY,
name VARCHAR(255),
start_time TIMESTAMPTZ NOT NULL,
end_time TIMESTAMPTZ NOT NULL,
status ENUM('scheduled','active','ended','cancelled'),
created_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE TABLE flash_sale_items (
sale_id UUID NOT NULL,
sku_id VARCHAR(50) NOT NULL,
flash_price DECIMAL(10,2) NOT NULL,
original_price DECIMAL(10,2) NOT NULL,
total_stock INT NOT NULL,
sold_count INT DEFAULT 0,
PRIMARY KEY (sale_id, sku_id)
);
CREATE TABLE flash_sale_orders (
order_id UUID PRIMARY KEY,
sale_id UUID NOT NULL,
user_id UUID NOT NULL,
sku_id VARCHAR(50) NOT NULL,
quantity INT NOT NULL,
price DECIMAL(10,2),
status ENUM('reserved','paid','cancelled','expired'),
reservation_token VARCHAR(64),
created_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_sale_user (sale_id, user_id)
);Event Bus Design (Kafka)
Asynchronous event topics decouple the low latency purchase reservation lane from payment processing and the downstream Order Management System.
Topic: flash-order-events
Partitions: 64
Partition key: sale_id (preserves per-sale reservation ordering)
Retention: 3 days (allows replaying failed order creation workflows)
Producer: Flash Sale Service after Redis Lua reservation success
Event: { reservation_id, sale_id, user_id, sku_id, qty, admission_jwt_jti, timestamp }
Consumer groups:
1. order-creator: idempotent INSERT into PostgreSQL flash_sale_orders
2. payment: processes authorization charges within the 10-minute reservation window
3. notification: dispatches confirmation email and push alerts on payment success
On payment failure or timeout: INCR stock back in Redis as a compensating action.
Sync path: Redis Lua reservation completes in under 50 ms, while order creation proceeds asynchronously via Kafka.
DLQ: flash-order-events-dlq, triggering alerts when consumer lag exceeds 30 seconds during active sales.Fault Tolerance
| Concern | Solution |
|---|---|
| Redis crash mid-sale | Redis Cluster with WAIT 1 for replica acknowledgment, safety buffer (load 990/1000), and post-sale reconciliation. |
| Overselling | Lua script executes atomically with remaining stock verification and immediate undo via INCRBY. Failover reconciliation covers replica lag. |
| Payment timeout | Reservation TTL set to 10 minutes. Expired reservations trigger a compensating stock increment. |
| Double purchase | Idempotency key enforcement, user_limit validation, and device validation execute in the atomic Lua path. |
| 1M page loads | CDN pre-rendered static landing pages with zero origin load until the user clicks Buy. |
| Bot sniping | Multi-layered defense with CAPTCHA, device fingerprinting, rate limiting, and minimum account age checks. |
| Queue fairness | Redis Sorted Set uses the monotonic queue position as its score to enforce deterministic FIFO ordering. |
Specific: Redis Data Loss Mid-Sale
If a primary Redis node crashes, an asynchronous replica may miss the last 1 to 2 seconds of decrements, potentially allowing up to 10 to 20 units to be oversold.
Mitigation strategies include:
- WAIT command: Calling
WAIT 1 0immediately following the Lua script waits for at least 1 replica to acknowledge the write before the client is confirmed. The design assumes approximately 1 ms of additional latency for this acknowledgment path. - Safety buffer: For a 1,000-unit promotional allocation, load only 990 units into Redis. Reserve the remaining 10 units as an internal buffer to absorb replication discrepancies.
- Post sale reconciliation: Tally confirmed orders in PostgreSQL against physical warehouse stock in the Inventory Management System. If confirmed orders exceed available stock, cancel the latest timestamped orders and issue immediate customer compensation.
Additional Considerations
Interview Walkthrough
- 25-minute pacing strategy
Prioritize core admission control and Lua atomicity before deep-diving into multi-region edge caches.
- Contention on finite inventory: protect downstream services from a 1000x spike (5 min)
- CDN static pre-rendered sale page before queue (6 min)
- Virtual queue gates purchase requests with admission token (5 min)
- Redis Lua atomic inventory decrement on buy (5 min)
- Circuit breaker on payment and order services (4 min)
- Frame the entire problem as a contention pattern on finite inventory where the primary goal is protecting downstream services from a 1000x traffic spike at T=0.
- Lead with CDN and Edge Delivery: serve a static pre-rendered sale page so 1M page loads never touch origin servers.
- Gate purchase requests through a virtual queue that admission-controls at a calculated drain rate using Back-of-the-Envelope Estimation.
- Decrement inventory atomically in Redis via a Lua script, avoiding PostgreSQL queries for stock checks during the sale window.
- Apply Circuit Breaker and Retries and Bulkheads on the checkout path so a slow payment provider does not cascade into total failure.
- Plan for oversell risk: asynchronous replication lag can lose 1 to 2 seconds of decrements, so employ a safety buffer and post-sale reconciliation to resolve any excess reservations.
- Reject early clicks with server-side
start_timevalidation regardless of client clock skew. - Common pitfall: letting all 1M users hit the inventory API simultaneously without queue gating or edge caching.
Engineering Trade-offs
Throughput vs Fairness in Flash Sales
Flash sales balance fairness against throughput. A virtual waiting room absorbs traffic spikes while preventing origin saturation and reducing inventory contention.
Pre-Rendered Static Pages: Surviving the Traffic Spike
Serving static landing pages from CDN edge nodes absorbs millions of requests with zero origin load. Client-side JavaScript manages the countdown and reveals the Buy button at T=0. Only admitted purchase API calls reach backend clusters, preventing catastrophic server outages.
Countdown Precision and Clock Drift
The client fetches server time once on initial load, computes local clock drift, and adjusts the local countdown accordingly. The backend API validates against authoritative server clocks, rejecting any attempt submitted before sale.start_time.
Virtual Queue vs Direct Connection
No Queue: 500K req/sec all hit API + Redis directly. Redis handles it. BUT: API servers, load balancers, network with 500K concurrent connections --> likely crash. Result: errors, unfair (network latency determines winners). Virtual Queue: All users enqueued. Admitted at a controlled rate, up to 10K/sec. The effective rate is tuned to stock availability and downstream capacity. API servers handle the admitted rate rather than the full 500K req/sec burst. Result: stable FIFO admission and controlled downstream load. Trade-off: users wait in the queue, and the sale can sell out before every queued user is admitted.
Why Not Rely Solely on Auto-Scaling?
Cloud auto-scaling typically requires 2 to 5 minutes to launch and initialize compute instances. Flash sales surge from 0 to 500K requests per second in less than 1 second, and available stock often depletes within 30 seconds. Pre-provisioning infrastructure 24 hours prior, combined with edge caching and queue admission, is essential.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.