Interview Setup
Interview Prompt
Design inventory management for e-commerce: track stock per SKU per warehouse, support reservations during checkout, prevent overselling, and handle flash-sale spikes.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Read vs write QPS? | Handling 200K stock checks/sec vs 10K reservations/sec requires aggressive caching on the read path. |
| Flash sale peak? | Handling 500K ops/sec on hot SKUs requires a Redis atomic execution path rather than PostgreSQL row locks. |
| Reservation TTL? | A 10-minute checkout hold with automated expiry releases uncommitted stock safely. |
| Oversell tolerance? | Zero tolerance for overselling, enforced via database CHECK constraints and atomic DECR operations. |
Scope
In scope
- Stock check API
- Reserve / confirm / release
- Warehouse-level inventory
- Flash sale mode
- Audit log
- Low-stock alerts
Out of scope (state explicitly)
- WMS pick and pack
- Supplier procurement
- Shopping cart lifecycle
Functional Requirements
Start by confirming what inventory guarantees the business requires. Inquire whether you must prevent overselling at all costs, how multi-warehouse allocation operates, and whether flash-sale traffic spikes are in scope. Inventory management integrates closely with the Shopping Cart System and the Order Management System across the commerce lifecycle.
- Track stock levels: Real-time quantity tracking per SKU across warehouses, stores, and channels
- Stock reservation: Temporarily hold inventory for pending orders (soft lock) before payment confirmation
- Stock decrement: Atomically reduce stock on confirmed purchase; prevent overselling
- Multi-warehouse: Track inventory per warehouse/fulfillment center with inter-warehouse transfers
- Replenishment alerts: Notify when stock falls below reorder point (low-stock threshold)
- Batch updates: Bulk ingest from suppliers, warehouse scans, POS systems
- Stock adjustments: Manual adjustments for damage, theft, counting discrepancies
- Inventory holds: Reserve stock for flash sales, bundles, pre-orders before they go live
- Multi-channel sync: Sync available stock across website, mobile app, marketplace (Amazon, eBay), and physical stores
- Audit trail: Immutable log of every stock change with reason, actor, and timestamp
Non-Functional Requirements
Interviewers focus heavily on consistency and overselling prevention in this design. Establish the two-phase reserve and confirm pattern early, as it decouples fast customer checkout from durable ledger commits. Review Distributed Transactions: 2PC vs Saga to structure the reservation coordination.
- Strong Consistency: Stock counts MUST be accurate: overselling is unacceptable
- Low Latency: Stock check in < 10 ms; reservation in < 50 ms
- High Throughput: Handle 100K+ stock operations/sec during flash sales
- Availability: 99.99%: inventory service is in the critical path for every checkout
- Durability: Stock data never lost; every mutation is logged
- Idempotency: Duplicate requests (network retries) must not double-decrement stock
- Scalability: 100M+ SKUs across 1,000+ warehouses
Capacity Estimations
Analyze data volumes before deciding on storage architecture. SKU counts and warehouse fan-out dictate PostgreSQL storage requirements, while peak flash-sale QPS justifies an in-memory Redis atomic path on top.
| Metric | Calculation | Value |
|---|---|---|
| Total SKUs | Given (assumption documented in value) | 100M |
| Warehouses / fulfillment centers | Given (assumption documented in value) | 1,000 |
| SKU-warehouse combinations | Given | ~500M (not every SKU in every warehouse) |
| Stock check queries / sec | From Stock check queries / day ÷ 86400 (+ peak factor in value) | 200K (product page views trigger stock check) |
| Reservation requests / sec | From Reservation requests / day ÷ 86400 (+ peak factor in value) | 10K (add to cart / checkout) |
| Confirmed decrements / sec | From Confirmed decrements / day ÷ 86400 (+ peak factor in value) | 5K (orders placed) |
| Flash sale peak | Given (peak load assumption) | 500K stock ops/sec for hot SKUs |
| Stock record size | Given | ~200 bytes |
| Total data | 500M x 200B | 100 GB |
Architecture Diagram
In the room: lead with the two-phase reserve-then-confirm checkout flow, because overselling is the critical failure mode interviewers probe deepest.
Walk through the read vs write split first. Product browse pages query a cached sellable count, whereas checkout transitions stock through explicit reserve and confirm phases in PostgreSQL. High-velocity flash events bypass database row locks entirely by using Redis atomic counters. Structuring this as two distinct lanes ensures browse traffic never contends with checkout locks.
Component Deep Dives
Stock Reservation Flow: Two-Phase Pattern
Two-Phase Reservation Lifecycle
The two-phase reservation flow forms the architectural backbone for consistency, ensuring that checkout reservations never result in overselling.
Normal catalog traffic relies on cached PostgreSQL sellable counts, whereas flash sale traffic shifts to Redis atomic decrements and reconciles back to the durable database ledger asynchronously after the event.
Event Bus Design (Kafka)
Stock mutations fan out asynchronously to ensure the reservation API maintains p99 latencies under 50ms, decoupling channel updates and audit streams from the critical checkout path.
Topic: stock-changes
Partitions: 128
Partition key: sku_id (preserves per-SKU mutation ordering)
Retention: 7 days (channel sync replay)
Replication factor: 3, min.insync.replicas: 2
Producer: Inventory Service after PostgreSQL commit (transactional outbox)
Event: { event_id, sku_id, warehouse_id, delta, new_qty, reason, timestamp }
Consumer groups:
1. channel-sync: push stock to Amazon/eBay/Shopify within seconds
2. analytics: ClickHouse inventory movement reporting
3. low-stock-alerts: sku below reorder point to trigger procurement notification
Topic: reservation-events (reserved, confirmed, released) consumed by Order Service saga
Topic: low-stock-alerts: notify procurement team
Sync path: GET /stock reads Redis cache-aside; reserve uses PostgreSQL + Redis DECR
Async path: channel sync and analytics never block stock check < 10ms
DLQ: stock-changes-dlq; alert when consumer lag > 30sFlash Sale Concurrency: Redis Counter
Flash Sale Concurrency: Redis Counter
When interviewers probe spike traffic, explain that PostgreSQL row locks collapse under concurrent transactions. We offload atomic decrements to Redis and reconcile back to the database later.
Problem: Flash sale of 1,000 units. 500K users hit "Buy" simultaneously.
PostgreSQL: SELECT ... FOR UPDATE → row-level lock → 500K waiting → timeout/crash
Solution: Use Redis as the fast path for stock decrements.
Pre-sale setup:
SET flash_stock:{sku_id} 1000
On purchase attempt:
local remaining = redis.call('DECR', 'flash_stock:SKU-FLASH-1')
if remaining >= 0 then
-- SUCCESS: user got one
-- Async: write to PostgreSQL, create order
return 'reserved'
else
-- SOLD OUT
redis.call('INCR', 'flash_stock:SKU-FLASH-1') -- undo the DECR
return 'sold_out'
end
Why this works:
Redis DECR is atomic (single-threaded) → no race conditions
500K DECR operations/sec → Redis handles easily (100K ops/sec per shard)
After Redis confirms → async write to PostgreSQL (no lock contention)Multi-Warehouse Stock Allocation
Multi-Warehouse Stock Allocation Strategies
Multi-warehouse allocation is a common interview probe. Clarify whether you optimize for shipping costs, delivery speed, or regional inventory balance. Amazon's hybrid scoring model demonstrates a strong senior design approach.
User orders SKU-123. It's available in 3 warehouses:
WH-NYC: 50 units (200 miles from user)
WH-CHI: 30 units (700 miles)
WH-LAX: 100 units (2500 miles)
Allocation strategies:
1. Nearest warehouse (minimize shipping cost + time):
Sort warehouses by distance to delivery address
Pick closest with stock → WH-NYC ✓
2. Load-balanced (prevent one warehouse from depleting):
Pick warehouse with highest stock level → WH-LAX
3. Cost-optimized (minimize total fulfillment cost):
cost = shipping_cost + handling_cost + last_mile_cost
Consider: shipping zone, carrier rates, warehouse labor cost
Pick minimum cost
4. Hybrid ⭐ (Amazon's approach):
score = w1 x (1 / distance) + w2 x stock_level + w3 x (1 / cost)
+ w4 x delivery_speed_guarantee
Pick highest score
Split shipment:
Order has 3 items: SKU-A (only in WH-NYC), SKU-B (only in WH-LAX), SKU-C (both)
→ Split into 2 shipments: {SKU-A, SKU-C} from WH-NYC, {SKU-B} from WH-LAX
→ Minimize number of shipments while respecting stock constraints
→ NP-hard problem at scale → greedy heuristic: maximize items per shipmentAPI Design
Inventory Endpoints
Check Stock
GET /api/v1/inventory/{sku_id}/stock?warehouse_id=WH-NYC
Response: 200 OK
{
"sku_id": "SKU-123",
"warehouse_id": "WH-NYC",
"available": 50,
"reserved": 5,
"sellable": 45,
"low_stock": false,
"reorder_point": 10
}Reserve Stock
POST /api/v1/inventory/reserve
Idempotency-Key: "order-uuid-abc"
{
"items": [
{"sku_id": "SKU-123", "quantity": 2, "warehouse_id": "WH-NYC"}
],
"ttl_seconds": 600,
"user_id": "user-uuid"
}
Response: 200 OK
{
"reservation_id": "res-uuid",
"status": "reserved",
"expires_at": "2026-03-14T11:10:00Z",
"items": [{"sku_id": "SKU-123", "reserved_qty": 2, "warehouse_id": "WH-NYC"}]
}Confirm Reservation
POST /api/v1/inventory/confirm
{
"reservation_id": "res-uuid",
"order_id": "order-uuid"
}
Response: 200 OK
{ "status": "confirmed" }Bulk Stock Update
POST /api/v1/inventory/bulk-update
{
"warehouse_id": "WH-NYC",
"updates": [
{"sku_id": "SKU-123", "available": 200, "reason": "supplier_shipment"},
{"sku_id": "SKU-456", "available": 0, "reason": "discontinued"}
]
}
Response: 200 OK
{ "updated": 2, "failed": 0 }Common Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
PostgreSQL: Source of Truth (Sharded by sku_id)
CREATE TABLE inventory (
sku_id VARCHAR(50) NOT NULL,
warehouse_id VARCHAR(50) NOT NULL,
available INT NOT NULL DEFAULT 0 CHECK (available >= 0),
reserved INT NOT NULL DEFAULT 0 CHECK (reserved >= 0),
reorder_point INT DEFAULT 10,
max_stock INT DEFAULT 1000,
last_replenished TIMESTAMP,
updated_at TIMESTAMP DEFAULT NOW(),
PRIMARY KEY (sku_id, warehouse_id),
CHECK (reserved <= available)
);
CREATE TABLE reservations (
reservation_id UUID PRIMARY KEY,
sku_id VARCHAR(50) NOT NULL,
warehouse_id VARCHAR(50) NOT NULL,
user_id UUID NOT NULL,
order_id UUID,
quantity INT NOT NULL,
status ENUM('active', 'confirmed', 'released', 'expired') DEFAULT 'active',
expires_at TIMESTAMPTZ NOT NULL,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_status_expires (status, expires_at) WHERE status = 'active',
INDEX idx_sku (sku_id, warehouse_id)
);
CREATE TABLE stock_audit_log (
log_id BIGSERIAL PRIMARY KEY,
sku_id VARCHAR(50) NOT NULL,
warehouse_id VARCHAR(50) NOT NULL,
change_type ENUM('reserve', 'confirm', 'release', 'adjust', 'replenish', 'return'),
quantity_change INT NOT NULL,
previous_available INT,
new_available INT,
reason TEXT,
actor_id VARCHAR(50),
idempotency_key VARCHAR(64),
created_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_sku_time (sku_id, created_at DESC)
);Redis: Hot Path Cache and Flash Sale
# Sellable stock cache (read path)
stock:{sku_id}:{warehouse_id} → INT (sellable = available - reserved)
TTL: 60 seconds (refreshed from PostgreSQL)
# Flash sale atomic counters
flash_stock:{sale_id}:{sku_id} → INT
No TTL (managed explicitly)
# Reservation TTL keys
reservation:{reservation_id} → JSON { sku_id, warehouse_id, quantity, user_id }
TTL: 600 (10 minutes)
# Idempotency (prevent duplicate reserve/confirm)
idempotency:{key} → response JSON
TTL: 86400Kafka Topics
Topic: stock-changes (every mutation → consumed by channel sync, analytics, alerts) Topic: reservation-events (reserved, confirmed, released → consumed by order service) Topic: low-stock-alerts (sku dropped below reorder point → notify procurement)
Fault Tolerance
| Concern | Solution |
|---|---|
| Overselling | Reservation pattern + PostgreSQL CHECK constraint + Redis DECR atomic |
| Reservation leak (never confirmed/released) | TTL-based expiry + cron cleanup job every 1 min |
| Redis ↔ PostgreSQL divergence |
|
| Double decrement (retry) | Idempotency key on every reserve/confirm request |
| Warehouse system offline |
|
| Flash sale thundering herd |
|
| Database failover |
|
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50 and arch75 depth unless interviewing for a staff-level role.
- Reserve, confirm, and release lifecycle (5 min)
- PostgreSQL source of truth with row-level locks and check constraints (6 min)
- Hot-SKU execution path: Redis atomic counters for flash sales (5 min)
- Multi-channel synchronization via Kafka stock-change events (5 min)
- Safety buffer strategies: pausing channels when stock drops below 5 (4 min)
- Start with the reserve, confirm, and release lifecycle, ensuring inventory never drops negative and checkout remains atomic.
- Explain PostgreSQL as the source of truth using
SELECT ... FOR UPDATEand CHECK constraints across available and reserved columns. - Cover the hot-SKU path where Redis fronts flash-sale writes with Lua atomic decrements, updating PostgreSQL asynchronously for durable auditability.
- Walk through multi-channel synchronization, pushing stock-change events through Kafka to Amazon, eBay, and Shopify within seconds instead of relying on periodic polling.
- Mention safety buffers where stock below 5 sets channel availability to 0, because phantom out-of-stock states are preferable to overselling during propagation lag.
- Discuss sharding by SKU hash so that high-velocity products map to dedicated database partitions rather than contending on a single table partition.
- Avoid the common pitfall of checking stock only during cart addition without re-validating at checkout, which allows concurrent users to claim the last unit.
Engineering Trade-offs
PostgreSQL vs DynamoDB for Inventory
Inventory correctness fundamentally trades optimistic concurrency control against checkout latency, contrasting PostgreSQL relational locks with Redis atomic counters.
PostgreSQL ⭐:
✓ ACID transactions: critical for reserve → confirm atomicity
✓ CHECK constraints: available >= 0, reserved <= available
✓ SELECT ... FOR UPDATE: row-level locking for safe concurrent updates
✓ Rich queries for analytics: "total stock across all warehouses"
✗ Single-writer bottleneck for hot SKUs (millions of updates to same row)
Scaling: shard by sku_id hash → each shard handles subset of SKUs
Hot SKU (flash sale): Redis fronts the hot path; PostgreSQL is async
DynamoDB:
✓ Horizontal scalability (no sharding management)
✓ Conditional writes: UpdateExpression with ConditionExpression
SET available = available - 1 WHERE available >= 1
✓ Auto-scaling for burst traffic
✗ No multi-row transactions (25-item limit)
✗ Eventual consistency by default (need strongly consistent reads)
✗ More expensive at high write throughput
Decision: PostgreSQL for most e-commerce (ACID is critical).
DynamoDB if: AWS-native, massive scale (100M+ SKUs), can tolerate 25-item txn limit.
Both: Redis in front for read-heavy + flash-sale write-heavy paths.Inventory: Push vs Pull for Channel Sync
Push (event-driven) ⭐:
Stock changes → Kafka event → Channel Sync Service → update Amazon/eBay/Shopify
Latency: 1-5 seconds from stock change to channel update
✓ Near real-time; prevents overselling across channels
✗ Requires reliable event delivery (Kafka at-least-once)
Pull (polling):
Each channel polls our API every 30 seconds for stock updates
✓ Simple; channel controls frequency
✗ 30-second stale window → overselling risk
✗ Wasteful: 90% of polls return "no change"
Hybrid:
Push for critical changes (stock < 10 → urgent)
Pull as fallback reconciliation (every 5 minutes)
For safety buffer: when stock < 5, set channel stock = 0
Prevents overselling during push delivery lag
"Phantom out-of-stock" is better than oversellingReview
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.