Interview Setup
Interview Prompt
Design an ecommerce platform like Amazon. Users browse a product catalog, search with filters, add items to cart, checkout, and pay. Support inventory management so overselling does not occur.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Catalog size and peak search QPS? | Evaluating 500M SKUs with 200K peak search QPS determines whether a distributed search index is required instead of SQL full text search. |
| Inventory model: single warehouse or multi warehouse distribution? | Multi warehouse distribution adds reservations per fulfillment center and ship from nearest routing logic during checkout. |
| Cart reservation TTL: how long should the system hold reserved inventory? | Use a 15 minute reservation window as the planning assumption. Shorter windows can increase checkout abandonment, while longer windows hold scarce inventory for too long during high demand sales. |
| Payment processing: authorize during checkout or capture upon shipment? | Use authorization during checkout and capture after shipment as the planning assumption. This directly shapes order state transitions and the compensation flow for failed or cancelled orders. |
Scope
In scope
- Product catalog and category hierarchy
- Search with facets and filters
- Shopping cart with inventory reservation
- Order state machine covering creation, payment, fulfillment, and shipping
- Payment orchestration with failure compensation
Out of scope (state explicitly)
- Recommendation engine implementation
- Review and rating system implementation
- Warehouse management and internal fulfillment systems
- Detailed fraud detection and scoring, assuming basic rule checks exist
Functional Requirements
When discussing requirements with your interviewer, clarify catalog browsing, cart operations, checkout consistency, and inventory boundaries first. Flash sale concurrency and multi seller marketplace support serve as common senior follow-ups.
- Product Catalog: Browse, search, and filter products with rich details, images, specifications, and seller metadata
- Shopping Cart: Add, modify, and remove items with persistent cross session synchronization
- Checkout and Payment: Place orders using multiple payment gateways with idempotent authorization and capture flows
- Order Management: Track complete order lifecycles through placement, payment, fulfillment, shipping, and delivery
- Inventory Management: Real time stock tracking with reservation guarantees to prevent overselling
- Search and Discovery: Full text search with faceted filtering across categories, price brackets, brands, and customer ratings
- Personalized Recommendations: Display co-purchase suggestions such as frequently bought together items and related categories, while recommendation generation remains a separate system
- Reviews and Ratings: Display verified purchaser reviews and score distributions, while review creation and moderation remain a separate system
- Seller Management: Seller onboarding, SKU catalog management, inventory allocation, and sales performance dashboards
- Notifications: Multi channel delivery for order confirmations, shipping updates, and delivery alerts
Non-Functional Requirements
Interviewers prioritize checkout data consistency and inventory correctness on this problem. Catalog read scalability is critical, but robust write path invariants differentiate senior candidates.
- High Availability: 99.99% uptime for catalog browsing and cart access because downtime directly degrades gross merchandise value
- Low Latency: Search query responses under 200 milliseconds and product page loads under 500 milliseconds at p99
- Massive Scalability: Support 500M SKUs, 100M daily active users, and 10M orders per day
- Strong Consistency: Inventory reservation and decrement require strict transactional correctness. Payment processing uses idempotent provider calls, durable state transitions, and Saga compensation because the external gateway is not part of the database transaction
- Eventual Consistency: Search index updates, review aggregations, and recommendation models tolerate asynchronous synchronization
- Flash Sale Resilience: Absorb 100x traffic spikes and peak loads of 10K orders per second without stock race conditions
- Idempotent Operations: Use idempotency keys and durable state checks so network retries do not create duplicate charges or duplicate inventory reservations
Capacity Estimations
Establishing capacity math before defining the inventory architecture is essential. Product catalog volume and peak order write concurrency during flash sales dictate whether optimistic locking or distributed in memory counters are mandatory.
| Metric | Calculation | Value |
|---|---|---|
| SKUs | Given | 500M |
| DAU | Given | 100M |
| Search QPS | Given | 200K |
| Orders / day | Given | 10M |
| Orders / sec | 10M ÷ 86,400 | ~116 avg, ~10K peak during flash sales |
| Product page views / day | 100M DAU x ~50 views | 5B |
| Cart operations / day | 100M DAU x ~5 ops | 500M |
| Avg order value | Given | $50 |
| Daily GMV | Given | $500M |
Architecture Diagram
In an interview setting, separate read heavy catalog browsing from write intensive, consistency critical checkout before drawing individual microservice boundaries.
The architecture divides cleanly into read optimized discovery paths and strongly consistent transaction paths. Product catalog data is served through global CDNs, Redis caches, and Elasticsearch clusters. Cart state resides in low latency Redis session stores until checkout promotion. The checkout workflow coordinates through an Order Service that reserves inventory transactionally and authorizes payment through a Saga workflow. Command Query Responsibility Segregation (CQRS) ensures catalog search performance remains insulated from high volume order write spikes, while fulfillment and notification services operate asynchronously over Kafka.
Component Deep Dives
Product Catalog Service
Catalog browsing, shopping cart workflows, and order fulfillment represent isolated architectural paths with fundamentally distinct latency and consistency profiles. The Product Catalog Service manages the authoritative product catalog, category hierarchies, and SKU variant specifications.
- Document Data Model: Uses MongoDB or DynamoDB because product categories have highly heterogeneous schemas where apparel requires size and color, electronics requires technical specifications, and books require ISBN and author data.
- Multi Tier Caching: Caches hot product records in Redis to serve product detail pages accessed thousands of times per minute.
- Entity Schema: Encapsulates title, description, CDN image URLs, base pricing, seller identifiers, dynamic attribute key-values, and nested category paths.
- Seller Association: Maintains seller profiles and seller to SKU relationships in the account and catalog domains. Detailed seller onboarding workflows stay outside the checkout path.
- Seller Offers: Stores seller specific price, availability, and fulfillment attributes separately from the canonical product record, so checkout can resolve the selected offer and verify its inventory. The selected seller and offer are carried into the cart and order records.
Search Service (Elasticsearch)
Product search synchronizes from the primary database via Change Data Capture to Elasticsearch, ensuring that browse traffic never queries transactional databases directly.
- Inverted Index Structure: Indexes product titles, descriptions, brand names, category hierarchies, and search keywords.
- Faceted Aggregations: Generates multi dimensional facet counts across category trees, price ranges, customer review ratings, stock availability, and verified brands.
- Composite Ranking: Combines BM25 text relevance with sales velocity signals, average rating scores, and sponsored product boosts. Seller offer selection can be resolved after product retrieval without duplicating the canonical product document for every seller.
- Typeahead Autocomplete: Powers prefix and fuzzy matching on product titles and brand names through completion suggesters.
Cart Service
Manages user shopping carts in low latency Redis memory. Cart state remains session scoped until checkout creates durable order records.
- Authenticated Sessions: Stores active carts in Redis hashes keyed by
cart:{user_id}with a 30-day sliding TTL and Append Only File persistence. - Guest Sessions: Persists unauthenticated carts in browser local storage and merges them into Redis upon account login.
- Price Revalidation: Captures price at the time an item is added to the cart, but always re-verifies current authoritative catalog prices during checkout initiation.
- Live Stock Verification: Checks stock availability when opening the cart and prompts the user if items have become scarce.
User Service
Maintains customer account profiles, saved shipping addresses, and payment references with aggressive caching on read heavy authorization paths.
- Relational Storage: Uses PostgreSQL for customer profile records, multiple address books, saved payment provider tokens, and account preferences.
- Authentication Boundary: An Authentication Service or identity provider issues signed JWT tokens. The API Gateway validates those tokens, while the User Service manages profile lifecycle and CRUD operations.
- Checkout Binding: Associates cart sessions and completed orders with authoritative user identifiers, orchestrating guest account merges seamlessly.
Inventory Service: Concurrency and Correctness
Enforces strict stock integrity and prevents overselling through atomic decrement operations and durable time bounded reservations.
Standard Operations: Pessimistic locking or atomic condition updates within relational database transactions:
BEGIN TRANSACTION;
SELECT available_quantity FROM inventory WHERE product_id = ? AND seller_id = ? FOR UPDATE;
UPDATE inventory SET available_quantity = available_quantity - ? WHERE product_id = ? AND seller_id = ?;
COMMIT;High Concurrency Flash Sales: Preloads available inventory counters into Redis memory. An atomic Lua check and decrement operation absorbs admission traffic without serializing every request on the primary database. The durable Inventory Service still confirms each admitted reservation against the primary database.
Pricing Service
Computes the authoritative checkout total from current catalog prices, promotions, taxes, and shipping rules.
- Price Quote: Creates a short lived quote with a version or timestamp so the order records the exact price used for authorization.
- Promotion Validation: Rechecks coupon eligibility, expiration, account limits, and minimum order values during checkout rather than trusting the cart snapshot.
- Order Snapshot: Persists unit prices, discounts, tax, shipping, and total on the order so later catalog price changes do not alter an existing order.
Order Service: Saga Based Orchestration
Coordinates distributed order checkout workflows across inventory, pricing, payment, and notification boundaries without two phase commit locks.
1. Validate Cart --> Cart Service (verify line items and quantities) 2. Calculate Total --> Pricing Service (apply promotions, taxes, and shipping fees) 3. Reserve Inventory --> Inventory Service (hold stock with 15 minute TTL) 4. Persist Order --> Order Service (create order record in PAYMENT_PENDING state) 5. Authorize Payment --> Payment Service (request authorization hold from gateway) 6. Commit Inventory --> Inventory Service (finalize the reservation and transition the order to PAYMENT_AUTHORIZED) 7. Start Fulfillment --> Order Service (emit the fulfillment event after durable authorization and inventory commit) 8. Capture Payment --> Payment Service (capture after shipment confirmation and transition the order to PAID) 9. Dispatch Alerts --> Notification Service (send confirmation email and SMS asynchronously) Compensating Actions: - If Order persistence fails after reservation: Release the inventory hold (Step 3) - If Payment authorization fails: Release the inventory hold and cancel the pending order - If inventory commit fails after payment authorization: Void the authorization, release the reservation if possible, and mark the order cancelled - If payment capture fails after shipment: Retry idempotently and reconcile with the payment provider before marking the order PAID
Payment Service
Integrates with external payment service providers and enforces strict idempotency keys to prevent duplicate charges during network retries. Provider webhooks update payment state asynchronously, and a reconciliation job polls stale payment intents when the provider response is ambiguous.
- Provider Abstraction: Integrates with external gateways such as Stripe, Adyen, and PayPal through unified client adapters.
- Idempotency Guarantees: Enforces a unique client supplied
idempotency_keyfor every charge attempt to ensure safe retries over intermittent networks. - Authorization and Capture: Places an authorization hold during checkout and captures funds only after the warehouse confirms shipment. Provider webhooks and reconciliation handle ambiguous or delayed payment outcomes.
Recommendation Service
Generates personalized product suggestions through offline co-purchase mining, serving precomputed recommendations directly from cache at query time.
- Collaborative Filtering: Runs batch Apache Spark jobs to mine historical purchase patterns and generate frequently bought together associations.
- Content Based Filtering: Correlates similar products by matching catalog attributes such as category taxonomy, brand affinity, and price range.
- Real Time Session Signals: Analyzes recent browsing events within active user sessions to dynamically adjust homepage recommendations.
- Low Latency Serving: Pre-computes top recommendation candidate lists and stores them in Redis to achieve p99 response times under 50 milliseconds.
Async Event Pipeline (Kafka)
Decouples synchronous checkout processing from downstream fulfillment systems while buffering high throughput traffic spikes during sales events.
- Core Event Streams: Partitions dedicated topics for
order-events,product-events,inventory-events, andclick-events. - Consumer Ecosystem: Coordinates independent consumers including Elasticsearch indexers, ClickHouse analytics pipelines, recommendation feature builders, fraud detection engines, and notification dispatchers.
- System Decoupling: Intermittent downstream outages in notification or analytics systems do not block customer order placement because the order state and outbox record are committed synchronously before asynchronous consumers process the event.
Event Bus Design (Kafka)
Topic configurations, partitioning strategies, and consumer group bindings govern asynchronous event propagation across the platform.
topics:
product_events:
partitions: 128
partition_key: product_id
ordering: "Guaranteed per product_id partition order"
retention: 7 days
producers:
- Catalog database change stream or CDC connector (create, update, delete)
consumers:
- Elasticsearch indexer
- Recommendation pipeline
- Catalog cache invalidator
order_events:
partitions: 64
partition_key: order_id
retention: 30 days
retention_purpose: "Fulfillment tracking and compliance replay"
producers:
- Order Service transactional outbox relay (state machine transitions)
consumers:
- Notification Service
- Analytics pipeline (ClickHouse)
- Fraud check engine
- Warehouse fulfillment trigger
deduplication:
idempotency_key: "composite(order_id, transition, version)"
inventory_events:
partitions: 64
partition_key: "product_id + seller_id"
ordering: "Guaranteed per inventory item partition order"
producers:
- Inventory Service transactional outbox relay (reserve, release, commit)
consumers:
- Redis hot-stock synchronization
- Elasticsearch in_stock facet updater
- Periodic reconciliation job
user_click_events:
partitions: 64
partition_key: user_id
producers:
- Cart Service
- Frontend telemetry beacons
consumers:
- Recommendation pipeline
- Data lake (S3)
cluster_configuration:
replication_factor: 3
min_insync_replicas: 2
producer_idempotence: true
dead_letter_queues:
- order-events-dlq
- product-events-dlq
max_retries_before_dlq: 3
execution_paths:
sync_path: "validate cart, calculate authoritative total, reserve inventory, persist order and outbox record, return HTTP 201"
async_path: "outbox relay publishes order-events after commit, followed by asynchronous fulfillment triggers, Elasticsearch indexing, notifications, analytics ingestion, and recommendation refresh"API Design
Client API Type Definitions
RESTful API contracts and TypeScript client interfaces for product search, catalog retrieval, cart management, and order checkout. The domain interfaces define product queries, faceted responses, cart states, and checkout order payloads.
// Domain models and client interfaces for product discovery, cart, and checkout
export interface ProductSummary {
productId: string;
title: string;
price: number;
currency: string;
discountPrice?: number;
rating: number;
reviewCount: number;
thumbnailUrl: string;
inStock: boolean;
}
export interface SearchProductsRequest {
query: string;
category?: string;
brand?: string;
priceMin?: number;
priceMax?: number;
sortBy?: "relevance" | "price_asc" | "price_desc" | "rating";
page?: number;
cursor?: string; // use for deep pagination at scale
pageSize?: number;
}
export interface SearchProductsResponse {
products: ProductSummary[];
totalMatches: number;
page: number;
nextCursor?: string | null;
facets: {
categories: Array<{ name: string; count: number }>;
brands: Array<{ name: string; count: number }>;
priceRanges: Array<{ range: string; count: number }>;
};
}
export interface SellerOffer {
sellerId: string;
price: number;
discountPrice?: number;
currency: string;
inStock: boolean;
fulfillmentType?: "platform" | "seller";
}
export interface ProductDetail extends ProductSummary {
description: string;
imageUrls: string[];
sellerId?: string;
offers?: SellerOffer[];
categoryPath: string[];
attributes: Record<string, string | string[]>;
}
export interface CartItem {
productId: string;
sellerId: string;
quantity: number;
addedPrice: number;
currentPrice: number;
}
export interface CartState {
userId: string;
items: CartItem[];
subtotal: number;
updatedAt: string;
}
export interface PlaceOrderRequest {
shippingAddressId: string;
paymentMethodId: string;
promoCode?: string;
idempotencyKey: string;
}
export type OrderStatus =
| "created"
| "payment_pending"
| "payment_authorized"
| "fulfilling"
| "shipped"
| "paid"
| "delivered"
| "cancelled"
| "refunded";
export interface OrderResponse {
orderId: string;
status: OrderStatus;
subtotal: number;
discount: number;
tax: number;
shipping: number;
total: number;
createdAt: string;
}
export interface ECommerceClient {
searchProducts(request: SearchProductsRequest): Promise<SearchProductsResponse>;
getProduct(productId: string): Promise<ProductDetail>;
getCart(): Promise<CartState>;
updateCartItem(productId: string, quantity: number): Promise<CartState>;
placeOrder(request: PlaceOrderRequest): Promise<OrderResponse>;
getOrder(orderId: string): Promise<OrderResponse>;
}Search Products
Executes full text query matching across inverted indices and computes real time facet aggregations.
GET /api/v1/products/search?q=wireless+headphones&category=electronics&price_min=20&price_max=200&sort=relevance&page=1 HTTP/1.1
Host: api.ecommerce.com
Authorization: Bearer <user_token>
HTTP/1.1 200 OK
Content-Type: application/json
{
"products": [
{
"product_id": "prod_102938",
"title": "Sony WH-1000XM5 Noise-Canceling Headphones",
"price": 349.99,
"discount_price": 279.99,
"currency": "USD",
"rating": 4.7,
"review_count": 12543,
"in_stock": true
}
],
"total_matches": 1420,
"page": 1,
"next_cursor": "eyJjYXRlZ29yeSI6ImVsZWN0cm9uaWNzIiwib2Zmc2V0IjoyMH0=",
"facets": {
"brands": [
{ "name": "Sony", "count": 42 },
{ "name": "Bose", "count": 31 }
],
"price_ranges": [
{ "range": "20-50", "count": 120 },
{ "range": "200-500", "count": 85 }
]
}
}Get Product
Fetches comprehensive product specifications, current pricing, and live inventory state.
GET /api/v1/products/prod_102938 HTTP/1.1
Host: api.ecommerce.com
HTTP/1.1 200 OK
Content-Type: application/json
{
"product_id": "prod_102938",
"title": "Sony WH-1000XM5",
"price": 349.99,
"discount_price": 279.99,
"currency": "USD",
"rating": 4.7,
"in_stock": true,
"attributes": {
"color": "Black",
"connectivity": "Bluetooth 5.2"
}
}Place Order
Initiates the checkout saga with a client generated idempotency key that is persisted with the order. Repeating the same request returns the existing order instead of creating a second checkout workflow:
POST /api/v1/orders HTTP/1.1
Host: api.ecommerce.com
Authorization: Bearer <user_token>
Idempotency-Key: 7b8c9d0e-1f2a-4b3c-9d8e-5a6b7c8d9e0f
Content-Type: application/json
{
"shipping_address_id": "addr-uuid",
"payment_method_id": "pm-uuid",
"promo_code": "SAVE10"
}
HTTP/1.1 201 Created
Content-Type: application/json
{
"order_id": "order-uuid",
"status": "payment_pending",
"total": 299.99
}Common Error Responses
Standardized HTTP error envelopes returned across API Gateway endpoints:
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
MongoDB: Product Catalog
Document store model accommodating dynamic attributes across diverse merchandise categories. The overall storage architecture uses MongoDB for catalog attributes, Redis for ephemeral cart sessions and hot caches, and PostgreSQL for ACID order, line item, inventory, and reservation transactions.
{
"_id": "product-uuid",
"title": "Sony WH-1000XM5 Headphones",
"category_path": ["Electronics", "Audio", "Headphones"],
"brand": "Sony",
"price": 349.99,
"discount_price": 279.99,
"attributes": {
"color": "Black",
"connectivity": "Bluetooth 5.2"
},
"rating": 4.7,
"review_count": 12543,
"offers": [
{
"seller_id": "sel_54321",
"price": 279.99,
"currency": "USD",
"in_stock": true
}
]
}PostgreSQL: Inventory & Reservations
Relational schema enforcing strict nonnegative quantity constraints and expiration timestamps for stock holds.
CREATE TABLE inventory (
product_id UUID NOT NULL,
seller_id UUID NOT NULL,
available_quantity INT NOT NULL CHECK (available_quantity >= 0),
PRIMARY KEY (product_id, seller_id)
);
CREATE TABLE reservations (
reservation_id UUID PRIMARY KEY,
order_id UUID,
user_id UUID NOT NULL,
product_id UUID NOT NULL,
seller_id UUID NOT NULL,
quantity INT NOT NULL CHECK (quantity > 0),
idempotency_key VARCHAR(64) NOT NULL,
expires_at TIMESTAMP NOT NULL,
status VARCHAR(20) NOT NULL DEFAULT 'active' CHECK (status IN ('active', 'committed', 'released')),
UNIQUE (user_id, idempotency_key, product_id, seller_id)
);
CREATE INDEX idx_reservation_expiry
ON reservations (status, expires_at);PostgreSQL: Orders (ACID required)
Transactional table managing customer order lifecycles, financial totals, and settlement statuses. Payment authorization and capture remain separate states.
CREATE TABLE orders (
order_id UUID PRIMARY KEY,
user_id UUID NOT NULL,
status VARCHAR(24) NOT NULL
CHECK (status IN ('created', 'payment_pending', 'payment_authorized', 'fulfilling', 'shipped', 'paid', 'delivered', 'cancelled', 'refunded')),
idempotency_key VARCHAR(64) NOT NULL,
payment_reference VARCHAR(128),
payment_status VARCHAR(20) NOT NULL
CHECK (payment_status IN ('pending', 'authorized', 'captured', 'voided', 'failed', 'refunded')),
subtotal DECIMAL(12,2) NOT NULL,
discount DECIMAL(12,2) NOT NULL DEFAULT 0,
tax DECIMAL(12,2) NOT NULL DEFAULT 0,
shipping DECIMAL(12,2) NOT NULL DEFAULT 0,
total DECIMAL(12,2) NOT NULL,
currency CHAR(3) NOT NULL,
created_at TIMESTAMP NOT NULL,
updated_at TIMESTAMP NOT NULL,
UNIQUE (user_id, idempotency_key)
);
CREATE INDEX idx_orders_user_created
ON orders (user_id, created_at DESC);
CREATE TABLE order_items (
order_item_id UUID PRIMARY KEY,
order_id UUID NOT NULL,
product_id UUID NOT NULL,
seller_id UUID NOT NULL,
quantity INT NOT NULL CHECK (quantity > 0),
unit_price DECIMAL(12,2) NOT NULL
);
CREATE INDEX idx_order_items_order
ON order_items (order_id);PostgreSQL: Transactional Outbox
The Order Service records each state transition and its domain event in the same database transaction, then an outbox relay publishes the event to Kafka asynchronously.
CREATE TABLE outbox_events (
event_id UUID PRIMARY KEY,
aggregate_type VARCHAR(50) NOT NULL,
aggregate_id UUID NOT NULL,
event_type VARCHAR(100) NOT NULL,
aggregate_version BIGINT NOT NULL,
payload JSONB NOT NULL,
created_at TIMESTAMP NOT NULL,
published_at TIMESTAMP,
UNIQUE (aggregate_type, aggregate_id, aggregate_version)
);
CREATE INDEX idx_outbox_pending
ON outbox_events (created_at)
WHERE published_at IS NULL;Redis: Cart & Flash Sale
In-memory data structures support session cart operations and flash sale admission control:
# Cart session storage (HSET cart:{user_id} {product_id} {item_json})
HSET cart:usr_102938 prod_98721 '{"quantity": 2, "price": 349.99, "seller_id": "sel_54321"}'
EXPIRE cart:usr_102938 2592000
# Flash sale inventory counter
SET flash:stock:sel_54321:prod_98721 1000
# Admission uses the atomic Lua check-and-decrement script defined in the inventory trade-off sectionFault Tolerance
| Concern | Solution |
|---|---|
| Overselling | The primary database enforces inventory invariants with transactional updates and nonnegative quantity constraints. Redis provides flash sale admission control, while durable reservation confirmation and reconciliation keep database stock authoritative |
| Order saga failure | The Saga executes idempotent compensating actions for completed steps that support compensation |
| Payment failure | Retry with an idempotency key. If authorization fails, cancel the pending order and release the inventory reservation. If the provider result is ambiguous, reconcile the provider state before compensating |
| Inventory sync (Redis vs DB) | Publish Kafka inventory events for asynchronous cache updates and run periodic reconciliation against the primary database |
| Search index lag | A 1 to 5 second discovery lag is typical, with longer lag possible during backlog. Checkout always revalidates authoritative price and live inventory |
| Flash sale stampede | Redis admission control, per user rate limits, a virtual waiting room, and Kafka buffering absorb bursts without making Redis the inventory source of truth |
Flash Sale Architecture
Managing extreme traffic surges and stock contention when thousands of users compete for limited inventory. The critical failure modes are inventory divergence, payment timeouts after stock holds, and abandoned reservation cleanup.
- Cache Prewarming: Preload available stock quantities (such as 1,000 units) into Redis before the sale event begins.
- Purchase Rate Limiting: Enforce a strict ceiling of at most one checkout attempt per user account for promotional SKUs.
- Virtual Waiting Room: Direct excess concurrent users to a queued waiting screen with real time queue position updates before allowing checkout access.
- Atomic Redis Decrement: Execute
DECR flash:stock:{seller_id}:{product_id}in Redis. If the returned value is greater than or equal to zero, issue a provisional reservation token. The durable Inventory Service then confirms the reservation against the primary database. A negative result is rejected and the counter is repaired. - Asynchronous Order Ingestion: Queue validated order tokens into Kafka so durable workers can smooth bursts and apply bounded concurrency without exhausting database connections.
Additional Considerations
Multi Seller / Marketplace Model
Architectural mechanisms required when an order spans independent third party merchants. These include multi seller settlement, logistics fulfillment, dynamic promotions, and connected system design considerations.
- Fulfillment Partitioning: A single customer order can contain items from multiple independent merchants, requiring separate shipment tracking numbers and independent fulfillment lifecycles.
- Automated Payment Splitting: The platform retains a commission fee between 15% and 30%, automatically routing the remaining balance to the respective seller accounts upon delivery confirmation.
Warehouse & Fulfillment
Coordinates logistics from warehouse receipt through final carrier delivery.
- Geographic Order Routing: Evaluates multiple regional fulfillment centers to route line items to the nearest warehouse holding sufficient stock.
- Lifecycle State Tracking: The fulfillment subsystem monitors warehouse state transitions across picking, packing, carrier handover, in transit milestones, and verified delivery.
Price and Promotion Engine
Computes dynamic pricing and validates coupons during checkout.
- Dynamic Pricing Signals: Evaluates competitor pricing feeds, seasonal demand patterns, and real time inventory velocity to calculate price updates.
- Promotion Types: Supports fixed currency discounts, percentage deductions, free shipping thresholds, and bundled buy one get one offers.
- Checkout Verification: Validates coupon eligibility, expiration timestamps, account usage limits, and minimum order values before final payment authorization.
Interview Walkthrough
A structured guide for navigating a 45 minute system design interview on ecommerce architectures.
- 25 minute cut
Focus on core catalog and checkout paths, deferring multi warehouse and edge details unless requested.
- Functional and non functional requirements with catalog vs checkout separation (3 min)
- Fast checkout walkthrough covering cart session, inventory reservation, and payment authorization (7 min)
- CQRS pattern separating read heavy catalog queries from write heavy order processing (6 min)
- Inventory reservation mechanics using optimistic locking or Redis atomic counters (5 min)
- Elasticsearch faceted search synchronized via Change Data Capture (4 min)
- 60 second checkout summary: Walk through adding an item to the Redis cart, reserving inventory with a 15 minute TTL during checkout initiation, authorizing payment through an external gateway, persisting the order record, and committing inventory. Highlight that temporary search index lag is acceptable for browsing, but checkout must always verify live stock against the primary write database.
- Open by contrasting massively read heavy catalog browsing with write heavy, consistency critical checkout, establishing early that these two distinct traffic shapes require isolated architectures.
- Explain CQRS where PostgreSQL manages orders and inventory while Elasticsearch and Redis handle product discovery and facet filtering.
- Address inventory race conditions explicitly by walking through row locks or Redis Lua check-and-decrement scripts, because overselling is the primary failure mode interviewers probe.
- Detail how Change Data Capture pipelines stream database commits into Kafka to keep Elasticsearch and Redis caches updated asynchronously.
- Frame the shopping cart as a session scoped Redis hash with asynchronous persistence, cleanly isolated from persistent order records.
- Highlight the common architectural pitfall of using a single monolithic database for catalog browsing, cart management, and order processing, which inevitably causes high volume sales browsing to exhaust checkout database connections.
Related Problems
Ecommerce platform architectures intersect directly with these specialized system designs and foundational distributed systems concepts:
- Flash Sale System explores high concurrency inventory decrement scripts, Redis token buckets, and virtual waiting room queues under extreme traffic spikes.
- Payment Gateway details distributed idempotency keys, payment service provider integrations, and two phase authorization and capture settlement.
- Recommendation System examines offline collaborative filtering, real time candidate generation, and low latency feature serving.
- Review and Rating System analyzes asynchronous sentiment analysis, spam moderation pipelines, and rolling score aggregations.
- Fraud Detection System details real time rule evaluation, velocity checks, and machine learning scoring on payment transactions.
- Distributed Transactions: 2PC vs Saga provides foundational comparisons between blocking two phase commits and orchestrator-driven compensating saga workflows.
- Caching Patterns and Invalidation explores cache aside, write through, and TTL expiration strategies across distributed caching layers.
- Microservices Patterns: Strangler, Saga, and Outbox examines the transactional outbox pattern and message relay mechanisms for atomic database writes and event publishing.
- Scaling 0 to 1M Users establishes foundational scaling tiers from monolithic databases to decoupled microservices and read replicas.
Engineering Trade-offs
Why MongoDB or DynamoDB for Product Catalog (Not MySQL)?
Balancing consistency guarantees, latency SLAs, and partition tolerance across product catalogs, inventory management, distributed transactions, and search infrastructure. The first comparison evaluates document databases against relational schemas for highly variable product category attributes.
Product data is inherently heterogeneous:
- Apparel attributes: size, color, fabric, fit
- Electronics attributes: RAM, CPU, screen size, GPU, battery life
- Book attributes: ISBN, author, publisher, page count, language
Relational Database Approaches (MySQL / PostgreSQL):
Option A: Single wide table with 200+ nullable columns
Result: Sparse data layout, high storage overhead, and less efficient indexing.
Option B: Entity Attribute Value (EAV) pattern
Result: Poor query latency, weak schema-level type enforcement, and multiple JOIN operations per attribute lookup.
Option C: Native JSON column
Result: Better schema flexibility, but query optimization and secondary indexing become harder to manage at very large catalog scale.
Document Database Approach (MongoDB / DynamoDB):
- Each product record encapsulates only its relevant attributes without NULL padding.
- Category expansions do not require relational table schema migrations, although new indexes may still need separate rollout work.
- Rich queries over nested attribute structures: db.products.find({"attributes.RAM": "16GB"})
- Secondary indexes support high cardinality filtered searches.
// Apparel document structure:
{
"product_id": "prod_102",
"title": "Running Crewneck",
"category": "clothing",
"attributes": { "size": ["S", "M", "L", "XL"], "color": "Navy", "fabric": "Cotton" }
}
// Electronics document structure:
{
"product_id": "prod_504",
"title": "Developer Laptop Pro",
"category": "electronics",
"attributes": { "RAM": "32GB", "CPU": "M3 Max", "screen": "16-inch", "GPU": "30-core" }
}
Workloads Requiring Relational Databases:
- Orders: ACID transaction guarantees spanning orders, line items, and payment references.
- Inventory: Strict consistency with CHECK (quantity >= 0) constraints.
- User accounts: Relational integrity across users, delivery addresses, and saved payment profiles.
Recommendation:
Use PostgreSQL or MySQL for transactional order and inventory domains, and MongoDB or DynamoDB for the product catalog.Inventory Management: The Hardest Problem
Analyzing concurrency control mechanisms to prevent overselling under high checkout contention:
-- Concurrency challenge:
-- 100 users click "Buy Now" simultaneously for the last remaining item in stock.
-- Without concurrency controls, 100 orders are created and 99 customers are disappointed.
-- Approach 1: Pessimistic Row Locking (SELECT FOR UPDATE)
BEGIN TRANSACTION;
SELECT available_quantity FROM inventory WHERE product_id = 'prod_123' AND seller_id = 'seller_456' FOR UPDATE;
-- Verify: available_quantity >= 1
UPDATE inventory SET available_quantity = available_quantity - 1 WHERE product_id = 'prod_123' AND seller_id = 'seller_456';
COMMIT;
-- Trade-offs:
-- Guaranteed correctness under high concurrency.
-- Lock contention serializes 100 concurrent requests on a single row.
-- For multi item orders, update inventory rows in a deterministic product order to reduce deadlock risk.
-- Risk of database connection pool exhaustion under sustained contention.
-- Approach 2: Atomic Conditional Decrement (Optimistic) [Recommended for standard operations]
UPDATE inventory
SET available_quantity = available_quantity - 1
WHERE product_id = 'prod_123' AND seller_id = 'seller_456' AND available_quantity >= 1;
-- Execution logic:
-- If rows_affected = 1: Stock successfully reserved, proceed to payment authorization.
-- If rows_affected = 0: Insufficient stock, reject checkout with an out of stock response.
-- Trade-offs:
-- No explicit application held row lock, eliminating lock wait bottlenecks.
-- Higher throughput for this single row conditional update. Larger multi item transactions can still encounter contention, so use a deterministic update order.
-- Every reservation request still incurs a relational database write.-- Approach 3: Redis In-Memory Admission Control [Recommended for high velocity flash sales]
-- Setup: Preload available inventory:
SET flash:stock:seller_456:prod_123 1000
-- Use one Lua script so check and decrement are atomic.
-- If stock is available, decrement and return the remaining count.
-- If stock is exhausted, return -1 without changing the counter.
local stock = redis.call('GET', KEYS[1])
if not stock or tonumber(stock) <= 0 then
return -1
end
return redis.call('DECR', KEYS[1])
-- Application flow:
-- Returned value >= 0: Admit the request and issue a provisional reservation token.
-- Durable Inventory Service then confirms the token with the primary database.
-- Durable reservation failure: invalidate the token and reconcile the Redis counter.
-- Returned value -1: Reject the request as sold out.
-- Trade-offs:
-- Sub-millisecond admission latency, sustaining over 100,000 requests per second at the Redis gate.
-- Redis memory state and database state can temporarily drift during failures, so Redis is not the final source of truth.
-- Requires durable reservation processing, idempotent requests, and periodic reconciliation to synchronize Redis state with the primary database.
-- Recommendation:
-- Use Approach 2 for standard catalog purchasing.
-- Use Approach 3 for limited-quantity flash sales and high velocity promotional drops.Saga Pattern vs 2PC for Distributed Transactions
Compares distributed coordination patterns across autonomous microservice databases.
Order placement spans multiple distributed services:
1. Inventory Service (reserve stock)
2. Order Service (create the order in PAYMENT_PENDING)
3. Payment Service (authorize payment)
4. Inventory Service (commit the reservation)
5. Order Service (move the order to PAYMENT_AUTHORIZED and start fulfillment)
6. Payment Service (capture payment after shipment)
7. Notification Service (send confirmation)
Two Phase Commit (2PC):
Coordinator asks all participants: "Can you commit?"
All respond YES: Coordinator broadcasts "COMMIT"
Any responds NO: Coordinator broadcasts "ABORT"
Limitations:
- Blocking: If the coordinator crashes after the prepare phase, participants can hold locks until recovery.
- High latency caused by synchronous multi phase coordination.
- Tight coupling across heterogeneous microservice databases.
Saga Pattern (Recommended):
Each step commits locally and defines an explicit compensating action for a later failure.
Step 1: Reserve Inventory <-- Compensate: Release Inventory
Step 2: Create Order (PENDING) <-- Compensate: Cancel Order
Step 3: Authorize Payment <-- Compensate: Void Authorization
Step 4: Commit Inventory + Start Fulfillment <-- Compensate: Void Authorization + release inventory if required
Step 5: Capture Payment after Shipment <-- Compensate: Refund if capture must be reversed
Step 6: Send Notification <-- No compensation required
Failure Scenario:
If Step 3 fails:
1. Cancel the pending order if required.
2. Release the inventory reservation.
3. Return a descriptive payment failure response to the user.
Coordination Models:
Orchestration (Recommended for order workflows):
Order Service acts as the centralized orchestrator.
It invokes each service sequentially and executes compensating transactions upon failure.
Benefits: Explicit state machine, clear operational tracing, centralized error handling.
Trade-off: The orchestrator coordinates all steps and requires durable execution state.
Choreography (Event-driven):
Services publish events to Kafka, and downstream services react autonomously.
Benefits: Loosely coupled service architecture.
Trade-offs: Cross service workflows become harder to trace, and compensations across event handlers become more complex.Search: Elasticsearch vs Database Full Text Search
Compares native relational text indexing with dedicated distributed search clusters.
MySQL FULLTEXT Index:
Strengths:
- Simple operational model without additional cluster infrastructure.
Limitations:
- Limited relevance scoring and ranking customizability.
- Not designed to sustain low latency faceted aggregations across brand, price range, and rating at this catalog scale.
- Lacks fuzzy matching for typographical errors (for example, matching "iphne" to "iPhone").
- No autocomplete prefix suggesters.
- Query latency becomes harder to keep predictable at 500M SKUs under high concurrent search traffic.
Elasticsearch (Recommended):
Strengths:
- BM25 relevance scoring tuned with term frequency and document length normalization.
- High performance multi dimensional faceted search using bucket aggregations.
- Fuzzy matching, synonym dictionaries, and language-specific stemming.
- Low latency autocomplete using prefix completion suggesters.
- Horizontal scalability through shard distribution across worker nodes.
- Near real time indexing with typical document visibility within 1 to 5 seconds from the database write.
Trade-offs:
- Additional infrastructure footprint and cluster management overhead.
- Not a primary source of truth, requiring continuous synchronization from transactional databases.
Synchronization Architecture:
Catalog database changes stream through Kafka using Change Data Capture or an equivalent change stream.
Elasticsearch consumer workers ingest events and index documents asynchronously.
Typical end-to-end sync latency is 1 to 5 seconds under normal load.Cart Design: Redis vs Database vs Client Side
Compares session storage tiers for shopping cart persistence, latency, and recovery.
Client Side Storage (localStorage):
Strengths:
- Zero backend server compute or storage load.
- Fully operational during offline periods.
Limitations:
- Cart is lost upon device switch, browser cache clearance, or private browsing sessions.
- No cross device synchronization.
- Vulnerable to client tampering, requiring mandatory price revalidation.
Best for: Unauthenticated guest browsing before session creation.
Relational Database (PostgreSQL or MySQL):
Strengths:
- Persistent across sessions and devices.
- Strict ACID transaction guarantees.
Limitations:
- Incurs a database write for every cart item addition, modification, or removal.
- Higher resource overhead for ephemeral sessions where many carts are abandoned.
Best for: Carts requiring stronger durability or recovery guarantees than the session cache provides.
Redis Hash (Recommended):
Strengths:
- Sub-millisecond read and write latency.
- Append Only File (AOF) persistence can preserve state across process restarts when configured durably.
- Automatic expiration using TTL (auto-expires abandoned carts after 30 days).
- Native hash data structure: HSET cart:{user_id} {product_id} {quantity}
Limitations:
- In-memory storage still requires replication and careful cluster sizing. A separate durable cart store is only needed when business requirements exceed Redis recovery guarantees.
Best for: Authenticated users at high scale.
Hybrid Architecture (Industry Standard):
1. Guest users: Cart state persists in browser localStorage.
2. Authentication: Guest cart merges into Redis session store upon user login.
3. Logged-in browsing: Redis serves as the primary low latency cart store.
4. Checkout initiation: Line item prices, discounts, and inventory stock are re-validated against the primary database.Why Separate Read and Write Databases (CQRS)?
Isolates read heavy catalog browsing from write heavy order processing to maintain predictable conversion performance.
Without CQRS:
A single database handles both catalog browsing (reads) and checkout processing (writes).
Problem: During sales events, spikes in checkout write traffic saturate database connections,
causing product browsing pages to slow down and reducing customer conversion.
With CQRS (Command Query Responsibility Segregation):
Write Model:
PostgreSQL clusters manage orders and inventory with ACID transactions and CHECK constraints.
Read Model:
Elasticsearch handles full text search and faceted filtering.
Redis caches hot product detail pages.
MongoDB read replicas serve catalog browsing queries.
Synchronization:
Transactional writes record domain events in an outbox, and an outbox relay publishes them to Kafka so read models update asynchronously.
Benefits:
- Reads and writes scale independently according to traffic asymmetry. In this baseline, 5B product page views vs 10M daily orders represents roughly 500 product page views per order, which illustrates the much heavier browsing workload.
- Each database technology is optimized for its specific access pattern.
- Write database load during peak flash sales is isolated from catalog read capacity, reducing the chance that checkout spikes degrade browsing latency.
Trade-offs:
- Eventual consistency between write and read models introduces a transient sync lag.
- Increased architectural and operational complexity.
Mitigating the Consistency Gap:
When a product price is updated, Kafka normally synchronizes Elasticsearch within 1 to 5 seconds.
To prevent checkout with stale catalog data, the checkout orchestrator always re-fetches
authoritative pricing and stock directly from the write database before charging the customer.Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.