Interview Setup
Interview Prompt
Design an order management system supporting order placement, lifecycle tracking through multi-warehouse fulfillment, cancellations, returns, and distributed saga coordination across payment and inventory.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Peak order placement throughput? | Average 115 orders per second scaling to 10K during promotional flash sales determines database write capacity. |
| Synchronous vs asynchronous payment handling? | A distributed saga model uses payment event publishing to advance states and triggers compensating steps on failure. |
| Support for split shipments? | Orders containing items across multiple warehouses require dynamic splitting into independent shipment records. |
| Returns and refund scope? | A reverse saga orchestrates return label generation, warehouse receipt confirmation, and payment refund processing. |
Scope
In scope
- Order placement and cancellation
- Finite state machine
- Payment saga orchestration
- Shipment tracking
- Returns and refunds
- Idempotency safeguards
Out of scope (state explicitly)
- Inventory service internals
- Payment gateway integration
- Warehouse pick and pack execution
Functional Requirements
Align with your interviewer on the order lifecycle scope. Inquire whether split multi-warehouse shipments, return workflows, and carrier tracking integration fall within the required design boundaries.
- Place order: Convert active cart items from the Shopping Cart System into a confirmed order with shipping addresses, payment methods, and delivery tiers.
- Order lifecycle management: Progress orders through a finite state machine: placed, payment_confirmed, processing, shipped, out_for_delivery, delivered, returned, and cancelled.
- Multi-item split shipments: Split single orders across multiple warehouses or third-party sellers into independent shipment tracking records.
- Real-time order tracking: Ingest carrier status updates (FedEx, UPS, USPS) via webhooks and polling to surface live transit checkpoints.
- Returns and refunds: Provide customer return initiation, generate printable return labels, and authorize automated refunds upon warehouse receipt.
- Order history and search: Enable customers and support agents to search, filter, and review historic orders.
- Customer notifications: Dispatch email, SMS, and push alerts at every meaningful lifecycle transition.
- Automated invoice generation: Generate immutable PDF invoices stored in object storage upon payment authorization.
- Order cancellation: Permit full or partial order cancellations prior to physical warehouse dispatch.
- Address and preference modifications: Allow shipping address modifications before orders reach the warehouse processing cutoff.
Non-Functional Requirements
Interviewers focus primarily on ACID state transition guarantees and saga orchestration across payment and inventory services. Emphasize idempotency during order placement, because network retries are inevitable at 10M daily orders.
- Strong Consistency: Enforce strict ACID guarantees across order states to eliminate lost orders and phantom charges.
- High Availability: Achieve 99.99% uptime because checkout degradation immediately blocks commercial revenue.
- Low Latency: Return initial order placement confirmations in under 2 seconds, including inventory hold and payment handoff.
- Scalability: Support 10M+ orders daily and handle 50M+ status inquiries, scaling to 10K orders per second during events like a Flash Sale System.
- Idempotency: Ensure accidental double submissions or client retry storms never create duplicate orders or bill customers twice.
- Comprehensive Auditability: Record every state mutation in an append-only, immutable audit trail.
- Long-Term Durability: Retain finalized order records for 7+ years to comply with statutory accounting and tax regulations.
Capacity Estimations
Examine throughput and capacity numbers before proposing relational database partitioning schemes. Daily order volume and status read frequencies establish the read-to-write ratio, while multi-year regulatory retention mandates a tiered archiving strategy.
| Metric | Calculation | Value |
|---|---|---|
| Orders / day | Given (assumption documented in value) | 10M |
| Orders / sec | 10M ÷ 86400 | ~115 (peak 10K during sales) |
| Order status checks / day | Given (assumption documented in value) | 50M |
| Avg items per order | Given (typical workload assumption) | 3 |
| Order record size | Given | ~5 KB (with items, addresses, payment) |
| Storage / day | 10M x 5 KB | 50 GB |
| Storage / year | Given | ~18 TB |
Architecture Diagram
In the room: say saga with compensating steps rather than 2PC, because payment, inventory, and shipping cannot share one distributed lock.
Walk your interviewer through the distributed saga workflow. Orders serve as the authoritative system of record in PostgreSQL, while payment, inventory, and shipping coordinate through asynchronous event topics with an immutable state log for every transition, and I draw each downstream service as an asynchronous compensating step rather than a 2PC participant.
Component Deep Dives
Order State Machine
Lifecycle State Transitions
Every saga progression maps directly to a validated transition in the state machine. State transitions represent append-only facts, and invalid state jumps are rejected at the API gateway rather than patched silently in the database.
Event Bus Design (Kafka)
Order state changes publish to an event streaming bus using the transactional outbox pattern. Downstream consumers for warehouse fulfillment, notification dispatch, and search indexing operate asynchronously, completely decoupled from the synchronous place-order path. Learn more in Message Queues Fundamentals.
Topic: order-events
Partitions: 128
Partition key: order_id (preserves per-order state transition ordering)
Retention: 30 days (allows saga replay and regulatory audit compliance)
Replication factor: 3, min.insync.replicas: 2
Producer: Order Service following PostgreSQL transaction commit (transactional outbox pattern)
Event: { order_id, user_id, status, items[], payment_id, timestamp }
Consumer groups:
1. notification: email, SMS, and push alerts on confirmed, shipped, and delivered states
2. shipping: generates shipping labels and assigns carriers upon payment confirmation
3. search-indexer: CDC to Elasticsearch for customer service dashboard queries
4. analytics: ClickHouse order funnel metrics
5. recommendation: co-purchase association mining
Topic: payment-events (success or failure) -> Order Service saga compensation
Topic: shipment-events (carrier webhooks) -> Tracking Service
Topic: return-events (initiated, received, refunded) -> Inventory release workflows
Sync path: POST /orders saga (inventory reservation and payment authorization) in under 500 ms.
Async path: notifications, warehouse label generation, and search indexing never block order confirmation.
DLQ: order-events-dlq, triggering alerts when consumer lag exceeds 60 seconds.Order Placement: Saga Pattern
Distributed Saga Orchestration
The saga pattern is the central architecture topic for distributed order workflows. Walk through each sequential forward step and its corresponding compensating transaction before detailing service boundaries. For a detailed breakdown of 2PC vs Saga, review Distributed Transactions: 2PC vs Saga and Microservices Patterns: Saga and Outbox.
Step 1 records the initial order entity, Step 2 reserves inventory in the Inventory Management System, Step 3 captures payment, and Step 4 confirms the order. If any step fails, compensating transactions execute in reverse order to release holds and restore consistency.
Placing an order involves multiple services. Use SAGA pattern (not 2PC):
Step 1: Create order (Order Service)
INSERT order with status = 'PLACED'
Step 2: Reserve inventory (Inventory Service)
For each item: reserve stock at selected warehouse
If ANY item out of stock --> compensate: cancel order, release other reservations
Step 3: Charge payment (Payment Service)
Authorize + capture payment
If payment fails --> compensate: release all inventory reservations, mark order PAYMENT_FAILED
Step 4: Confirm order (Order Service)
Update status = 'PAYMENT_CONFIRMED'
Publish order-confirmed event
Step 5 (async): Generate invoice, send confirmation email, update analytics
Saga compensation (rollback):
If Step 3 fails (payment declined):
- Undo Step 2: release inventory reservations
- Undo Step 1: mark order as PAYMENT_FAILED
- Notify user: "Payment failed. Your items are still in your cart."
If Step 2 fails (out of stock):
- Undo Step 1: mark order as CANCELLED
- Notify user: "Sorry, [item] is no longer available."
Why SAGA (not distributed transaction / 2PC)?
2PC: requires all services to hold locks simultaneously --> blocks everything if one is slow
SAGA: each step commits independently; compensating transactions undo on failure
At e-commerce scale, SAGA is the industry standard (Amazon, Shopify, etc.)Split Shipments
Multi-Warehouse Fulfillment
Multi-warehouse orders represent a common real-world operational requirement. Explain how parent order states derive dynamically from child shipment states, ensuring partial shipments and cancellations compose cleanly.
Split Shipments: Multi-Warehouse Example: An order with 3 items sourced from different warehouses is automatically split: - Item A: only available at WH-NYC - Item B: only available at WH-LAX - Item C: available at both warehouses, allocated to WH-NYC due to regional proximity Result: 2 independent shipments: - Shipment 1 (WH-NYC): Item A + Item C -> FedEx tracking ID 12345 - Shipment 2 (WH-LAX): Item B -> UPS tracking ID 67890 Each shipment has its own tracking, carrier, state machine, and EDD. Parent order state is derived: - If any shipment SHIPPED -> order = PARTIALLY_SHIPPED - All shipments DELIVERED -> order = DELIVERED - Partial cancellation cancels one shipment independently.
Idempotent Order Placement
Client-Generated Idempotency Keys
Network timeouts during checkout are guaranteed at massive scale, so interviewers expect client-generated idempotency tokens rather than relying solely on server-side deduplication.
Problem: User clicks "Place Order" and network times out. User clicks again.
Without idempotency: two identical orders created, charged twice.
Solution:
Client generates idempotency_key (UUID) on checkout page load.
Both clicks send same idempotency_key.
Server:
1. Check Redis: GET idempotency:{key}
If exists --> return cached order (already processed)
2. Check PostgreSQL: SELECT order_id FROM orders WHERE idempotency_key = ?
If exists --> return existing order
3. If neither exists --> proceed with order creation
4. After creation: SET idempotency:{key} {order_id} in Redis (TTL 24h)
Result: exactly one order, regardless of retries.API Design
Order Endpoints
The order API exposes idempotent resource operations for order creation, tracking queries, customer-initiated cancellations, and returns.
Place Order
POST /api/v1/orders
Idempotency-Key: "order-abc-123"
{
"items": [
{"sku_id": "SKU-123", "quantity": 2, "price": 29.99},
{"sku_id": "SKU-456", "quantity": 1, "price": 49.99}
],
"shipping_address": { "street": "...", "city": "...", "zip": "...", "country": "US" },
"payment_method_id": "pm-uuid",
"delivery_preference": "standard"
}
Response: 201 Created
{
"order_id": "order-uuid",
"status": "payment_pending",
"estimated_delivery": "2026-03-18",
"total": 109.97,
"shipments": [
{"shipment_id": "ship-1", "items": ["SKU-123","SKU-456"], "warehouse": "WH-NYC"}
]
}Get Order Status
GET /api/v1/orders/{order_id}
Response: 200 OK
{
"order_id": "order-uuid",
"status": "shipped",
"placed_at": "2026-03-14T10:00:00Z",
"items": [...],
"shipments": [
{"shipment_id": "ship-1", "status": "in_transit", "carrier": "FedEx",
"tracking_number": "794644790132", "estimated_delivery": "2026-03-18",
"tracking_url": "https://fedex.com/track?id=794644790132"}
],
"payment": { "method": "Visa ending 4242", "amount": 109.97, "status": "captured" }
}Cancel Order
POST /api/v1/orders/{order_id}/cancel
{ "reason": "changed_mind" }
Response: 200 OK
{ "status": "cancelled", "refund_amount": 109.97, "refund_status": "processing" }Initiate Return
POST /api/v1/orders/{order_id}/return
{
"items": [{"sku_id": "SKU-123", "quantity": 1, "reason": "defective"}],
"return_method": "mail"
}
Response: 200 OK
{
"return_id": "ret-uuid",
"return_label_url": "https://s3.../return-label.pdf",
"refund_amount": 29.99,
"refund_status": "pending_return_receipt"
}Common Error Responses
Structured error contracts returned for payment declines, inventory exhaustion, and invalid state transitions.
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 402 Payment Required: account balance or payment method has insufficient funds 502 Bad Gateway: payment gateway provider timeout, poll transaction status endpoint
Data Model
PostgreSQL: Source of Truth
Normalized relational schemas persist parent orders, line items, warehouse shipments, returns, and the immutable transition audit log.
CREATE TABLE orders (
order_id UUID PRIMARY KEY,
user_id UUID NOT NULL,
status VARCHAR(30) NOT NULL DEFAULT 'placed',
subtotal DECIMAL(10,2),
tax DECIMAL(10,2),
shipping_cost DECIMAL(10,2),
total DECIMAL(10,2),
currency CHAR(3) DEFAULT 'USD',
shipping_address JSONB,
payment_method_id VARCHAR(64),
payment_status VARCHAR(20),
idempotency_key VARCHAR(64) UNIQUE,
placed_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_user (user_id, placed_at DESC),
INDEX idx_status (status)
);
CREATE TABLE order_items (
item_id BIGSERIAL PRIMARY KEY,
order_id UUID NOT NULL REFERENCES orders(order_id),
sku_id VARCHAR(50) NOT NULL,
quantity INT NOT NULL,
unit_price DECIMAL(10,2),
total_price DECIMAL(10,2),
shipment_id UUID,
status VARCHAR(20) DEFAULT 'active',
INDEX idx_order (order_id)
);
CREATE TABLE shipments (
shipment_id UUID PRIMARY KEY,
order_id UUID NOT NULL REFERENCES orders(order_id),
warehouse_id VARCHAR(50),
carrier VARCHAR(20),
tracking_number VARCHAR(64),
status VARCHAR(30) DEFAULT 'pending',
shipped_at TIMESTAMPTZ,
delivered_at TIMESTAMPTZ,
INDEX idx_order (order_id),
INDEX idx_tracking (tracking_number)
);
CREATE TABLE order_state_log (
log_id BIGSERIAL PRIMARY KEY,
order_id UUID NOT NULL,
from_state VARCHAR(30),
to_state VARCHAR(30) NOT NULL,
actor VARCHAR(50),
reason TEXT,
created_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_order (order_id, created_at)
);
CREATE TABLE returns (
return_id UUID PRIMARY KEY,
order_id UUID NOT NULL,
user_id UUID NOT NULL,
status VARCHAR(20) DEFAULT 'initiated',
refund_amount DECIMAL(10,2),
reason TEXT,
return_label_url TEXT,
created_at TIMESTAMPTZ DEFAULT NOW(),
INDEX idx_order (order_id)
);Redis: Caching and Idempotency
Fast in-memory structures cache order tracking summaries for status polling and enforce 24-hour idempotency deduplication.
# Order status cache (avoids direct database queries for polling)
order_status:{order_id} --> Hash { status, tracking, updated_at }
TTL: 300
# Idempotency token cache (prevents duplicate checkout execution)
idempotency:{key} --> order_id (if already processed)
TTL: 86400
# User recent orders list (for rapid profile loading)
user_orders:{user_id} --> List of order_ids (last 20)
TTL: 3600Kafka Event Streaming Specifications
State changes stream across dedicated event partitions to synchronize payment, shipping, and reverse return lifecycles.
Topic: order-events (placed, confirmed, shipped, delivered: captures state transitions) Topic: payment-events (payment authorization outcome consumed by order orchestrator) Topic: shipment-events (carrier checkpoint updates from tracking webhooks) Topic: return-events (return initiated, parcel received, refund authorized)
Fault Tolerance
| Concern | Solution |
|---|---|
| Duplicate order placement | Idempotency keys generated per checkout session enforced via database UNIQUE constraints. |
| Payment succeeds but order update fails | Saga compensation pattern where payment events reconcile into order state and background workers resolve stranded charges. |
| Inventory reserved but payment fails | Compensating release action triggered on payment failure, backed by a TTL-based auto-release timer. |
| State machine corruption | Strict transition validation enforced in application logic, verified with an append-only order_state_log table. |
| Carrier webhook failure | Scheduled poller querying carrier tracking endpoints every 30 minutes for all active in-transit shipments. |
| Database node failure | PostgreSQL synchronous standby replication ensuring zero data loss failovers. |
Additional Considerations
Interview Walkthrough
- 25-minute pacing strategy
Prioritize state transitions and saga compensation mechanics before detailing carrier webhooks or analytics CDC pipelines.
- Order state machine: CREATED to PAYMENT_PENDING to SHIPPED (5 min)
- Saga orchestration over 2PC: reserve to pay to ship (6 min)
- Idempotency keys on place-order API (5 min)
- PostgreSQL ACID writes with Kafka for downstream events (5 min)
- Warehouse pick lists and email notifications via async consumers (4 min)
- Lead with the order state machine: CREATED to PAYMENT_PENDING to CONFIRMED to SHIPPED to DELIVERED, ensuring each transition is recorded as an immutable event rather than a silent database overwrite.
- Explain saga orchestration over 2PC: reserve inventory, charge payment, and create shipments, with compensating transactions automatically rolling back on failure.
- Cover idempotency keys on order placement: at 10M orders per day, even a 1% retry rate creates 100K potential duplicate charges without deduplication.
- Walk through PostgreSQL as the write path for ACID state transitions, accompanied by Elasticsearch as the read path for customer-service search via change data capture.
- Mention Kafka events for downstream consumers such as warehouse pick lists, email notifications, and analytics, none of which block the synchronous order confirmation response.
- Discuss cancellation rules: only allow cancellations prior to the SHIPPED state, triggering inventory release and payment refund as compensating steps.
- Common pitfall: attempting two-phase commit across inventory, payment, and shipping, where one slow dependency locks all participants and the coordinator turns into a single point of failure.
Engineering Trade-offs
Consistency vs Availability in Order Orchestration
Distributed order systems balance strong consistency against overall availability, using saga compensating actions rather than rigid two-phase commit protocols.
Saga vs 2PC for Distributed Order Placement
2PC (Two-Phase Commit):
Coordinator asks all services: "Can you commit?"
All say yes --> "Commit." All commit atomically.
Problems at e-commerce scale:
- Coordinator is SPOF
- All services hold locks during prepare phase --> latency, blocking
- If any service is slow --> ALL are blocked
- Not supported across heterogeneous systems (different DBs, services)
Saga (Choreography or Orchestration):
Each step commits locally. Failures trigger compensating transactions.
Choreography: each service publishes event; next service reacts
Order created --> Inventory listens, reserves stock --> Payment listens, charges
Loose coupling but hard to debug (no central view of saga progress)
Orchestration (recommended): central orchestrator coordinates steps
Order Service calls Inventory, then Payment, then Shipping
On failure: orchestrator calls compensating actions in reverse order
Easy to monitor, debug, and modify
Trade-off:
2PC: strong consistency, but fragile and slow at scale
Saga: eventual consistency during saga execution, but resilient and fast
For e-commerce: Saga is the clear winner. Brief inconsistency during
the 2-second order placement window is acceptable.Order Search: Why Elasticsearch
Customer service needs: "Find all orders for user X with product Y shipped in March" PostgreSQL can do this, but: - Full-text search across product names, addresses is slow - Complex compound filters with pagination are expensive - At 10M orders/day, queries slow down without heavy indexing Elasticsearch: - Index order data (denormalized): order_id, user, items, status, dates - Support full-text search + filters + aggregations - Sub-200ms response for complex queries CDC pipeline: PostgreSQL --> Debezium --> Kafka --> ES consumer --> Elasticsearch Lag: < 5 seconds from order placed to searchable in ES Use PostgreSQL for: order placement, state transitions (ACID) Use Elasticsearch for: search, customer service dashboard, analytics queries
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.