Interview Setup
Interview Prompt
Design a surge pricing system like Uber that adjusts ride fares in real time based on supply (available drivers) and demand (ride requests) per geographic zone.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What geographic granularity is appropriate, city-wide or hexagonal cells? |
|
| How often should surge recalculate? | Every 60s balances freshness vs oscillation. 2M GPS updates/sec need streaming aggregation. |
| What happens when surge service is down? | Default 1.0x (underprice) keeps rides flowing; blocking rides is worse than lost margin. |
| Maximum surge cap and emergency override? | 8x cap prevents backlash; admin SET surge_cap:{city} 1.0 during disasters. |
Scope
In scope
- Real-time supply/demand tracking
- Dynamic multiplier calculation
- Zone-level pricing (H3)
- Price smoothing
- Passenger and driver communication
Out of scope (state explicitly)
- Payment processing and driver payouts
- Route optimization and ETA
- Driver background checks
Functional Requirements
Start by confirming geospatial granularity and recalculation frequency with your interviewer. Ask about surge caps, driver heatmaps, and whether predictive ML is in scope.
- Dynamic pricing: Adjust ride prices based on real-time supply (drivers) / demand (ride requests)
- Geospatial granularity: Different surge multipliers per geographic zone (H3 hexagons)
- Real-time computation: Recalculate surge every 60 seconds
- Surge display: Show multiplier before booking confirmation
- Surge caps: Maximum 8x; emergency caps (1x during disasters)
- Driver incentives: Show high-surge zone heatmap to attract supply
- Predictive surge: Forecast upcoming demand spikes using ML
- Smooth transitions: Gradual ramp up/down to avoid oscillation
Non-Functional Requirements
Your interviewer will care most about sub-10 ms surge lookups on every fare quote. Mention smoothing early because without it, surge oscillates and riders game the system.
- Low Latency: Surge lookup per zone in < 10 ms
- Freshness: Current conditions reflected within 2 minutes
- Scale: 500K active zones, 100K surge lookups/sec, 2M driver GPS/sec
- Availability: 99.99%: in critical path for every ride request
Capacity Estimations
Run this math before you pick H3 resolution. Active zones and GPS events per second tell you Flink compute; surge lookup QPS drives Redis memory for zone keys.
| Metric | Calculation | Value |
|---|---|---|
| Geographic zones (H3 res 7) | Given | ~10M globally |
| Active zones (with activity) | Given | ~500K |
| Recalculation frequency | Given (assumption documented in value) | Every 60 seconds |
| Ride requests / sec | From Ride requests / day ÷ 86400 (+ peak factor in value) | 100K |
| Driver location updates / sec | From Driver location updates / day ÷ 86400 (+ peak factor in value) | 2M |
| Surge lookups / sec | From Surge lookups / day ÷ 86400 (+ peak factor in value) | 100K |
Architecture Diagram
In the room: frame surge as a supply and demand feedback loop rather than price gouging, where the multiplier attracts drivers while dampening rider demand.
Walk your interviewer through the compute vs read split. Surge multipliers are recomputed every 60 seconds per H3 zone from live demand and supply ratios. Flink sliding windows smooth the signal, while Redis caches the current multiplier so fare quotes read a single key at checkout. This architecture separates the asynchronous streaming compute lane from the synchronous lookup path. For related tracking and dispatching infrastructure, see Real-Time Vehicle Tracking and Ride Hailing System (Uber), alongside the core streaming principles in Stream Processing Basics.
Component Deep Dives
Surge Calculation Algorithm
Surge Pricing Formula and Smoothing Functions
Next we walk through each component on the diagram. Starting with the piecewise surge formula, map the demand to supply ratio to an explicit multiplier with exponential smoothing.
For each zone (H3 cell, resolution 7 ~ 5.16 km^2), every 60 seconds:
demand = count(ride_requests in zone, last 5 minutes) / 5 (per-minute rate)
supply = count(available_drivers in zone, last 5 minutes) / 5
ratio = demand / max(supply, 1)
surge = piecewise_function(ratio):
ratio <= 1.0: 1.0 (no surge)
ratio 1.0-1.5: 1.0 + (ratio - 1.0) * 0.5
ratio 1.5-3.0: 1.25 + (ratio - 1.5) * 1.0
ratio > 3.0: min(ratio, 8.0)
Smoothing (prevent oscillation):
smoothed = 0.7 * previous_surge + 0.3 * calculated_surge
Without smoothing:
T=0: high demand -> surge 3x -> riders cancel -> demand drops -> surge 1x
T=1: riders see 1x -> rush back -> surge 3x -> cycle repeats
With smoothing: gradual change over 3-5 minutes. Stable UX.H3 Hexagons & Boundary Blending
Geospatial Indexing: Uber H3 Hexagonal Grid
H3 hexagons outperform square grids for geospatial pricing due to uniform neighbor adjacency. Boundary blending ensures crossing a single street boundary does not abruptly jump pricing from 1.0x to 3.0x.
Grid squares: corner cells have different distances from center than edges. H3 hexagons: all 6 neighbors equidistant. Better circle approximation. Resolution 7 (~5 km^2): surge zones Resolution 9 (~175m): precise driver matching Boundary blending: rider_surge = 0.6 * zone_surge + 0.4 * avg(neighbor_surges) Prevents hard surge boundaries (crossing one street changes price dramatically).
Predictive Surge (ML)
Predictive Demand Modeling and Supply Incentives
Predictive surge is a staff-level extension that positions supply incentives ahead of demand spikes so multipliers stay manageable for riders.
Features: historical demand (same hour/day/week), weather, events (concert end time), real-time trend (demand increasing?), time of day. Model: XGBoost per-zone. Prediction horizon: 15-30 min. Use case: "Zone X will have 3x surge in 15 min" -> show drivers incentive to head there. Result: supply arrives BEFORE demand spike -> surge is lower -> better UX for everyone.
Flink Window Semantics
Window Semantics: Sliding vs Tumbling Windows
Tumbling vs sliding windows directly affect surge user experience, as counters resetting at fixed minute boundaries cause sudden price drops.
Tumbling window (5 min): counts reset every 5 min.
At minute 4:59 -> high surge. At minute 5:00 -> counter resets to 0 -> surge drops to 1x.
Sudden drops at window boundaries = bad UX.
Sliding window (5 min window, 1 min slide):
At any given second, the window covers the LAST 5 minutes.
Every 1 minute, Flink re-evaluates with updated counts.
No sudden resets. Smooth, continuous surge updates.
Implementation in Flink:
DataStream<RideRequest> requests = ...
requests
.keyBy(event -> event.zoneId)
.window(SlidingEventTimeWindows.of(
Time.minutes(5), Time.minutes(1)))
.aggregate(new DemandSupplyAggregator())
.map(new SurgeCalculator())
.addSink(new RedisSink());
Watermark strategy:
Allow 10-second out-of-orderness for late GPS events.
Events arriving > 10 seconds late are dropped (acceptable for surge accuracy).API Design
Surge Multiplier and Zone Heatmap APIs
Surge evaluation exposes lightweight endpoints for real-time fare quoting and geospatial heatmaps consumed by driver client applications.
GET /api/v1/surge?lat=37.7749&lng=-122.4194
Response: 200 OK
{
"zone_id": "872830926cfffff",
"surge_multiplier": 2.3,
"estimated_fare": { "base": 15.00, "surged": 34.50 },
"message": "Prices are 2.3x due to high demand",
"updated_at": "2026-03-14T11:00:00Z"
}
GET /api/v1/surge/heatmap?ne_lat=37.82&ne_lng=-122.35&sw_lat=37.70&sw_lng=-122.52
Response: 200 OK
{ "zones": [{ "zone_id": "...", "surge": 2.3, "center": [37.78, -122.41] }, ...] }Data Model
Storage Schemas: Cache, Warehouse, and Message Queue
Redis: Current Surge
surge:{zone_id} --> Hash { multiplier: 2.3, demand: 45, supply: 20, updated_at: ts }
TTL: 120 seconds (stale if not refreshed)
surge_cap:{city} --> FLOAT (admin override, e.g., 1.0 during emergency)
No TTL (manually removed)ClickHouse: Historical Surge
CREATE TABLE surge_history (
zone_id String, multiplier Float32, demand UInt32, supply UInt32,
timestamp DateTime, date Date MATERIALIZED toDate(timestamp)
) ENGINE = MergeTree() PARTITION BY toYYYYMM(timestamp)
ORDER BY (zone_id, timestamp);Event Bus Design (Kafka)
Topic: driver-locations
Partitions: 128
Partition key: h3_zone_prefix (co-locate zone demand/supply computation)
Retention: 24h
Topic: ride-requests
Partition key: h3_zone_prefix
Consumer: Flink surge-pipeline (single consumer group)
- 5-min sliding window, 1-min slide
- Count requests + available drivers per H3 zone
- Piecewise surge function + exponential smoothing + neighbor blending
- HSET surge:{zone} in Redis (TTL 120s)
Sync path: GET /surge?lat=&lng= → H3 lookup → Redis GET < 10ms
Async path: GPS and ride requests never block surge reads
Alert: surge > 5x for > 30 min or supply = 0 in zoneCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff 402 Payment Required: account balance or payment method has insufficient funds 502 Bad Gateway: payment gateway provider timeout, poll transaction status endpoint
Fault Tolerance
| Concern | Solution |
|---|---|
| Surge service down | Default to 1.0x: underprice rather than block rides |
| Stale surge data | TTL 120s; if expired, use historical pattern or 1.0x |
| Location data lag |
|
| Emergency events | Admin override: SET surge_cap:{city} 1.0: instant |
| Oscillation | Exponential smoothing prevents wild swings |
Additional Considerations
Interview Walkthrough
- 25-minute cut
Skip arch50/arch75 depth unless staff.
- Surge as supply-demand feedback loop, not gouging (5 min)
- H3 hex zones (~5 km²) with neighbor blending (6 min)
- Surge = f(demand/supply) with exponential smoothing (5 min)
- Pipeline: requests + heartbeats → Flink → Redis multiplier (5 min)
- Show multiplier before booking with wait-and-save (4 min)
- Frame surge as a supply-demand feedback loop rather than price gouging, explaining how the multiplier attracts drivers and dampens rider demand.
- Walk through H3 hex zones (resolution 7, ~5 km²) with neighbor blending to avoid cliff-edge pricing at zone boundaries.
- Explain the ratio: surge = f(demand/supply) with exponential smoothing (α ≈ 0.3) to prevent ping-pong oscillation.
- Cover real-time pipeline: ride requests + driver heartbeats → Flink windowed aggregation → Redis current multiplier per zone.
- Mention transparency: show multiplier before booking with a "wait and save" estimate when surge is decaying.
- Discuss ethical guardrails: auto-cap at 1.0x during detected emergencies, manual ops override, regulatory caps (NYC 2.5x).
- Common pitfall: updating surge instantly on every request, because drivers chase a spike that vanishes before they arrive, causing supply whiplash.
Engineering Trade-offs
Ethical Guardrails and Emergency Controls
Surge pricing balances revenue optimization with rider trust, requiring strict caps, smoothing, and emergency exemptions during crises.
Surge during emergencies (hurricane, attack): Auto-detect: demand > 10x normal AND news API reports emergency -> cap at 1.0x Manual override: ops can cap any city instantly Regulatory: some cities mandate caps (NYC: 2.5x during emergencies) Transparency: Show surge BEFORE booking (rider chooses to accept or wait) "Wait and save" option: "Surge likely to decrease in ~10 minutes" Fare estimate with surge shown prominently (no surprise at end) Why surge is necessary: 1. Incentivizes drivers to high-demand areas (supply response) 2. Reduces demand (riders who can wait, do wait) 3. Without surge: high-demand periods have ZERO drivers -> worse for everyone
Geospatial Resolution: Zone Size Trade-offs
Choosing the appropriate H3 resolution balances spatial granularity against sample size statistical stability.
Large zones (10 km^2): ✓ More data points per zone -> more accurate demand/supply estimate ✗ Masks hyperlocal demand (airport vs nearby residential) Small zones (0.5 km^2): ✓ Precise surge reflecting local conditions ✗ Fewer data points -> noisy, unreliable estimates ✗ "Surge boundary" problem: crossing one street changes price Sweet spot: H3 resolution 7 (~5 km^2) with neighbor blending. For airports/stadiums: use resolution 8 (~1 km^2) custom zones.
Driver Supply Response and Dynamic Feedback Loops
Surge pricing acts as an automated economic feedback loop where price signals stimulate supply rebalancing.
Surge creates a feedback loop:
High demand -> surge rises -> drivers see high-surge zone on heatmap
-> drivers drive to that zone -> supply increases -> surge decreases
Measurement: "supply elasticity to surge"
How many extra drivers appear per 1x increase in surge?
Typical: 1.5x surge attracts 30% more drivers within 10 minutes
3.0x surge attracts 100% more drivers within 15 minutes
Driver incentive push notification:
When zone surge > 2.0x AND duration > 3 minutes:
Push to drivers within 10 km:
"High demand in Downtown! Earn 2.3x fares. Estimated $45 for next ride."
Only push if driver is online, not in a ride, and hasn't been pushed in last 15 min
(prevent notification fatigue)
Surge decay on supply arrival:
As drivers arrive -> supply increases -> smoothed surge decreases
Important: decay must be gradual (smoothing factor 0.7)
If decay is too fast: drivers arrive, surge drops, drivers leave, surge rises again
(ping-pong effect)Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.