System Design Problem

Design a Surge Pricing System like Uber or Lyft

Commonly Asked By:UberLyftBoltGrab

Interview Setup

Interview Prompt

Design a surge pricing system like Uber that adjusts ride fares in real time based on supply (available drivers) and demand (ride requests) per geographic zone.

Clarifying Questions (ask before designing)

QuestionWhy it matters
What geographic granularity is appropriate, city-wide or hexagonal cells?
  • H3 resolution 7 gives ~10M zones globally
  • 500K active. City-wide misses local hotspots.
How often should surge recalculate?Every 60s balances freshness vs oscillation. 2M GPS updates/sec need streaming aggregation.
What happens when surge service is down?Default 1.0x (underprice) keeps rides flowing; blocking rides is worse than lost margin.
Maximum surge cap and emergency override?8x cap prevents backlash; admin SET surge_cap:{city} 1.0 during disasters.

Scope

In scope

  • Real-time supply/demand tracking
  • Dynamic multiplier calculation
  • Zone-level pricing (H3)
  • Price smoothing
  • Passenger and driver communication

Out of scope (state explicitly)

  • Payment processing and driver payouts
  • Route optimization and ETA
  • Driver background checks

Functional Requirements

Start by confirming geospatial granularity and recalculation frequency with your interviewer. Ask about surge caps, driver heatmaps, and whether predictive ML is in scope.

  • Dynamic pricing: Adjust ride prices based on real-time supply (drivers) / demand (ride requests)
  • Geospatial granularity: Different surge multipliers per geographic zone (H3 hexagons)
  • Real-time computation: Recalculate surge every 60 seconds
  • Surge display: Show multiplier before booking confirmation
  • Surge caps: Maximum 8x; emergency caps (1x during disasters)
  • Driver incentives: Show high-surge zone heatmap to attract supply
  • Predictive surge: Forecast upcoming demand spikes using ML
  • Smooth transitions: Gradual ramp up/down to avoid oscillation

Non-Functional Requirements

Your interviewer will care most about sub-10 ms surge lookups on every fare quote. Mention smoothing early because without it, surge oscillates and riders game the system.

  • Low Latency: Surge lookup per zone in < 10 ms
  • Freshness: Current conditions reflected within 2 minutes
  • Scale: 500K active zones, 100K surge lookups/sec, 2M driver GPS/sec
  • Availability: 99.99%: in critical path for every ride request

Capacity Estimations

Run this math before you pick H3 resolution. Active zones and GPS events per second tell you Flink compute; surge lookup QPS drives Redis memory for zone keys.

MetricCalculationValue
Geographic zones (H3 res 7)Given~10M globally
Active zones (with activity)Given~500K
Recalculation frequencyGiven (assumption documented in value)Every 60 seconds
Ride requests / secFrom Ride requests / day ÷ 86400 (+ peak factor in value)100K
Driver location updates / secFrom Driver location updates / day ÷ 86400 (+ peak factor in value)2M
Surge lookups / secFrom Surge lookups / day ÷ 86400 (+ peak factor in value)100K

Architecture Diagram

In the room: frame surge as a supply and demand feedback loop rather than price gouging, where the multiplier attracts drivers while dampening rider demand.

Walk your interviewer through the compute vs read split. Surge multipliers are recomputed every 60 seconds per H3 zone from live demand and supply ratios. Flink sliding windows smooth the signal, while Redis caches the current multiplier so fare quotes read a single key at checkout. This architecture separates the asynchronous streaming compute lane from the synchronous lookup path. For related tracking and dispatching infrastructure, see Real-Time Vehicle Tracking and Ride Hailing System (Uber), alongside the core streaming principles in Stream Processing Basics.

Loading...

Component Deep Dives

Surge Calculation Algorithm

Surge Pricing Formula and Smoothing Functions

Next we walk through each component on the diagram. Starting with the piecewise surge formula, map the demand to supply ratio to an explicit multiplier with exponential smoothing.

For each zone (H3 cell, resolution 7 ~ 5.16 km^2), every 60 seconds:

  demand = count(ride_requests in zone, last 5 minutes) / 5  (per-minute rate)
  supply = count(available_drivers in zone, last 5 minutes) / 5
  
  ratio = demand / max(supply, 1)
  
  surge = piecewise_function(ratio):
    ratio <= 1.0:  1.0 (no surge)
    ratio 1.0-1.5: 1.0 + (ratio - 1.0) * 0.5
    ratio 1.5-3.0: 1.25 + (ratio - 1.5) * 1.0
    ratio > 3.0:   min(ratio, 8.0)

  Smoothing (prevent oscillation):
    smoothed = 0.7 * previous_surge + 0.3 * calculated_surge
    
    Without smoothing:
      T=0: high demand -> surge 3x -> riders cancel -> demand drops -> surge 1x
      T=1: riders see 1x -> rush back -> surge 3x -> cycle repeats
    With smoothing: gradual change over 3-5 minutes. Stable UX.

H3 Hexagons & Boundary Blending

Geospatial Indexing: Uber H3 Hexagonal Grid

H3 hexagons outperform square grids for geospatial pricing due to uniform neighbor adjacency. Boundary blending ensures crossing a single street boundary does not abruptly jump pricing from 1.0x to 3.0x.

Grid squares: corner cells have different distances from center than edges.
H3 hexagons: all 6 neighbors equidistant. Better circle approximation.
  Resolution 7 (~5 km^2): surge zones
  Resolution 9 (~175m): precise driver matching

Boundary blending:
  rider_surge = 0.6 * zone_surge + 0.4 * avg(neighbor_surges)
  Prevents hard surge boundaries (crossing one street changes price dramatically).

Predictive Surge (ML)

Predictive Demand Modeling and Supply Incentives

Predictive surge is a staff-level extension that positions supply incentives ahead of demand spikes so multipliers stay manageable for riders.

Features: historical demand (same hour/day/week), weather, events (concert end time),
real-time trend (demand increasing?), time of day.

Model: XGBoost per-zone. Prediction horizon: 15-30 min.
Use case: "Zone X will have 3x surge in 15 min" -> show drivers incentive to head there.
Result: supply arrives BEFORE demand spike -> surge is lower -> better UX for everyone.

Flink Window Semantics

Window Semantics: Sliding vs Tumbling Windows

Tumbling vs sliding windows directly affect surge user experience, as counters resetting at fixed minute boundaries cause sudden price drops.

Tumbling window (5 min): counts reset every 5 min.
  At minute 4:59 -> high surge. At minute 5:00 -> counter resets to 0 -> surge drops to 1x.
  Sudden drops at window boundaries = bad UX.

Sliding window (5 min window, 1 min slide):
  At any given second, the window covers the LAST 5 minutes.
  Every 1 minute, Flink re-evaluates with updated counts.
  No sudden resets. Smooth, continuous surge updates.

Implementation in Flink:
  DataStream<RideRequest> requests = ...
  requests
    .keyBy(event -> event.zoneId)
    .window(SlidingEventTimeWindows.of(
        Time.minutes(5), Time.minutes(1)))
    .aggregate(new DemandSupplyAggregator())
    .map(new SurgeCalculator())
    .addSink(new RedisSink());

Watermark strategy:
  Allow 10-second out-of-orderness for late GPS events.
  Events arriving > 10 seconds late are dropped (acceptable for surge accuracy).

API Design

Surge Multiplier and Zone Heatmap APIs

Surge evaluation exposes lightweight endpoints for real-time fare quoting and geospatial heatmaps consumed by driver client applications.

HTTP
GET /api/v1/surge?lat=37.7749&lng=-122.4194
Response: 200 OK
{
  "zone_id": "872830926cfffff",
  "surge_multiplier": 2.3,
  "estimated_fare": { "base": 15.00, "surged": 34.50 },
  "message": "Prices are 2.3x due to high demand",
  "updated_at": "2026-03-14T11:00:00Z"
}

GET /api/v1/surge/heatmap?ne_lat=37.82&ne_lng=-122.35&sw_lat=37.70&sw_lng=-122.52
Response: 200 OK
{ "zones": [{ "zone_id": "...", "surge": 2.3, "center": [37.78, -122.41] }, ...] }

Data Model

Storage Schemas: Cache, Warehouse, and Message Queue

Redis: Current Surge

surge:{zone_id}  --> Hash { multiplier: 2.3, demand: 45, supply: 20, updated_at: ts }
TTL: 120 seconds (stale if not refreshed)

surge_cap:{city}  --> FLOAT (admin override, e.g., 1.0 during emergency)
No TTL (manually removed)

ClickHouse: Historical Surge

SQL
CREATE TABLE surge_history (
    zone_id String, multiplier Float32, demand UInt32, supply UInt32,
    timestamp DateTime, date Date MATERIALIZED toDate(timestamp)
) ENGINE = MergeTree() PARTITION BY toYYYYMM(timestamp)
  ORDER BY (zone_id, timestamp);

Event Bus Design (Kafka)

Topic: driver-locations
  Partitions: 128
  Partition key: h3_zone_prefix (co-locate zone demand/supply computation)
  Retention: 24h

Topic: ride-requests
  Partition key: h3_zone_prefix

Consumer: Flink surge-pipeline (single consumer group)
  - 5-min sliding window, 1-min slide
  - Count requests + available drivers per H3 zone
  - Piecewise surge function + exponential smoothing + neighbor blending
  - HSET surge:{zone} in Redis (TTL 120s)

Sync path: GET /surge?lat=&lng= → H3 lookup → Redis GET < 10ms
Async path: GPS and ride requests never block surge reads
Alert: surge > 5x for > 30 min or supply = 0 in zone

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
402 Payment Required: account balance or payment method has insufficient funds
502 Bad Gateway: payment gateway provider timeout, poll transaction status endpoint

Fault Tolerance

ConcernSolution
Surge service downDefault to 1.0x: underprice rather than block rides
Stale surge dataTTL 120s; if expired, use historical pattern or 1.0x
Location data lag
  • Use last known positions
  • degrade gracefully
Emergency eventsAdmin override: SET surge_cap:{city} 1.0: instant
OscillationExponential smoothing prevents wild swings

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50/arch75 depth unless staff.

    • Surge as supply-demand feedback loop, not gouging (5 min)
    • H3 hex zones (~5 km²) with neighbor blending (6 min)
    • Surge = f(demand/supply) with exponential smoothing (5 min)
    • Pipeline: requests + heartbeats → Flink → Redis multiplier (5 min)
    • Show multiplier before booking with wait-and-save (4 min)
  • Frame surge as a supply-demand feedback loop rather than price gouging, explaining how the multiplier attracts drivers and dampens rider demand.
  • Walk through H3 hex zones (resolution 7, ~5 km²) with neighbor blending to avoid cliff-edge pricing at zone boundaries.
  • Explain the ratio: surge = f(demand/supply) with exponential smoothing (α ≈ 0.3) to prevent ping-pong oscillation.
  • Cover real-time pipeline: ride requests + driver heartbeats → Flink windowed aggregation → Redis current multiplier per zone.
  • Mention transparency: show multiplier before booking with a "wait and save" estimate when surge is decaying.
  • Discuss ethical guardrails: auto-cap at 1.0x during detected emergencies, manual ops override, regulatory caps (NYC 2.5x).
  • Common pitfall: updating surge instantly on every request, because drivers chase a spike that vanishes before they arrive, causing supply whiplash.

Engineering Trade-offs

Ethical Guardrails and Emergency Controls

Surge pricing balances revenue optimization with rider trust, requiring strict caps, smoothing, and emergency exemptions during crises.

Surge during emergencies (hurricane, attack):
  Auto-detect: demand > 10x normal AND news API reports emergency -> cap at 1.0x
  Manual override: ops can cap any city instantly
  Regulatory: some cities mandate caps (NYC: 2.5x during emergencies)

Transparency:
  Show surge BEFORE booking (rider chooses to accept or wait)
  "Wait and save" option: "Surge likely to decrease in ~10 minutes"
  Fare estimate with surge shown prominently (no surprise at end)

Why surge is necessary:
  1. Incentivizes drivers to high-demand areas (supply response)
  2. Reduces demand (riders who can wait, do wait)
  3. Without surge: high-demand periods have ZERO drivers -> worse for everyone

Geospatial Resolution: Zone Size Trade-offs

Choosing the appropriate H3 resolution balances spatial granularity against sample size statistical stability.

Large zones (10 km^2):
  ✓ More data points per zone -> more accurate demand/supply estimate
  ✗ Masks hyperlocal demand (airport vs nearby residential)

Small zones (0.5 km^2):
  ✓ Precise surge reflecting local conditions
  ✗ Fewer data points -> noisy, unreliable estimates
  ✗ "Surge boundary" problem: crossing one street changes price

Sweet spot: H3 resolution 7 (~5 km^2) with neighbor blending.
For airports/stadiums: use resolution 8 (~1 km^2) custom zones.

Driver Supply Response and Dynamic Feedback Loops

Surge pricing acts as an automated economic feedback loop where price signals stimulate supply rebalancing.

Surge creates a feedback loop:

  High demand -> surge rises -> drivers see high-surge zone on heatmap
  -> drivers drive to that zone -> supply increases -> surge decreases

Measurement: "supply elasticity to surge"
  How many extra drivers appear per 1x increase in surge?
  Typical: 1.5x surge attracts 30% more drivers within 10 minutes
  3.0x surge attracts 100% more drivers within 15 minutes

Driver incentive push notification:
  When zone surge > 2.0x AND duration > 3 minutes:
    Push to drivers within 10 km:
    "High demand in Downtown! Earn 2.3x fares. Estimated $45 for next ride."
    
  Only push if driver is online, not in a ride, and hasn't been pushed in last 15 min
  (prevent notification fatigue)

Surge decay on supply arrival:
  As drivers arrive -> supply increases -> smoothed surge decreases
  Important: decay must be gradual (smoothing factor 0.7)
  If decay is too fast: drivers arrive, surge drops, drivers leave, surge rises again
  (ping-pong effect)

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...