System Design Problem

Design a Real-Time Vehicle Tracking System

Commonly Asked By:UberLyftGrabDoorDash

Interview Setup

Interview Prompt

Design a real time vehicle tracking system for fleet operators and delivery companies. Ingest GPS from millions of vehicles, show live positions on a dashboard, store route history for replay, and push ETA updates to end customers.

Clarifying Questions (ask before designing)

QuestionWhy it matters
How often do vehicles report location and over what protocol?At 10M vehicles, 5 second updates over MQTT require roughly 21 TB/day, while the assumed HTTP request overhead alone is roughly 690 TB/day.
How fresh must dashboard updates be?3-second freshness requires Redis latest position writes and Pub/Sub, not polling PostgreSQL.
How long must historical trails be retained?90 days raw + 2 years aggregated drives TimescaleDB hot tier vs ClickHouse cold tier.
Do we need geofence alerts and speed violations?Event detection adds Flink stateful processing beyond simple position storage.

Scope

In scope

  • MQTT ingestion for 10M vehicles
  • Latest position in Redis with Pub/Sub to dashboards
  • Trail persistence in TimescaleDB
  • Fleet dashboard with 500K concurrent viewers
  • Historical route replay
  • Movement status detection (moving/idle/offline)

Out of scope (state explicitly)

  • Turn by turn navigation and map tile rendering (covered in Map Rendering and Navigation)
  • Full scale polygon geofencing platform (covered in Geofencing Service)
  • Vehicle maintenance scheduling

Functional Requirements

Clarify coordinate update frequency and subscriber personas across riders, drivers, and operations dispatchers. The system requires GPS telemetry ingestion, real time map positioning, comprehensive trip history playback, and integration with an ETA service for customer facing delivery estimates.

In the room: distinguish this problem from ETA Calculation Service, because vehicle tracking focuses on live coordinate ingestion, current state, and alerting rather than the ETA algorithm itself.

  • Live tracking: Display real time location of vehicles on a map (fleet management, delivery tracking)
  • Location ingestion: Ingest GPS coordinates from thousands/millions of vehicles at configurable intervals
  • Trip tracking: Track active trips with start, waypoints, and end, and record the full route trail
  • Geofence alerts: Trigger alerts when vehicles enter/exit defined zones
  • Historical playback: Replay a vehicle's route over any past time period
  • Speed/idle alerts: Detect speeding, excessive idling, harsh braking events
  • Fleet dashboard: Real time overview of entire fleet: active, idle, offline vehicles
  • ETA for deliveries: Show customer the live position and ETA of their delivery
  • Multi tenant: Support multiple fleet operators on the same platform

Non-Functional Requirements

Position updates every few seconds should become visible on the map within the 3-second SLO while supporting millions of concurrent active trips.

  • Real time: Location updates visible on dashboard within 3 seconds of device report
  • High Throughput: Handle 2M location updates/sec (10M vehicles x 5s interval)
  • Scalability: Support 10M+ simultaneously tracked vehicles
  • Durability: No location point accepted by the ingest path is lost after durable acceptance (required for compliance, insurance, disputes)
  • Low Latency: Map updates in < 3 seconds, with dashboard aggregations in < 500 ms
  • Availability: 99.99%
  • Data Retention: Raw location trails retained for 90 days, with aggregated data retained for 2+ years
  • Bandwidth Efficiency: Minimize cellular data usage for vehicle trackers

Capacity Estimations

GPS points per second and subscriber fan out drive Kafka and WebSocket sizing.

MetricCalculationValue
Active vehiclesGiven (assumption documented in value)10M
Avg location update intervalGiven (typical workload assumption)5 seconds
Location updates / day2M/sec x 86400~173B events
Location updates / sec10M vehicles ÷ 5 sec interval~2M
Location point sizeGiven (assumption documented in value)100 bytes
Raw data / day2M x 100B x 86400~17 TB
Concurrent dashboard viewersGiven (peak load assumption)500K
Persistent client connections500K dashboard WebSocket + 10M vehicle MQTT10.5M total
Telemetry payload throughput2M updates/sec x 100B~200 MB/sec
Kafka logical retention2M x 100B x 172800 sec~34.6 TB for 48h before replication
Kafka replicated retention34.6 TB x RF 3~103.7 TB before broker overhead
Redis latest state logical payload10M vehicles x ~200B~2 GB before Redis overhead
Historical queries / secGiven (peak workload assumption)10K

Architecture Diagram

Vehicle tracking is a write heavy telemetry pipeline. Devices publish GPS coordinates every few seconds over MQTT, Kafka absorbs traffic surges, Flink validates and processes the stream, Redis holds the latest position for low latency reads, and TimescaleDB stores the historical trail. WebSocket gateways shard client subscriptions by geohash so one map viewport does not trigger global fan out. Geofence events are consumed as a stream, while customer facing ETA computation is delegated to ETA Calculation Service. Map presentation and polygon geofencing integrate with Map Rendering & Navigation and Geofencing Service.

In the room: comparing MQTT with HTTP at 5 second intervals across 10M vehicles demonstrates that persistent connections eliminate massive TCP and TLS handshake overhead.

Loading...

Component Deep Dives

Streaming Telemetry Architecture

The write path is optimized for high volume streaming telemetry. MQTT keeps 10M persistent vehicle connections lean, Kafka buffers bursts, and Flink validates and fans out updates to Redis for live lookups and TimescaleDB for historical trails. Dashboard operators subscribe over WebSocket channels to receive real time updates without polling.

Telemetry flows from vehicle sensors through the gateway into streaming processors. The processors update live state in Redis and push deltas over WebSockets. Compare with streaming patterns in Stream Processing Basics.

Connection Gateway: Handling 10M Persistent Connections

MQTT vs HTTP for vehicle GPS: HTTP requires approximately 4 KB of overhead per update. Across 12 updates per minute and 10M vehicles, the assumed headers alone total roughly 690 TB per day. MQTT ⭐: Persistent connections with a 20-byte header and 100-byte payload require only 120 bytes per update, reducing the transport payload estimate to approximately 21 TB per day. QoS 1 provides at least once delivery. Reconnection logic handles tunnel transitions, while Last Will and Testament publishes an offline event after the broker detects an unexpected disconnect.

MQTT Broker Cluster (EMQX or VerneMQ): Using approximately 200K concurrent connections per broker as a starting capacity assumption gives about 50 broker nodes for 10M vehicles. Benchmark connection churn, message rate, memory usage, and failover behavior before fixing the production node count.

Location Processor: Updating Latest Position in Redis

A Flink streaming job consumes from the Kafka vehicle-locations topic, validates coordinates, deduplicates points by event_id, enriches telemetry with fleet metadata, writes to Redis via HSET under vehicle:{vehicle_id}, refreshes a separate vehicle_connectivity:{vehicle_id} key from broker session state, and publishes to geohash scoped Redis Pub/Sub channels for dashboard distribution. The latest position key can expire after 5 minutes as a stale telemetry guard, while connectivity state distinguishes GPS_LOST from a disconnected vehicle. Because Redis Pub/Sub is ephemeral, clients can refresh current state through the REST endpoint after reconnecting.

Movement Status Detection: Speeds above 5 km/h indicate moving. Speeds below 2 km/h for more than 5 minutes indicate parked. Speeds from 2 through 5 km/h are treated as idling unless a domain rule says otherwise. A separate connectivity key derived from MQTT session heartbeats distinguishes GPS_LOST from offline status when the vehicle remains connected but stops sending GPS points.

Trail Writer: Persisting Location History

Storage choice: TimescaleDB serves the recent 90 day trail tier with time based partitioning, 10 to 20 times compression, SQL querying, PostGIS integration, and continuous aggregates. ClickHouse serves the cold tier for 2 or more years of aggregate analytics.

Batch writing: Flink accumulates location points and executes bulk inserts into TimescaleDB every 5 seconds. Processing 2M points per second across a 5 second window yields 10M rows per batch, which distributes across 4 shards at 2.5M rows per shard, with an under 1 second completion target that must be validated by load testing.

Dashboard: Real Time Fleet View

Fleet operators connect via WebSockets to backend servers that subscribe to geohash scoped Redis Pub/Sub channels. Initial page loads use the fleet vehicle index in Redis to fetch current positions. Viewport filtering ensures servers transmit only vehicles falling within active map coordinates, while client side libraries handle clustering at lower zoom levels. When a dashboard reconnects after a missed Pub/Sub event, it refreshes current state through the REST API.

Event Bus Design (Kafka)

Topic: vehicle-locations
  Partitions: 512 (illustrative starting point for 2M updates/sec)
  Partition key: vehicle_id
  Retention: 48h (replay after TimescaleDB or Flink recovery)
  Replication factor: 3, min.insync.replicas: 2

Producer: MQTT broker bridge (idempotent producer)
  Event: { event_id, vehicle_id, lat, lng, speed, heading, timestamp, events[] }

Consumer groups:
  1. location-processor: Flink: validate, dedup, update vehicle:{id} in Redis (5-min TTL)
     Publish to fleet:{fleet_id}:geo:{geohash_prefix} for dashboard WebSocket fan out
  2. trail-writer: batch INSERT into TimescaleDB location_points hypertable
  3. alert-engine: geofence breach, speed > 300 km/h, idle > 30 min -> vehicle-alerts topic

Topic: vehicle-alerts (64 partitions) -> Notification Service (push, webhook)
Topic: vehicle-status-changes (64 partitions) -> fleet management integrations

Ingest path: broker ACK < 100ms after durable acceptance, with the Kafka bridge retrying independently
Hot read path: dashboard reads current position from Redis
Async path: trail history, alerts, and analytics never block GPS ingest
DLQ: vehicle-locations-dlq after 3 retries, with an alert when consumer lag > 60s
Recovery: replay Kafka to rebuild Redis latest state after cache loss

API Design

Telemetry and Fleet Management APIs

The platform accepts high throughput binary telemetry over MQTT while exposing REST and WebSocket endpoints for fleet operators to monitor status, review history, query aggregate statistics, and serve customer delivery tracking views.

Domain Types

TYPESCRIPT
interface VehicleLocationUpdate {
  eventId: string;
  vehicleId: string;
  lat: number;
  lng: number;
  speed: number;
  heading: number;
  timestamp: number;
  fuelLevel?: number;
  engineStatus?: "on" | "off";
  events?: string[];
}

interface DriverSummary {
  name: string;
  phone: string;
}

interface VehicleLocationResponse {
  vehicleId: string;
  lat: number;
  lng: number;
  speed: number;
  heading: number;
  status: "moving" | "idle" | "parked" | "offline" | "GPS_LOST";
  lastUpdated: string;
  driver?: DriverSummary;
}

interface DeliveryTrackingResponse {
  deliveryId: string;
  vehicleId: string;
  lat: number;
  lng: number;
  eta?: string;
  lastUpdated: string;
}

interface FleetOverviewResponse {
  fleetId: string;
  totalVehicles: number;
  moving: number;
  idle: number;
  parked: number;
  offline: number;
  alertsActive: number;
}

Send Location Update (MQTT)

MQTT Topic: vehicles/{vehicle_id}/location  QoS: 1
Payload schema: VehicleLocationUpdate (protobuf on the wire)
{
  "event_id": "evt-uuid",
  "vehicle_id": "v-uuid",
  "lat": 37.7749, "lng": -122.4194,
  "speed": 45.3, "heading": 180,
  "timestamp": 1710400000,
  "fuel_level": 0.65,
  "engine_status": "on",
  "events": ["harsh_brake"]
}

Get Vehicle Current Location

HTTP
GET /api/v1/vehicles/{vehicle_id}/location
Response: 200 OK
{
  "vehicle_id": "v-uuid",
  "lat": 37.7749, "lng": -122.4194,
  "speed": 45.3, "heading": 180,
  "status": "moving",
  "last_updated": "2025-03-14T10:23:45Z",
  "driver": {"name": "John", "phone": "+1..."}
}

Get Vehicle Route History

HTTP
GET /api/v1/vehicles/{vehicle_id}/trail?start=2025-03-14T08:00:00Z&end=2025-03-14T18:00:00Z&simplify=true

Get Customer Delivery Tracking

HTTP
GET /api/v1/deliveries/{delivery_id}/tracking
Response: 200 OK
{
  "delivery_id": "d-uuid",
  "vehicle_id": "v-uuid",
  "lat": 37.7749,
  "lng": -122.4194,
  "eta": "2026-09-23T18:45:00Z",
  "last_updated": "2026-09-23T18:32:10Z"
}

The tracking service combines the latest vehicle position with the ETA returned by ETA Calculation Service. A customer facing stream or WebSocket can push ETA changes without requiring the customer application to poll continuously.

Fleet Overview

HTTP
GET /api/v1/fleets/{fleet_id}/overview
Response: 200 OK
{
  "fleet_id": "f-uuid",
  "total_vehicles": 5000,
  "moving": 3200,
  "idle": 800,
  "parked": 700,
  "offline": 300,
  "alerts_active": 12
}

Common Error Responses

400 Bad Request: invalid input, missing required fields, or malformed JSON payload
401 Unauthorized: missing or invalid authentication token or API key
403 Forbidden: authenticated caller lacks required permissions for this resource
404 Not Found: requested resource ID does not exist
409 Conflict: duplicate write or version conflict, retry with a unique idempotency key
422 Unprocessable Entity: syntactically valid request failed semantic business validation
429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header
500 Internal Error: unexpected server failure, retry safely with an idempotency key
503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
440 Login Timeout: WebSocket connection session expired, client reconnect is required

Data Model

Redis: Latest Vehicle State

vehicle:{vehicle_id}  -> Hash { lat, lng, speed, heading, status, fleet_id, driver_id, fuel_level, last_updated, event_id }
TTL: 300 (stale telemetry guard)
vehicle_connectivity:{vehicle_id}  -> { state, last_heartbeat, session_epoch }
TTL: 300 (offline detection)
fleet_vehicles:{fleet_id}  -> SET of vehicle_ids
fleet_stats:{fleet_id}:moving  -> INT (Flink-maintained aggregate counter)
alerts:{vehicle_id}  -> LIST of active alert JSONs
dashboard_channel -> fleet:{fleet_id}:geo:{geohash_prefix}

TimescaleDB: Location Trail (90 Days)

SQL
CREATE TABLE location_points (
    event_id        UUID NOT NULL,
    vehicle_id      UUID NOT NULL,
    timestamp       TIMESTAMPTZ NOT NULL,
    lat             DOUBLE PRECISION,
    lng             DOUBLE PRECISION,
    speed           REAL,
    heading         SMALLINT,
    altitude        REAL,
    status          TEXT,
    odometer        REAL,
    fuel_level      REAL,
    events          JSONB,
    PRIMARY KEY (event_id, timestamp, vehicle_id)
);

CREATE INDEX idx_location_points_vehicle_time
    ON location_points (vehicle_id, timestamp DESC);

SELECT create_hypertable('location_points', 'timestamp',
    chunk_time_interval => INTERVAL '1 day',
    partitioning_column => 'vehicle_id',
    number_partitions => 4);

ALTER TABLE location_points SET (
    timescaledb.compress,
    timescaledb.compress_segmentby = 'vehicle_id',
    timescaledb.compress_orderby = 'timestamp DESC'
);

Kafka Topics

Topic: vehicle-locations  (512 partitions, 48h retention)
  Key: vehicle_id  Value: protobuf { lat, lng, speed, heading, timestamp, events[] }

Topic: vehicle-alerts  (64 partitions)
Topic: vehicle-status-changes  (64 partitions)

Fault Tolerance

ConcernSolution
MQTT broker failureCluster with session handoff and durable sessions. QoS 1 provides at least once delivery, and the bridge retries Kafka publication when necessary
Kafka lagConsumer group rebalancing and automatic Flink parallelism scaling
Redis failureRedis Cluster with replicas and AOF for fast recovery, followed by rebuilding latest state from the Kafka vehicle-locations topic if cache state is lost
TimescaleDB write failureKafka retains data for 48h, then replays the affected range after database recovery
Vehicle goes offlineUse MQTT session state or heartbeat TTL for connectivity detection, preserve the last known position, and distinguish offline from GPS_LOST when the MQTT session remains active
GPS drift/spoofingValidate speed against distance between points, timestamp monotonicity, and plausible movement. Flag or reject impossible readings

Handling Vehicle GPS Blackout (Tunnel)

Vehicle units buffer timestamps and points locally during signal outages. Upon reconnection, the tracker uploads the batch marked with a gap indicator. If gaps exceed 5 minutes, the server sets status to GPS_LOST while keeping the MQTT session active, displaying an advisory icon on the dashboard. Upon resume, the server can optionally interpolate the missing segment using known road geometry or a map matching service.

Out of Order GPS Points

In Redis, a Lua script evaluates timestamps to update stored state only if incoming readings are strictly newer than the cached timestamp. In TimescaleDB, the event_id plus timestamp and vehicle_id key rejects repeated inserts of the same telemetry event, while ORDER BY timestamp keeps history queries ordered. Flink streaming pipelines apply event time processing with bounded watermarks to evaluate speed violations accurately.

Additional Considerations

Interview Walkthrough

  • 25-minute cut

    Skip arch50 and arch75 depth unless interviewing for a staff level role.

    • Frame as a write heavy telemetry pipeline handling 2M GPS updates/sec at peak (5 min)
    • Ingestion flow: vehicles to MQTT brokers, Kafka buffering, and stream processors (6 min)
    • Hot path: Redis stores latest position per vehicle_id with short TTL (5 min)
    • Read path: dashboard clients subscribe via WebSockets to geohash sharded channels (5 min)
    • Staff topics: historical trails in ClickHouse, map tile overlays, and partition hot spots (4 min)
  • Frame the problem as a write heavy telemetry pipeline where 2M location updates/sec must not wait for TimescaleDB or alert processing before the vehicle receives its MQTT publish acknowledgement.
  • Walk through ingestion: vehicles connect to MQTT brokers, streaming into Kafka partitioned by vehicle_id, followed by Flink enrichment and geofence alerts.
  • Explain the hot path: Redis stores the latest position per vehicle using an atomic Lua script that discards out of order timestamps, while connectivity state distinguishes GPS_LOST from a disconnected vehicle.
  • Cover the read path: dashboard clients subscribe via WebSocket to a fan out layer reading from Redis Pub/Sub rather than querying TimescaleDB.
  • Describe tiered historical storage, keeping the latest 24 hours in the freshest raw form, applying TimescaleDB compression through 90 days, and generating ClickHouse aggregates for cold analytics.
  • Mention GPS blackout handling: buffer and replay on reconnect, interpolating gaps, and distinguishing GPS_LOST from true offline status.
  • Common pitfall: writing every GPS point synchronously to TimescaleDB on the ingestion path creates a storage bottleneck at 2M writes/sec and removes the buffering benefit of Kafka.

Engineering Trade-offs

Protocol and Storage Architecture Trade Offs

The platform balances network protocols, trajectory compression, and time series database tiering so it can support millions of connected vehicles economically. Review network transport trade offs in Network Protocols and compare with shared fleets in Bike-Sharing System.

MQTT vs WebSocket vs gRPC for Vehicle Communication

ProtocolOverheadBest For
MQTT ⭐2-4 bytes per messageIoT devices with constrained bandwidth
WebSocket2-14 bytes per frameBrowser based dashboards
gRPCHTTP/2 + protobufService to service, mobile apps

Decision: Vehicles to Server uses MQTT, Dashboard to Server uses WebSockets, and internal microservices communicate via gRPC.

Location Storage: Raw Points vs Compressed Trails

Start from the 17 TB/day raw point estimate. Douglas-Peucker simplification can reduce point count by 50%, with an illustrative 70% reduction on highways and 30% in dense city grids. Delta encoding can provide another 10 to 15 times compression on top of the time series storage format. The Hybrid ⭐ approach keeps the last 24 hours in the freshest raw form, uses TimescaleDB with 10 times compression from 1 to 90 days, and stores 90 or more days as ClickHouse aggregated segments. Under these assumptions, the total storage footprint is roughly 170 TB compared with 1.5 PB uncompressed.

Scalability: 10M Vehicles at 5-Second Intervals

A starting capacity plan for 2M updates/sec uses 50 MQTT broker nodes, 18 Kafka brokers with replication factor 3 and 512 partitions, 25 Flink TaskManagers, 20 Redis shards with 5 GB memory allocated per shard, 4 TimescaleDB shards, and 10 dashboard WebSocket servers. The 5 GB per Redis shard allocation provides substantial headroom above the roughly 2 GB logical payload estimate. Treat these as initial sizing assumptions and validate them with load tests.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...