Interview Setup
Interview Prompt
Design Google Calendar for 1B users: 100K calendar views/sec, 50K free/busy queries/sec, 10K event creates/sec, and 50B stored events (500 rules/user average with RRULE expansion). Cross-user scheduling, room booking, and reminder fan-out are in scope.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| Expand recurring events on write or on read? | A workload of 500 events per user with a weekly standup over two years generates 104 instances, meaning expand-on-read keeps a single row per series rather than 50B materialized rows. |
| Free/busy query: exact intervals or daily bitmask? | Handling 50K free/busy queries per second across five attendees requires 250K calendar reads without caching, whereas a daily bitmask (32 bits for 15-minute slots) delivers O(1) bitwise AND queries across attendees. |
| Sharing model: per-calendar ACL or per-event? | A per-calendar ACL is straightforward for shared team calendars, whereas per-event controls are necessary for busy-only private sharing. |
| How are reminders delivered at scale? | One billion users with two reminders per event in a peak morning window produce millions of push notifications per minute, requiring scheduled job fan-out rather than inline dispatch on create. |
| Timezone handling: store UTC or local time? | Recurring events like 9 AM every Monday in America/New_York shift their UTC offset during Daylight Saving Time, requiring the database to store the IANA timezone and expand instances using up-to-date tzdata. |
Scope
In scope
- Event CRUD with RRULE recurrence and exception handling (EXDATE, edit-this, and edit-following)
- Free/busy merge across N attendees with privacy filters (busy-only vs full detail)
- Room and resource booking with exclusion constraints to prevent double-booking
- Reminder scheduling and multi-channel notification fan-out
- CalDAV sync, timezone and DST correctness, and sharding by calendar_id
- Capacity estimation with explicit sizing calculations
Functional Requirements
Start by clarifying whether external calendar synchronization (via CalDAV and iCal) and conference room bookings are in scope, because they require exclusion constraints and cross-attendee free/busy queries. Verify RRULE recurrence support and the sharing privacy model (busy-only vs full detail) before detailing notification channels.
- Create, update, delete events: Title, time, location, description, and recurrence rules.
- Recurring events: Daily, weekly, monthly, yearly, and custom schedules matching RFC 5545 RRULE.
- Invite attendees: Send meeting invitations and track RSVP status.
- Free/busy lookup: Check availability across multiple people and calendars.
- Calendar sharing: Grant view-only, busy-only, or full edit access across organizations.
- Reminders and notifications: Push notifications and emails dispatched before events begin.
- Time zone handling: Store events in UTC while rendering in the viewing user's local timezone.
- Room and resource booking: Find and reserve conference rooms with guaranteed no-double-booking.
- Calendar synchronization: Provide CalDAV and iCal standards for external clients.
- Multiple calendars per user: Support separate personal, work, holiday, and secondary calendars.
Non-Functional Requirements
Stress-test conference room booking consistency and free/busy query latency across multiple attendees. PostgreSQL exclusion constraints avoid race conditions and outperform application-level optimistic locking for shared physical resources. Address DST edge cases and ensure reminder fan-out maintains strict time delivery under morning usage spikes.
- Consistency: Zero double-booking of rooms or conflicting exclusive physical resources.
- Availability: 99.99% uptime for the primary read and calendar viewing path.
- Low Latency: Calendar view loads under 200ms, and multi-attendee free/busy queries complete under 500ms.
- Sync Latency: Changes propagate to all connected client devices within 5 seconds.
- Scalability: Support 1B+ registered users and 50B+ stored event definitions.
Capacity Estimations
Recurring events cause severe storage explosion if every occurrence is materialized ahead of time. Storing RRULE rules and expanding them on read within bounded windows keeps storage manageable, while notification fan-out for 10B events dictates asynchronous queue sizing.
| Metric | Calculation | Value |
|---|---|---|
| Users | Given | 1B |
| Events per user (next 12 months) | Given | 500 (including recurring expansions) |
| Total events | Given | 500B (with recurring expansions) |
| Stored events (without expansion) | Given | 50B |
| Calendar views / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 100K |
| Event creates / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 10K |
| Free/busy queries / sec | Derived from daily volume ÷ 86400 (+ peak factor) | 50K |
Architecture Diagram
Separate event CRUD from free/busy aggregation and reminder delivery, as each exhibits distinct read and write patterns with different latency budgets. Event writes route through PostgreSQL with timezone-aware RRULE storage, while reads expand recurrence rules on demand for requested date windows.
The free/busy service merges busy intervals across N attendees using sweep-line algorithms or daily bitmasks in Redis, avoiding N+1 calendar database fetches at 50K queries per second. Conference room reservations leverage PostgreSQL exclusion constraints so overlapping bookings are rejected at the database engine level.
Reminder delivery executes asynchronously: on event creation, reminder tasks are enqueued to Kafka with fire-time offsets, allowing background workers to fan out push notifications and emails without slowing the synchronous create response path.
In the room
Clarify whether recurring events expand on read or write: storing RRULE rules and expanding at query time keeps 50B materialized instances out of your persistent database.
Component Deep Dives
We analyze recurrence storage and expansion first, followed by the free/busy merge algorithm, room booking concurrency, and timezone handling under daylight saving time shifts.
Recurring Events: Storage vs Expansion
Recurrence storage is a primary architectural decision: storing a single RRULE row avoids materializing 730 instances of a daily standup, but exception handling adds complexity during query expansion.
Wrong approach: Store every occurrence of "daily standup for 2 years" = 730 rows Right approach: Store ONE row with recurrence rule (RRULE) RRULE:FREQ=WEEKLY;BYDAY=MO,WE,FR;UNTIL=20261231 -> Expands at query time to show individual occurrences Exception handling: - "Edit this occurrence": store exception (overrides that single instance) - "Edit this and following": split into two recurrence rules - "Delete this occurrence": store exclusion date (EXDATE) Storage: 1 event row + exceptions Query: expand RRULE in [view_start, view_end] range at read time Cache: pre-expand next 30 days in Redis for fast calendar view
Free/Busy Query (Key Algorithm)
The free/busy scheduling algorithm merges busy intervals across attendees and inverts them to find mutual gaps, while precomputed daily bitmasks make this an O(1) bitwise operation at scale.
Input: "Find when Alice, Bob, and Carol are all free next Tuesday 9am-5pm" 1. Fetch all events for each person on Tuesday (from cache or DB) 2. Build busy intervals per person: Alice: [9:00-10:00, 11:00-12:00, 14:00-15:00] Bob: [9:30-10:30, 13:00-14:00] Carol: [10:00-11:30] 3. Merge all busy intervals -> union: [9:00-12:00, 13:00-15:00] 4. Invert -> free slots: [12:00-13:00, 15:00-17:00] 5. Filter by desired meeting duration Optimization: Pre-compute daily free/busy bitmask (one bit per 15-min slot) 8 hours x 4 slots/hr = 32 bits per day per person AND all bitmasks -> free slots in O(1) bitwise operation
Room Booking: PostgreSQL Exclusion Constraint
Room booking enforces a strict invariant where exclusion constraints at the database level prevent overlapping reservations, even under concurrent transactions.
-- Atomic room reservation (prevent double-booking)
BEGIN;
SELECT event_id FROM events
WHERE room_id = 'conf-room-A'
AND date = '2026-03-14'
AND (start_time < '11:00' AND end_time > '10:00')
FOR UPDATE;
-- If no rows returned, room is free
-- (FOR UPDATE cannot be used with COUNT/aggregates in PostgreSQL)
INSERT INTO events (room_id, start_time, end_time) VALUES ('conf-room-A', '2026-03-14 10:00:00+00', '2026-03-14 11:00:00+00');
COMMIT;
-- Alternative: Exclusion constraint (database-level guarantee)
-- EXCLUDE USING gist (room_id WITH =, tsrange(start_time, end_time) WITH &&)
-- Prevents overlapping time ranges for same room automaticallyTime Zone + DST Deep Dive
Daylight saving time shifts UTC offsets twice a year for recurring events, requiring the system to store IANA timezones and expand them dynamically with the latest tzdata rather than storing static UTC offsets.
Problem: "9 AM every Monday" in New York
Summer (EDT, UTC-4): 9 AM local = 13:00 UTC
Winter (EST, UTC-5): 9 AM local = 14:00 UTC
Storage: Store RRULE + original timezone ("America/New_York")
Expansion: At query time, expand RRULE using IANA tz database
NEVER store computed UTC offsets in the RRULE itself
Offsets change when DST rules change (governments update DST dates)
Must re-expand using latest IANA tzdata at render timeCalDAV Sync Protocol
CalDAV synchronization allows external clients to stay in sync via delta synchronization tokens, avoiding the overhead of re-downloading entire calendars on every poll.
When Client A edits an event, the server persists the update and increments the calendar version. Client B sends a GET request with its current sync-token, prompting the server to return only the changes committed since that token. This efficient delta synchronization transfers minimal diffs, resolving conflicts using last-write-wins based on server timestamps.
Event Bus Design (Kafka)
Kafka propagates calendar changes asynchronously to invalidate free/busy caches, trigger notifications, and push external sync updates off the critical event creation path.
Topic: calendar-changes Partitions: 64 (partition by calendar_id) Events: event_created, event_updated, event_deleted, rsvp_changed, recurrence_exception Retention: 7 days Producers: Event Service after PostgreSQL commit (EXCLUDE constraint validated) Consumers: Free/Busy cache invalidator, Notification Service, CalDAV sync push Topic: reminder-dispatch Partitions: 32 (partition by user_id) Events: reminder_due (15m / 1h / 1d before start) Consumers: Notification workers (push, email, SMS) Topic: free-busy-invalidation Fan-out when attendee accepts/declines on shared calendar Consumers: per-shard bitmask cache rebuild (Redis) Event CRUD path: validate TZ/DST -> PG transaction -> publish calendar-changes -> 201 Free/busy queries read Redis bitmask; invalidation triggered by calendar-changes consumer
API Design
Present a clean REST interface for application clients along with a dedicated free/busy endpoint. Avoid folding availability into the general event list to prevent N+1 database queries on every page view.
Calendar Endpoints
POST /api/events -> Create event
GET /api/events?start=...&end=...&calendar_id=... -> List events in range
PUT /api/events/{id} -> Update event (this/all/following)
DELETE /api/events/{id} -> Delete event (this/all/following)
POST /api/events/{id}/rsvp -> Accept/decline/tentative
GET /api/freebusy -> Query free/busy for list of users
POST /api/rooms/search -> Find available rooms for time slot
GET /api/calendars -> List user's calendars
POST /api/calendars/{id}/share -> Share calendar with user/groupCommon Error Responses
400 Bad Request: invalid input, missing required fields, or malformed JSON payload 401 Unauthorized: missing or invalid authentication token or API key 403 Forbidden: authenticated caller lacks required permissions for this resource 404 Not Found: requested resource ID does not exist 409 Conflict: duplicate write or version conflict, retry with a unique idempotency key 422 Unprocessable Entity: syntactically valid request failed semantic business validation 429 Too Many Requests: rate limit quota exceeded, client should honor Retry-After header 500 Internal Error: unexpected server failure, retry safely with an idempotency key 503 Service Unavailable: downstream dependency is unavailable or overloaded, retry with exponential backoff
Data Model
PostgreSQL Schema
CREATE TABLE events (
event_id UUID PRIMARY KEY,
calendar_id UUID NOT NULL,
creator_id UUID NOT NULL,
title TEXT,
description TEXT,
location TEXT,
start_time TIMESTAMPTZ NOT NULL,
end_time TIMESTAMPTZ NOT NULL,
timezone TEXT DEFAULT 'UTC',
is_all_day BOOLEAN DEFAULT FALSE,
recurrence TEXT,
room_id UUID,
status TEXT DEFAULT 'confirmed',
EXCLUDE USING gist (room_id WITH =, tsrange(start_time, end_time) WITH &&)
WHERE (room_id IS NOT NULL)
);
CREATE TABLE event_attendees (
event_id UUID REFERENCES events(event_id),
user_id UUID,
rsvp_status TEXT DEFAULT 'needs-action',
PRIMARY KEY (event_id, user_id)
);Redis Key Structure (Free/Busy & RRULE Expansion Caches)
freebusy:{user_id}:{date} -> 32-bit bitmask
calendar:{user_id}:{month} -> expanded events for monthFault Tolerance
Room Double-Booking Prevention
PostgreSQL exclusion constraints provide an ironclad, database-level guarantee. Even when two concurrent transactions attempt to book overlapping room slots simultaneously, the GiST index constraint immediately blocks the conflicting write.
Additional Considerations
System Comparisons & Cross-links
- Ticket Booking System: Whereas ticket booking manages inventory of discrete seats or tickets, a shared calendar manages dynamic, continuous time intervals where recurring event series are first-class entities.
- Notification System: The calendar functions as the authoritative upstream event source that schedules time-offset notification dispatches, delegating multi-channel fan-out and retry logic to the notification platform.
- Distributed Job Scheduler: While both systems parse recurring rules (RRULE vs cron expressions), a calendar system serves interactive, user-facing scheduling with real-time free/busy evaluations and timezone-aware DST shifts.
Key Architectural Nuances
- RRULE Complexity: Full RFC 5545 specification support involves complex edge cases with custom recurrence intervals, month-end rollovers, and multi-day exclusions (EXDATE).
- Time Zone Shifting: A recurring meeting set for 9 AM in New York maps to different UTC hours across the year due to Daylight Saving Time, making IANA timezone expansion mandatory.
- Cross-Organization Privacy: Federated free/busy lookups across corporate tenants must expose availability blocks while withholding private titles, attendee rosters, and descriptions.
- Physical Resource Invariants: High-contention conference rooms rely on database exclusion constraints to maintain strict single-occupancy guarantees.
Interview Walkthrough
- 25-minute interview progression
Focus on the core algorithmic trade-offs before diving into edge-case details.
- Expand recurring events on read vs write with RRULE caching (8 min)
- Free/busy aggregation via interval sweep-line and Redis bitmasks (9 min)
- Sharding topology by calendar_id with room booking isolation (8 min)
- Frame scheduling conflicts as interval overlap detection using database exclusion constraints or application-level locking.
- Explain free/busy aggregation across multiple calendars while preserving privacy through busy-only views.
- Address timezone handling: store timestamps in UTC, store original IANA timezones, and expand on display to handle DST transitions cleanly.
- Discuss recurring event expansion as lazy read-time generation backed by 30-day pre-computed Redis caches.
- Walk through meeting invite flows with attendee RSVP state machines and asynchronous notification pipelines.
- Highlight the common pitfall of naive local datetime comparisons across timezones without proper UTC normalization.
Engineering Trade-offs
Scalability: Sharding Calendar Data
Shard by calendar_id (~ user_id for primary calendar):
events table: shard by calendar_id
1B users / 256 shards = ~4M users per shard
Single-user operations: single shard query (direct lookup)
Cross-shard challenges:
1. Free/busy query for 10 attendees across 10 shards:
Solution: Pre-computed free/busy bitmasks in Redis
Key: freebusy:{user_id}:{date} -> 32-bit bitmask
No cross-shard DB queries needed
2. Shared calendar: Kafka topic calendar-changes -> each shard's consumer
Eventual consistency: updates visible within 2-5 seconds
3. Room booking: Separate rooms shard with exclusion constraintsOffline Sync Conflict Resolution
Strategy: Field-level merge with last-write-wins per field
Example:
Phone edit (offline, T=10:00): Changed title to "Team Standup v2"
Laptop edit (online, T=10:05): Changed time to 10:30 AM
Phone reconnects at T=10:15:
- title: phone="Team Standup v2" (T=10:00), server is older
-> Accept phone's title
- time: phone=10:00 AM (unchanged), server=10:30 AM (T=10:05)
-> Keep server's time
Same-field conflict: Last-write-wins by timestamp
Deletion conflict: Option A (delete wins) vs Option B (preserve with changes)
Recurring event conflict: Apply bulk changes first, then single exceptionsRRULE Expansion at Read-Time vs Materialized
Option 1: Expand at read time (recommended) Store: 1 row with RRULE ✓ Storage efficient: 1 row instead of 365+ rows per recurring event ✓ "Edit all future" is a single row update ✗ CPU cost at read time (~1ms per event) Mitigation: Redis cache of expanded events for next 30 days Option 2: Materialized occurrences Store: 365 rows for "daily standup for 1 year" ✓ Simple queries (no expansion logic) ✗ Storage explosion: 500B+ rows for all users ✗ "Edit all future" requires updating 200+ rows atomically Recommendation: Expand at read time + aggressive caching
Bitmask Granularity: 15-min vs Per-Minute
Fifteen-minute slots consume only 32 bits per day per user (amounting to 1.46 TB across 1B users for an entire year). Per-minute granularity demands 15x greater memory. Standard meeting intervals naturally align with 15-minute boundaries, making 15-minute bitmasks the optimal choice for production calendar platforms.
Exclusion Constraint vs Application-Level Locking
PostgreSQL exclusion constraints enforce strict reservation invariants at the database engine level, eliminating race conditions from concurrent transactions. While application-level locks provide custom error messaging flexibility, they introduce vulnerability to bugs and lock lease timeouts. Production architectures adopt a defense-in-depth approach by combining application validation with database exclusion constraints as the ultimate safety net.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.