Interview Setup
Interview Prompt
Design GitHub for 200M repositories, 100M developers, ~35K git read ops/sec peak, 100+ PB of Git object storage, and 50M open pull requests, separating web metadata from the Git RPC storage tier.
Clarifying Questions (ask before designing)
| Question | Why it matters |
|---|---|
| What scale are we designing for? | Millions of repos require a distributed storage tier, not just a single NFS mount. |
| Are we focusing on Git operations or the web UI? | Git operations such as push and pull are I/O and network intensive. the Web UI is a standard CRUD application. |
| How large are the repositories? | Large repos require delta compression, shallow clones, and specialized Git backend tuning. |
Scope
In scope
- Git Push and Pull operations over SSH/HTTPS
- Distributed storage of Git repositories
- Web interface for browsing code and commits
- High availability and disaster recovery
Out of scope (state explicitly)
- CI/CD pipelines and runners
- Advanced project management boards and workflow automation
- Detailed GitHub Actions implementation
Functional Requirements
Git hosting only: Gitaly storage, pull requests, and web UI. Full CI/CD platform capabilities (runners, DAG execution, and build artifacts) are covered in CI/CD Platform (GitLab CI / GitHub Actions).
Start by asking your interviewer whether CI/CD is in scope. For this discussion it is treated as a separate system, as covered in our CI/CD Platform design. Confirm monorepo scale, fork sharing requirements, and whether Git LFS is needed before splitting web metadata from Git RPC storage.
- Git Operations: Support
git push,git pull, andgit cloneover SSH and HTTPS. - Web Interface: Users can browse the repository file tree, view commit history, and read file contents.
- Pull Requests: Users can create PRs, diff code, and leave line-by-line comments.
- Issues & Collaboration: Issue tracking, starring, commenting, and forking repositories.
- Large File Support: Support Git LFS for large binaries so repository object storage is not dominated by frequently changing media files.
- Access Control: Enforce repository visibility, membership roles, branch protection, and transport authentication across both the web tier and Git RPC tier.
Non-Functional Requirements
Your interviewer will stress test the roughly 20 to 1 read to write ratio and whether packfile generation survives viral clone traffic. They will also probe why NFS failed and how Gitaly RPC on local NVMe moves Git traversal away from network mounted storage.
- High Availability: Source code should remain accessible to Git clients and downstream automation during individual node and region failures.
- High Read Throughput:
git cloneandgit fetch(reads) vastly outnumbergit push(writes). - Storage Efficiency: Repositories can grow to tens of gigabytes. Storage must be deduplicated and compressed.
- Data Integrity: Zero tolerance for source code corruption, with object verification, repository repair workflows, and integrity checks during replication.
- Security: Authenticate Git transport and web requests, authorize repository actions, and isolate private repository data from public cache paths.
Capacity Estimations
Size the Git RPC tier for peak clone traffic rather than average volume. A viral open source release can increase packfile generation by 100 fold while the push rate remains steady.
| Metric | Calculation | Value |
|---|---|---|
| Repositories | Given | 200M |
| Developers (DAU) | Given | 100M |
| Git fetch/clone ops/day | Given | 1B |
| Peak git reads/sec | 1B ÷ 86400 x 3 (peak factor) | ~35K reads/s |
| Git push ops/day | ~5% of reads | 50M |
| Peak git writes/sec | 50M ÷ 86400 x 3 | ~1.7K writes/s |
| Logical Git object storage | Given | 100+ PB deduplicated |
| Replicated physical storage | 100+ PB x 3 full replicas | ~300+ PB before other overhead |
| Open PRs | Given | 50M active |
| Web UI page views/sec | Given | 500K |
The read to write ratio is approximately 20 to 1 for Git operations. A viral open source release can spike clone traffic 100 fold. Because packfile generation can be CPU intensive and may take 500ms to 2s per clone, cache precomputed packfiles for popular public references and common clone profiles, then serve those immutable responses from the edge. Fork deduplication through object pools means 10K forks of a 3 GB repository can share the common object history, so storage growth is dominated by new objects and deltas rather than 30 TB of full copies.
Architecture Diagram
GitHub is two products sharing a URL: a web application for metadata such as pull requests, issues, and permissions, and a Git RPC storage tier for content addressed object graphs. Push and pull are I/O intensive protocol operations that should not run through a shared NFS mount. Instead, the storage node executes native Git operations against local NVMe and returns aggregated results over RPC.
Read traffic dominates at roughly a 20 to 1 clone to push ratio, and a viral open source release can increase packfile generation by 100 fold. The router maps each repository to a storage shard with three synchronous replicas inside the primary failure domain. PostgreSQL holds relational metadata and routing pointers, while Kafka distributes durable repository events to webhooks, search, pull request workflows, and CI integrations.
In the room
Split the web tier (PostgreSQL metadata, PRs) from the Git storage tier (Gitaly RPC on local NVMe) in your first sketch. Serving git from shared NFS is the classic pitfall, because random I/O over the network degrades clone latency at scale.
Component Deep Dives
1. Understanding Git Internals (Content-Addressable Storage)
The interview narrative moves from Git internals such as blobs, trees, commits, and refs through the NFS to RPC storage evolution, sharded routing, and the pull request merge path. Trial merges and diff caching prevent expensive recomputation on every page view across 50M open pull requests.
Git uses content addressed objects for blobs, trees, and commits, while refs are mutable pointers to object IDs. These objects form a directed acyclic graph that connects repository history, as discussed in Merkle Tree architectures.
Git uses a content addressed object database. Git objects are identified by SHA-1 or SHA-256 object IDs and stored as immutable objects whose identities are derived from their contents.
- Blob: The raw binary contents of a file. Filenames are NOT stored here.
- Tree: A directory listing. It maps filenames and permissions to Blob SHAs or other sub-Tree SHAs.
- Commit: Immutable metadata (author, message, timestamp, parent commit SHA) and a pointer to the root Tree SHA representing the state of the repository at that exact moment.
- Ref (Branch/Tag): A mutable, human-readable pointer (for example,
refs/heads/main) pointing to a specific Commit SHA. Pushing to a branch just updates this text file pointer.
2. The Storage Layer Evolution (NFS vs RPC) ⭐
The NFS to RPC evolution represents a core architectural pivot: moving the Git binary directly to the storage node ensures that random DAG traversals hit local NVMe rather than contending over network mounted disks, aligning with modern Block, File, and Object Storage design.
Originally, GitHub and GitLab stored raw .git folders on massive Network File Systems (NFS/NAS). Stateless Web workers would mount the NFS drive over the network and run Git binaries locally.
The Bottleneck: Git commands such as git log can perform many metadata and object accesses while traversing repository history. When those accesses cross a network file system boundary, latency and shared IOPS contention can become dominant bottlenecks for large repositories.
The RPC Solution (Gitaly / DGit):Move Git operations to the storage node. Instead of mounting a repository over the network, the Web tier sends an RPC request to the storage service. The storage service runs native Git operations against local SSD or NVMe and returns an aggregated result, removing per file network IOPS from the Git execution path.
3. Routing and Sharding (Git Router)
Repositories shard across many storage nodes, with a Git Router mapping each repository URL to its storage shard through a cached routing table. The routing decision also carries the current replica and fencing information needed to protect the write path.
A single massive cluster cannot hold all repositories. Repositories must be sharded across thousands of Storage Nodes.
- When a request arrives for
torvalds/linux.git, the Git Router acts as an L7 reverse proxy. - It looks up the repository location in a fast routing database (Redis or a highly cached PostgreSQL).
- It discovers that
linux.gitis hosted onStorage Node 402and seamlessly proxies the SSH/HTTPS connection or gRPC call there.
4. Pull Requests, Merge Conflicts, and Diff Caching
Pull request pages are read heavy, where asynchronous trial merges and diff caching by (base_sha, head_sha) pairs prevent expensive recomputation on every page view across 50M open pull requests. Learn more about effective caching tiers in Caching Patterns and Invalidation.
PR pages are read heavy, while merge computation can be reused for a given (base_sha, head_sha) pair. A background worker runs trial merges when the PR opens or the base branch advances. It caches clean diffs in Redis and records merge test state so the web UI does not block on git merge-tree.
PR merge conflict resolution: 1. User opens PR: feature-branch → main 2. Background worker runs trial merge: git merge-base main feature-branch git merge-tree $(merge-base) main feature-branch 3. If conflict: - Persist merge status and conflict metadata in the PR record. - Optionally store a synthetic merge result in a hidden ref for fast inspection. - UI shows "Resolve conflicts" with a 3-way diff (base, ours, theirs). - User resolves in the web editor or locally with git fetch and git merge main. 4. If clean merge: - Cache diff HTML in Redis with a key derived from (base_sha, head_sha). - Store the tested merge result metadata so the UI can render it without recomputing the merge on every page view. 5. On "Merge" click: - Branch protection checks run against the current base and head state. - The storage tier performs a compare-and-swap ref update. A fast-forward or merge commit is created as required. - A durable repository event triggers webhooks and downstream CI integrations asynchronously. 6. Scale: diff computation is proportional to the changed-file set in the chosen implementation. Cache by (base_sha, head_sha) pair.
5. Forking via Shared Object Pools
Forks can share repository objects through object pools and Git alternates, making fork creation mostly a metadata operation. Storage growth is then driven by new objects rather than duplicating the entire shared history.
If 10,000 users fork an open source repository (for example, the Linux kernel at ~3GB), duplicating the repository 10,000 times would waste 30 Terabytes of expensive SSD storage.
- Git Alternates / Object Pools: When you click "Fork", the new repository can reference a shared object pool instead of copying the existing Git objects. The alternate object store is referenced through
objects/info/alternates, and the pool lifecycle must ensure referenced objects remain available. - When a user pushes new content to their fork, only newly introduced Git objects need to be stored separately. Existing shared objects remain in the object pool.
- This deduplication makes fork creation an O(1) metadata operation in the common case and keeps incremental storage close to newly introduced objects.
6. Read Replica Offloading for Clone Traffic
Push traffic hits the current primary for linearizable ref updates. Clone and fetch requests can load balance across healthy secondaries when their replica lag is within the read policy. Read after write requests can be pinned to the primary until the secondary has caught up to the required repository version. This is critical during CI herds where thousands of runners clone the same commit simultaneously.
7. Git LFS and Large Binary Objects
Large binaries should not dominate the normal Git object graph. Git LFS stores lightweight pointer files in the repository and keeps the large payloads in dedicated object storage. The web tier authenticates access and applies repository or organization quotas, while the Git storage tier continues to optimize source code history.
8. Repository Integrity and Maintenance
Storage workers periodically verify object reachability and pack integrity. Repository maintenance can build commit graphs, reachability bitmaps, and multi pack indexes to reduce traversal and pack generation costs. Garbage collection runs only after reachability and replication checks so objects required by another healthy replica are not deleted prematurely.
API Design
Git Clone / Fetch (Read)
Standard REST APIs serve pull requests and issues, while Git code transfer uses Git Smart HTTP or SSH so the protocol can negotiate and transfer only the object set required by the client.
git clone https://github.com/user/repo.git 1. Client sends GET /user/repo.git/info/refs?service=git-upload-pack - Router authenticates the request and maps the repository to its storage shard. - Storage Node advertises references and capabilities for the repository. 2. Client sends POST /user/repo.git/git-upload-pack - Client negotiates the objects it already has with "want <commit-sha>" and "have <commit-sha>" lines. - Storage Node traverses the reachable object graph and generates a packfile containing the negotiated objects. - The packfile is streamed back to the client over the Git Smart HTTP connection. For SSH, the same upload-pack operation runs over the authenticated SSH channel rather than Smart HTTP.
Git Push (Write)
git push origin main 1. Client sends the git-receive-pack request. - Client sends a stream containing the new commits, trees, and blobs in a packfile, together with the requested ref update. - Storage Node validates the received objects in a quarantine area, verifies object and ref integrity, and checks branch protection policy before publishing the ref update. - The branch ref update uses the expected old object ID as a compare-and-swap condition so two concurrent pushes cannot silently overwrite each other. 2. The authoritative ref update commits to the repository write quorum. - A durable push event is recorded after the ref update and before downstream delivery is considered complete. - Kafka consumers asynchronously trigger webhooks, pull request updates, and CI integrations. The push acknowledgment is tied to the authoritative repository commit point after the required write quorum is satisfied. Asynchronous consumers then propagate the change to secondary systems.
Web Metadata API Contracts
Typed domain contracts keep repository metadata operations separate from the binary Git transport.
type RepositoryId = string & { readonly __brand: "RepositoryId" };
type PullRequestId = string & { readonly __brand: "PullRequestId" };
type RefName = string & { readonly __brand: "RefName" };
type ObjectId = string & { readonly __brand: "ObjectId" };
type RepositoryVisibility = "public" | "private" | "internal";
type PullRequestStatus = "OPEN" | "MERGED" | "CLOSED";
interface RepositorySummary {
repositoryId: RepositoryId;
ownerId: string;
name: string;
defaultBranch: RefName;
visibility: RepositoryVisibility;
}
interface CreatePullRequestRequest {
repositoryId: RepositoryId;
baseRepositoryId: RepositoryId;
baseRef: RefName;
headRepositoryId: RepositoryId;
headRef: RefName;
title: string;
description?: string;
}
interface MergePullRequestRequest {
pullRequestId: PullRequestId;
expectedBaseObjectId: ObjectId;
expectedHeadObjectId: ObjectId;
mergeMethod: "merge" | "squash" | "rebase";
}
interface PullRequestResponse {
pullRequestId: PullRequestId;
status: PullRequestStatus;
mergeable: boolean;
baseObjectId: ObjectId;
headObjectId: ObjectId;
}
interface CodeHostingMetadataService {
getRepository(repositoryId: RepositoryId): Promise<RepositorySummary>;
createPullRequest(request: CreatePullRequestRequest): Promise<PullRequestResponse>;
mergePullRequest(request: MergePullRequestRequest): Promise<PullRequestResponse>;
}Data Model
Git Object Data and Relational Metadata
A code hosting platform separates immutable Git object data in the storage tier from relational metadata in PostgreSQL. Refs, memberships, pull requests, and branch policy remain metadata even though they determine which Git objects are reachable.
// 1. Git Object Storage (On Disk / Gitaly)
// Git content is represented by immutable, content addressed objects.
// Object formats can use SHA-1 or SHA-256 object IDs depending on repository format.
Blob {
length: size_t
content: byte[] // The actual file contents
}
Tree {
entries: [
{ mode: 100644, type: "blob", sha: "abc1234...", name: "index.js" },
{ mode: 040000, type: "tree", sha: "def5678...", name: "src" }
]
}
Commit {
tree_sha: "xyz9876..."
parent_shas: ["prev123..."]
author: "Ronnie <ronnie@example.com>"
message: "Fix bug"
}
// Refs are mutable metadata pointers, not Git objects.
// Example: refs/heads/main -> commit object ID
// 2. Relational Metadata (PostgreSQL)
CREATE TABLE repositories (
id BIGSERIAL PRIMARY KEY,
owner_id UUID REFERENCES users(id),
name VARCHAR(255) NOT NULL,
storage_shard_id VARCHAR(50) NOT NULL,
default_branch VARCHAR(255) NOT NULL,
visibility VARCHAR(20) NOT NULL CHECK (visibility IN ('public', 'private', 'internal'))
);
-- Authoritative Git refs remain in the repository storage tier.
-- PostgreSQL may cache selected head object IDs for web metadata, but ref updates are committed by the Git storage leader.
CREATE TABLE repository_memberships (
repo_id BIGINT REFERENCES repositories(id),
user_id UUID REFERENCES users(id),
role VARCHAR(32) NOT NULL,
PRIMARY KEY (repo_id, user_id)
);
CREATE TABLE pull_requests (
id BIGSERIAL PRIMARY KEY,
head_repo_id BIGINT REFERENCES repositories(id),
base_repo_id BIGINT REFERENCES repositories(id),
base_branch VARCHAR(255) NOT NULL,
head_branch VARCHAR(255) NOT NULL,
status VARCHAR(20) NOT NULL CHECK (status IN ('OPEN', 'MERGED', 'CLOSED')),
base_object_id VARCHAR(64) NOT NULL,
head_object_id VARCHAR(64) NOT NULL
);
CREATE TABLE branch_protection_rules (
id BIGSERIAL PRIMARY KEY,
repo_id BIGINT REFERENCES repositories(id),
branch_pattern VARCHAR(255) NOT NULL,
require_reviews BOOLEAN NOT NULL DEFAULT FALSE,
require_status_checks BOOLEAN NOT NULL DEFAULT FALSE
);
CREATE TABLE issues (
id BIGSERIAL PRIMARY KEY,
repo_id BIGINT REFERENCES repositories(id),
author_id UUID REFERENCES users(id),
title VARCHAR(512) NOT NULL,
status VARCHAR(20) NOT NULL CHECK (status IN ('OPEN', 'CLOSED')),
created_at TIMESTAMP NOT NULL DEFAULT NOW()
);Fault Tolerance
| Failure Case | System Solution Design |
|---|---|
| Storage Node Failure | Each repository has three synchronous replicas inside its primary failure domain. A fenced leader owns ref updates, and a healthy replica can be promoted after it has confirmed the committed repository state. |
| Network Partition | A primary that loses quorum is fenced from accepting writes. Pushes temporarily fail until a new leader is elected with a current fencing epoch, preventing split brain writes. |
| Heavy CI Clone Traffic | Clone and fetch traffic load-balance across healthy secondaries when their replica lag satisfies the read policy. The primary remains available for pushes and ref updates. |
| Silent Object Corruption | Background integrity checks compare repository object reachability and checksums. Corrupt replicas are quarantined and repaired from a known good replica before re-entering service. |
Additional Considerations
Git LFS and Large Binary Objects
Git LFS keeps large binary payloads out of ordinary Git object transfer. The Git repository stores small pointer files, while large objects live in dedicated object storage behind authenticated URLs. This protects clone performance and keeps repository object maintenance focused on source code.
Repository Integrity and Maintenance
Storage workers periodically verify object reachability, pack integrity, and replica checksums. Garbage collection runs only after reachability analysis and replication safety checks, so unreachable objects are not deleted while another replica still needs them for recovery.
Security and Access Control
The Git transport must authenticate users through SSH keys or HTTPS credentials and authorize access against repository visibility and membership state before routing a request to storage. Private repository packfiles must not enter public edge caches, and authorization must be rechecked when a cached response could expose protected content.
Interview Walkthrough
- 25-minute cut
Skip deep architectural variants unless targeting staff level.
- Git RPC architecture vs network file systems (8 min)
- Git router routing and repository sharding (9 min)
- Relational metadata in PostgreSQL and object storage on NVMe (8 min)
- Split the web tier (PostgreSQL metadata, PRs, and issues) from the Git storage tier (Gitaly RPC on local NVMe).
- Explain content addressed storage: blobs point to trees, trees point to commits, and refs act as mutable pointers updated on push.
- Why NFS failed: random I/O over the network during Git log DAG traversal caused severe IOPS starvation. Moving the Git binary to the storage node resolved it.
- Router maps repository URLs to storage shards. The replication layer commits ref updates to a quorum before acknowledging successful pushes.
- Forks leverage object pools and alternates for instant O(1) creation where storage costs reflect deltas only.
- PR diffs: run trial merges in the background, cache diff outputs by (base_sha, head_sha) pair, and never recompute on individual page views.
- Common pitfall: attempting to serve Git repositories over shared NFS mounts, where network IOPS starvation destroys clone and fetch latency at scale.
Engineering Trade-offs
Strong Consistency vs. Eventual Consistency
The authoritative repository ref state must be strongly consistent within the write domain. If a developer pushes a commit to main and a CI server immediately clones the repository 100 milliseconds later, the read should observe the committed ref when the product promises read after write semantics. That does not require every geographic replica to be synchronous. The design can commit the ref update to a local quorum using a consensus protocol, then route immediate reads to the primary or a replica that has confirmed the required version. Cross region replicas can remain asynchronous with an explicit recovery point objective.
Monorepo vs. Multirepo Architecture
Standard Git remains effective for millions of ordinary repositories, but very large monorepos can make full clone and checkout workflows expensive. A 100GB repository can be impractical to materialize on every developer workstation. To support monorepos, systems can use partial clone, sparse checkout, and virtual file system approaches such as Scalar so the client downloads only the objects and files needed for the current workflow.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.