The workload changed shape, not just size
A human developer usually works in bursts: edit for minutes, commit occasionally, push a small number of branches, then wait for review. A coding agent can checkpoint after nearly every tool action, run many attempts at once, retry immediately, and keep working around the clock. The result is not simply more Git traffic. It is sustained write concurrency followed by a cascade of fetches, builds, scans, index updates, webhooks, and merge-queue work.
GitHub put unusually concrete numbers behind that shift in an October 6 architecture post. It says monthly Git activity rose from 218.2 billion events in September 2025 to 473.3 billion in August 2026. Developers and agents made 7.38 billion commits in September, more than five times the prior-year volume. Pushes rose from 0.69 billion to 3.35 billion per month, pull-request merges approached four times their earlier volume, and GitHub Actions ran 3.26 billion times in September. Those are GitHub's first-party measurements, not an independent benchmark, but they explain the design pressure.
The key distinction is between independent work and publication. Thousands of agents can create immutable objects on separate branches without disagreeing. They collide when they try to update the same branch, tag, merge queue, or release ref. A fast object upload does not make a stale compare-and-swap safe. A fast clone does not resolve contention on refs/heads/main. The scalability ceiling sits where a distributed service must publish one ordered truth.
Agent-scale Git is not solved by making clones faster. The hard part is publishing concurrent writes durably and consistently, then absorbing every read those writes trigger.
Public discussion of the exact architecture post remains small, so there is no basis for claiming broad consensus. The more visible community signal is reliability anxiety around recent Git operations and Actions incidents. That context matters because a migration cannot trade availability and auditability for a benchmark. It does not prove that agent traffic caused any incident.
Why one replica can be both protection and drag
GitHub's current repository layer is called Spokes. The platform keeps complete repository copies on local disks across several fileservers, five by default in the current description. Local disks provide low-latency native Git access. Replication provides durability and lets reads spread across hosts. A reference-changing push uses a three-phase protocol and a quorum so the web interface, API clients, and CI see a consistent repository state.
client push
-> route to repository replica set
-> transfer and validate objects on replicas
-> propose reference update
-> reach quorum across durable copies
-> publish the new ref
-> trigger CI, scans, indexes, webhooks, and caches
This design joins two concerns. The replicas are the authoritative durable copies, and they are also the machines that serve reads. Adding a replica supplies more read capacity, but the new replica also participates in writes. A push can be bounded by the slowest required participant. Losing a replica removes read capacity; losing quorum blocks publication. What protects the data also constrains the write path.
The tradeoff was rational for human-scale write patterns. GitHub's older Spokes engineering notes describe why application-level replication and quorum improved durability, geographic resilience, and read locality. They also make the limiting mechanism clear: reference updates must be serialized, and long-distance or overloaded participants increase latency. Agent fleets turn an edge case into a normal workload.
| Pressure | Human-heavy pattern | Agent-heavy pattern | Infrastructure consequence |
| Checkpoint rate | Occasional meaningful commits | Commit after small verified actions | Push latency enters the agent's inner loop |
| Branch count | A few active branches per person | Many parallel attempts and worktrees | More refs, objects, leases, and cleanup |
| Shared publication | Review-paced merges | Automated queues and rapid retries | Hot-ref contention and retry storms |
| Read fan-out | Developer clones and periodic CI | Every push starts builds, scans, indexes | One write produces thousands of reads |
| Traffic rhythm | Business-hour and release peaks | Sustained multi-region automation | Peak provisioning wastes capacity off-peak |
Narrow consensus and let everything else scale independently
GitHub's announced approach applies a familiar distributed-systems rule: coordinate only the state that needs agreement. Git objects are content-addressed and immutable. If an object with a particular identifier is present and valid, workers do not need a global order for its arrival. References are mutable names that point into the object graph. Publishing a reference must be atomic, conditional on the expected old value, and unable to expose a missing graph.
+----------------------+
push objects ------> authoritative object | Azure Blob Storage
| storage |
+----------+-----------+
|
validate graph| and policy
v
+----------+-----------+
update ref --------> small coordinated | compare-and-swap / ordered log
| publication path |
+----------+-----------+
|
published version / invalidation
v
+----------------+ +----------------+ +----------------+
| read worker + | | read worker + | | maintenance / |
| local cache | | local cache | | compaction |
+----------------+ +----------------+ +----------------+
In the new description, Azure Blob Storage becomes the authoritative durable layer. Lightweight compute workers cache repository data and serve Git requests. Read capacity can expand without adding a durable replica to every push. Losing a worker becomes closer to a cache miss than a durability event: another worker can fetch authoritative objects as traffic arrives. Compaction and garbage collection move to separate workers against durable storage instead of competing with live pushes and fetches.
The interesting optimization is the critical path. Object storage, connectivity checks, secret scanning, and other validation remain mandatory, but much of that work can proceed in parallel. The publication path should contain only what must agree before clients can observe the new ref. GitHub says internal benchmarks reached up to 35 times higher write throughput. Treat that as a directional first-party result until workload shape, percentile latency, fault injection, storage cost, and comparison boundaries are published.
Decoupling does not erase coordination. It relocates and sharpens it. A branch protection rule, required review, signed-commit policy, hidden ref, legal hold, audit record, or merge queue must still observe the exact transaction that publishes state. If those controls are evaluated on a stale cache or outside the atomic boundary, throughput improves while governance silently weakens.
Treat reference publication as the transaction
A safe push can acknowledge object receipt before publication, but it cannot report success until every object reachable from the proposed ref is durable, valid, authorized, and addressable. A reference lease prevents an agent from overwriting work it did not see. Policy evaluation must bind to the proposed old and new object IDs, not merely to a branch name.
def publish_ref(repo, ref, expected_old, proposed_new, actor, request_id):
receipt = durable_objects.require(repo, reachable_from=proposed_new)
graph = validate_connectivity(repo, proposed_new)
policy = evaluate_controls(
repo=repo,
ref=ref,
old=expected_old,
new=proposed_new,
actor=actor,
object_receipt=receipt,
graph=graph,
)
if not policy.allowed:
return reject(policy.reasons)
result = refs.compare_and_swap(
repo=repo,
ref=ref,
expected=expected_old,
replacement=proposed_new,
idempotency_key=request_id,
)
audit.append(result, policy.version, receipt.id)
invalidate_caches(repo, ref, result.version)
emit_downstream_events_once(result.transaction_id)
return result
The idempotency key matters when clients retry after a timeout. Without it, the service may emit duplicate webhooks, CI runs, or billing events even if the reference ends in the correct state. The version carried into cache invalidation prevents an older invalidation from overwriting knowledge of a newer ref. The audit link binds authorization, object durability, publication, and downstream fan-out into one trace.
Hot refs need admission control. If a merge queue submits faster than the ref can serialize updates, blind immediate retries make the queue less stable. Use bounded concurrency, exponential backoff with jitter, explicit stale-lease responses, and a queue that can recompute on the newest head. Agent clients should understand conflict as a normal state transition, not an invitation to hammer the same push.
Budget the write and its fan-out together
Average repository requests conceal the bottleneck. Capacity planning should model a hot repository or organization, distinguish object bytes from ref updates, and carry every published write into derived workload. A small commit can be expensive if it triggers thousands of identical fetches and heavyweight scans.
published_writes_per_second = agent_count * checkpoints_per_agent_second * publish_ratio
ref_utilization = published_writes_per_second * p99_ref_transaction_seconds
derived_reads_per_second = published_writes_per_second * (
ci_jobs_per_push
+ security_scans_per_push
+ index_consumers_per_push
+ webhook_consumers_per_push
)
object_ingress = unique_object_bytes_per_second
cache_origin_rate = derived_reads_per_second * cache_miss_ratio
maintenance_budget = object_ingress * compaction_amplification
A ref utilization near one means the shared branch is continuously busy before retries. The remedy may be fewer checkpoints, branch-local buffering, merge batching, or a different publication topology, not more read workers. Derived reads should be deduplicated where semantics allow: multiple consumers asking for the same packfile or tree can share cache work, but authorization and visibility must remain scoped.
| Metric | Why it matters | Bad interpretation |
| p50 / p95 / p99 push latency | Agents amplify tail latency through tight loops | Reporting only average latency |
| Ref conflicts and stale leases | Shows actual publication contention | Treating every rejected push as platform failure |
| Objects uploaded per published ref | Separates speculative work from accepted history | Counting all commits as productive output |
| Derived reads per publication | Sizes CI, scan, and index fan-out | Budgeting only Git client traffic |
| Cache hit by object type and age | Reveals origin-storage pressure | Using one global hit ratio |
| Recovery time after worker loss | Tests the storage/compute separation | Assuming cache workers are disposable |
Migrate the service, not the user-visible semantics
GitHub must rebuild while repositories keep moving. That makes compatibility an invariant, not a launch-day task. Existing clients, branch protections, merge queues, hooks, audit streams, visibility rules, legal holds, and recovery procedures must behave the same across old and new storage paths. A successful data copy with divergent authorization is a failed migration.
Start with an immutable inventory: repository identity, refs, object reachability, alternates, shallow and partial-clone state, large-file pointers, hidden namespaces, policy configuration, hook versions, retention requirements, and current placement. Backfill objects into durable storage and prove hashes and reachability. Then shadow read requests and compare results without serving them to users.
phase 0 inventory + invariants + rollback owner
phase 1 backfill immutable objects; verify hashes and reachability
phase 2 shadow reads; compare refs, packs, permissions, and latency
phase 3 mirror accepted writes; verify exactly-once downstream events
phase 4 canary authoritative reads for low-risk repositories
phase 5 canary ref publication with rapid per-repository rollback
phase 6 expand by workload class, not random percentage alone
phase 7 retire old authority only after recovery and audit drills pass
Workload-class canaries are essential. A dormant repository, a giant monorepo, a fork network, a repository with heavy Git LFS use, a regulated enterprise, and an agent fleet exercise different paths. Random percentage rollout may delay the rare but dangerous cases until traffic is already large.
Dual write is not automatically safer. If two systems can independently publish refs, operators have created two authorities. Prefer one publication authority with a mirrored verifier, or define an explicit ordering and reconciliation protocol. Rollback must include ref state, not only object data, and it must explain what happens to writes accepted after the cutover boundary.
| Invariant | Verification | Stop condition |
| No published ref points to a missing object | Reachability walk plus sampled fresh fetch | Any missing or corrupt reachable object |
| Reference updates remain atomic | Concurrent lease and delete/update replay | Split view or non-serializable result |
| Authorization is identical | Old/new decision diff on users, tokens, apps | Any privilege expansion |
| Downstream events are exactly once by transaction | Event-ledger reconciliation | Duplicate or missing release-critical event |
| Audit evidence remains complete | Trace request to object receipt, policy, ref, events | Unlinked or mutable decision evidence |
| Rollback preserves accepted work | Regional and storage fault drill | Unowned write gap or recovery ambiguity |
Failure modes hidden by a throughput win
| Failure | What it looks like | Control |
| Object/ref race | A ref becomes visible before all reachable objects are readable | Durability receipt and connectivity gate inside publication |
| Stale authorization | A cache serves data after visibility or token scope changes | Versioned policy decisions and scoped invalidation |
| Hot-ref retry storm | Agents repeatedly fail compare-and-swap and increase contention | Queueing, jitter, recompute, and per-actor rate limits |
| Event multiplication | A timed-out push produces duplicate CI, scans, or webhooks | Idempotency key and transaction-bound event ledger |
| Cache recovery avalanche | Worker loss sends a synchronized burst to origin storage | Warm pools, request coalescing, origin budgets, staged recovery |
| Maintenance starvation | Object growth outpaces compaction and retention cleanup | Dedicated workers, backlog SLO, admission control |
| Benchmark mismatch | 35x synthetic writes, little gain on protected production refs | Representative policies, object sizes, faults, and tail latency |
| Control drift | Branch rules, hooks, holds, or audit fields differ by path | Decision diff and migration invariant suite |
A reliability incident during the same week as an architecture announcement is tempting narrative material. Resist the shortcut. GitHub's October 7 incident affected Git operations, pull requests, and Actions, and users expressed frustration about status reporting. Public evidence does not establish that agent traffic or the new architecture caused it. Use the incident to test observability and migration expectations, not to invent causality.
Agent-scale repository readiness checklist
- Measure sustained and burst object uploads, ref updates, conflicts, retries, and derived reads by repository.
- Separate immutable object durability from mutable reference publication in the system model.
- Keep compare-and-swap, reachability, authorization, protected-branch policy, and audit evidence in one transaction contract.
- Move garbage collection and compaction off request-serving workers, with a backlog SLO and failure owner.
- Make retries idempotent across publication, webhooks, CI, scanning, indexing, and billing.
- Rate-limit and queue hot-reference writers; return actionable stale-lease responses to agent clients.
- Test cache loss, origin throttling, regional faults, slow storage, worker churn, and partial invalidation.
- Canary by workload class and governance profile, not only by random traffic percentage.
- Shadow and diff reads, writes, permissions, hooks, holds, and audit records before authority moves.
- Preserve an explicit rollback point and reconcile every write accepted after cutover.
Frequently asked questions
Why do coding agents stress Git differently?
They can checkpoint and retry inside a tight tool loop, run many branches concurrently, and trigger continuous CI and scanning. The system sees sustained writes and derived reads instead of a mostly human rhythm.
Why not add more replicas?
Replicas help reads, but if every durable replica joins each write quorum, more read capacity increases push work and tail latency. Decouple authoritative storage from disposable serving workers.
What must remain coordinated?
The mutable reference publication must remain atomic, conditional, authorized, and connected to durable objects. Immutable object transfer, many validations, caching, and maintenance can proceed independently as long as publication cannot expose an invalid graph.
Does object storage make Git semantics weaker?
It should not. Storage placement is an implementation detail. Reference ordering, reachability, permissions, hooks, branch rules, auditability, and accepted-write durability remain part of the user-visible contract.
Is GitHub's 35x result a general benchmark?
No. GitHub reports an internal result of up to 35 times higher write throughput. Without the exact workload, percentile latency, failure model, cost, and policy path, it is evidence for the architecture direction, not a universal performance prediction.
Sources and further reading
Sources were checked on October 8, 2026. Current scale and benchmark figures are attributed to GitHub; community threads are used as discussion and reliability signals, not proof of architectural cause.