Capacity
Serving a 62 GB patch to four million clients
On 17 July a customer shipped the largest single object we have ever cached. This is what the network did with it.
The object was 62.4 GB. The announcement went out at 17:00 UTC and the client auto-updater started pulling immediately, which is the worst possible traffic shape: a near-vertical ramp with no geographic smoothing, because everyone's client checked in at the same second.
Ninety seconds after the announcement we were pushing 71.4 Tbps. The origin — a single pair of servers behind a 10 Gbps uplink — peaked at 340 Mbps. That ratio is the entire product, so it is worth explaining how it happens.
Slices, not files
The edge does not think in files. It thinks in 8 MB slices. A 62.4 GB object is 7,988 slices, each independently cacheable and independently fetchable. When the first client in Frankfurt asked for byte 0, the edge asked the regional shield for slice 0 only.
Meanwhile the prefetcher noticed a sequential access pattern and started pulling slices 1 through 24 in parallel. By the time the client had consumed the first 8 MB, the next 192 MB were already local.
# what one edge machine saw in the first minute
slices_requested 1_284_991
slices_fetched 7_988 # exactly the object
collapse_ratio 160.9x
origin_requests 0 # shield absorbed all of itRequest collapsing is the whole trick
Four hundred thousand clients in one metro asked for slice 0 within the same 200 milliseconds. The edge issued one upstream request and parked the rest on a wait queue keyed by slice identity. When the slice landed, every waiter was served from memory.
The interesting engineering is not in the fast path. It is in making sure that 400,000 simultaneous waiters do not cost you 400,000 goroutines and 12 GB of queue overhead.
We cap the wait queue at 64k entries per slice and shed the remainder with a 30 ms retry hint. In practice the queue never got past 41k, because the upstream fetch completed in 18 ms.
Where it got uncomfortable
Two regions did not behave. São Paulo saturated its transit at 17:04 because a peering session had been down since the previous week and nobody had noticed — the traffic had simply been riding transit at three times the cost. Mumbai ran out of NVMe write bandwidth, not read bandwidth, because the admission policy was still writing slices it would never serve again.
| Region | Peak egress | Origin fetches | p95 TTFB |
|---|---|---|---|
| eu-central | 24.1 Tbps | 7,988 | 9 ms |
| us-east | 18.7 Tbps | 7,988 | 12 ms |
| ap-northeast | 11.2 Tbps | 7,988 | 14 ms |
| ap-south | 6.4 Tbps | 7,988 | 38 ms |
| sa-east | 3.1 Tbps | 7,988 | 61 ms |
What we changed
- Peering sessions now page when they flap, not when they stay down — the old alert only fired on transition.
- The admission policy considers write cost, so a slice that is unlikely to be re-read is served from memory and never written.
- Prefetch depth is now a function of measured client throughput rather than a constant.