USDC Logo

Use cases · Reference architectures

Three architectures, one way of thinking about a deployment.

USDC plans deployments in a repeatable way. Each reference architecture runs the same seven blocks in the same order, states only sourced figures and closes with what it does not solve.

01Hero
02Situation
03Constraint
04What USDC deploys
05How it works
06For the facility
07Does not solve
DISCUSS YOUR DEPLOYMENT

Start with the workload,
then design the
infrastructure around it.

Bring the near-term number you can defend and the interconnect date you are working toward. USDC will plan the campus around both.

Ready to discuss
your deployment?

Let's plan your campus
with confidence.

Discuss Your Deployment
SOURCES

Vera Rubin NVL72 Facility Planning Summary (NVIDIA)  •  DigiPowerX Cerebras CS4 Business Case.
Reference IT loads are published planning figures, not USDC measured results. Campus pod counts are illustrative.

Start with one pod and grow to a cluster without redesigning the build.

Most buyers need two to three megawatts now and cannot commit to twenty. The pod is the unit of purchase, so the second and sixth pod land on the same design as the first. USDC prepares the shared campus once; compute is added in pod increments as demand arrives.

NVIDIA Vera Rubin NVL72 reference IT load
~2.8MW
NVIDIA Vera Rubin NVL72 reference IT load
Cerebras CS4 reference IT load
2.5MW
Cerebras CS4 reference IT load
Pods in an illustrative campus
4–6PODS
Illustrative 10–15 MW campus range
CAMPUS SCHEMATIC · NOT TO SCALESTAGE 1 / 7CAMPUS PERIMETERPOWER BUSCOOLING HEADERUTILITYINTERCONNECTCONTROL PLANEONE PER CAMPUSCOOLINGPLANT + HEADERSSUBSTATIONSIZED FOR CAMPUSHEADERSSUPPLY / RETURNPOD 01IT ROWPOD 02 · PLANNEDPOD 02IT ROWPOD 03 · PLANNEDPOD 03IT ROWNETWORK SKID1 PER 5 IT PODSPOD 04 · PLANNEDPOD 05 · PLANNEDPOD 06 · PLANNEDPod 01 is built and energized first — revenue before the campus is finished.
01 / 07
THE SITUATION

A funded near term and an
unforecastable curve.

An AI company has a funded workload for the next twelve months and a demand curve after that which nobody can forecast honestly. It needs two to three megawatts of capacity now.

01 NOW · COMMITTED
2–3MW

Funded near-term workload. The number the company can defend today.

● Funded & active
02 FUTURE · UNKNOWN
?UNCERTAIN

Demand curve uncertain. A build-to-suit asks for a shell sized for the whole curve.

○ Unforecastable scale
CAPACITY FORECAST TRAJECTORYCOMMITMENT BOUNDARY VS BUILD-TO-SUIT SHELL
Funded demand (0–12 mo) Diverging cone (12 mo+) BTS Shell Sizing
BUILD-TO-SUIT SHELL · SIZED ONCE, OCCUPIED IN 18–24 MONTHS2–3 MW COMMITMENT0 – 12 MONTHS · FUNDED12 MONTHS + · NOT FORECASTABLETIME →
0–12 MO: KNOWN WORKLOAD (HIGH CONFIDENCE)
12+ MO: UNPREDICTABLE SCALE (CAPITAL AT RISK)
THE BUILD-TO-SUIT REALITY

A build-to-suit data center asks the company to commit to a shell sized for the whole curve, then wait eighteen to twenty four months to occupy it.

THE REALITY

The company signs for what it can defend today, but the facility it signed into was designed once—around one power topology, one cooling loop and one rack density.

02 / 07

The Constraint

The real value of modularity is repeatable engineering.

Modularity is usually sold as a speed argument. Speed is real, but it is not the important part. A modular pod makes the engineering decision repeatable: if the pod is the unit of design, the tenth pod is the same engineering as the first and adding capacity stops being a redesign.

Traditional expansionDesigned once

  • Fixed hall, sized for the whole curve
  • Fixed cooling assumptions
  • Fixed density assumptions
  • Possible redesign during expansion

Repeatable pod architectureDesigned per pod

  • Shared campus framework, sized for the end state
  • Repeatable pod design
  • Incremental capacity
  • Silicon flexibility by pod

03 / 07

What USDC Deploys

The site is prepared for the campus.

One utility interconnect, one perimeter, one cooling plant with headers sized for the end state, one control plane. Only the pods that are needed are built and energized.

Prepare the site for the campus. Build only the pods required today.

Trace a system

Hover or press a system to trace its path through the campus.

PWRCOOLNETUTILITY INTERCONNECTONE INTERCONNECT FOR THE WHOLE CAMPUS • NOT MODULARSHAREDSUBSTATION / SITE INFRASTRUCTUREPAD, YARD, PERIMETER SIZED FOR THE END STATESHAREDCOOLING PLANT + HEADERSONE PLANT • HEADERS SIZED FOR THE CAMPUSSHAREDCONTROL PLANEONE CONTROL PLANE ACROSS EVERY POD ON THE SITESHAREDPOD 01BUILT, ENERGIZEDIT PODPOD 02SAME DESIGNIT PODPOD 03SILICON PER PODIT PODPOD 04 • PLANNEDNOT YET PAID FORPOD 05 • PLANNEDNOT YET PAID FORPOD 06 • PLANNEDNOT YET PAID FORNETWORK SKIDONE PER FIVE IT PODS • TURNS ISLANDS INTO ONE FABRICFABRICSHARED ONCE • BUILT PER POD10–15 MW CAMPUS≈ 4–6 PODS+ 1–2 SKIDS

04 / 07

How It Works

Three phases. One design.

The shared site work is carried by USDC, not by the tenant. Each phase adds only the capital for what is added.

Phase 01 · Single Pod Energization
Plan View · System Schematic
SHARED CAMPUS INFRASTRUCTURE · SIZED FOR END STATESUBSTATION15 MW GRID FEEDCONTROL PLANEUNIFIED TELEMETRYCOOLING PLANTCENTRAL CHILLERSPOD 01BUILT & ENERGIZEDIT PODPOD 02 · PLANNEDPAD & HEADERS READYPOD 03 · PLANNEDPAD & HEADERS READY

Phase 01. One pod is built and energized. The substation, pad, yard and cooling headers are already sized for the full campus, so nothing is re-engineered later.

05 / 07

REFERENCE ENVELOPE

Two pod configurations
USDC has planned against.

Both figures come from named engineering documents, not from a marketing estimate. They define the envelope a pod must accommodate.

01
POD REFERENCE

NVIDIA Vera Rubin NVL72

IT LOADAbout
2.8MW
PHYSICAL UNIT

14 IT racks plus 2 network and storage racks in a single row, network racks in the middle.

SOURCE

Vera Rubin NVL72 Facility Planning Summary

SINGLE ROW LAYOUT
14 IT racks2 N/W / storage
Single row
02
POD REFERENCE

Cerebras CS4

IT LOADAbout
2.5MW
PHYSICAL UNIT

11 containers, about 3,970 sq ft.

SOURCE

DigiPowerX Cerebras CS4 Business Case

CONTAINER LAYOUT
11 containers
≈ 3,970 sq ft
MULTI-POD LAYOUT

Rows of IT pods placed contiguously across the length, with one network skid per five IT pods. A ten to fifteen megawatt campus is therefore four to six pods and one to two network skids.

06 / 07
What This Means for the Facility

Capital tracks demand instead of leading it.

The argument that separates a repeatable pod campus from a hall built once.

01

Capacity tracks demand

Capacity is added in pod increments rather than hall increments, so capital tracks demand instead of leading it.

02

Repeatable engineering

The power and cooling design at pod six is the design at pod one. Nothing is re-engineered mid growth.

03

Silicon flexibility

Silicon is chosen per pod. A later pod can hold a different accelerator generation, or a different vendor, without disturbing the pods already running.

04

Less stranded build

If the demand curve flattens, there is no stranded shell. The pods that were never built were never paid for.

07 / 07
ENGINEERING CONSTRAINT

What This Does
Not Solve

Modular compute does not eliminate
site-level constraints.

The shared elements have to be sized for the end state on day one. Substation capacity, water, land and the utility interconnect are not modular and interconnect queues are measured in quarters or years depending on the ISO. Pod modularity removes the compute commitment risk. It does not remove the interconnect lead time.

WHERE THE CAMPUS CONVERSATION STARTS

Start the campus conversation
with the interconnect date,
not the pod schedule.

SITE-LEVEL ELEMENTS  •  FIXED VS MODULAR
INTERCONNECTUtility interconnect is not modular.
Queue times depend on the ISO.
FIXED
LANDParcel and perimeter are
committed once.
FIXED
WATERWater is not modular and must be
secured for the end state.
FIXED
SUBSTATIONCapacity must be planned for the
full campus on day one.
FIXED
SHARED INFRACooling plant, headers and control
plane are sized for the end state.
SIZED ONCE
LEAD TIMEPod modularity does not remove
interconnection lead time.
FIXED
COMPUTEPods are added in increments.
Commitment risk is removed here
and only here.
MODULAR
DISCUSS YOUR DEPLOYMENT

Start with the workload,
then design the
infrastructure around it.

Bring the near-term number you can defend and the interconnect date you are working toward. USDC will plan the campus around both.

Ready to discuss
your deployment?

Let's plan your campus
with confidence.

Discuss Your Deployment
SOURCES

Vera Rubin NVL72 Facility Planning Summary (NVIDIA)  •  DigiPowerX Cerebras CS4 Business Case.
Reference IT loads are published planning figures, not USDC measured results. Campus pod counts are illustrative.

A pod is two machines, a prefill sidecar and a decode floor.

Serving a model has two phases with opposite hardware appetites. Splitting them inside one pod lets each phase run on the silicon it actually needs and lets power and cooling be provisioned per role.

Prefill
Compute-bound
Prefill · runs near sustained TDP
Decode
Bandwidth-bound
Decode · bursty, low arithmetic utilisation
Transfer
Non-blocking
KV cache transfer, GPU memory to GPU memory

REQUEST PATH·ONE POD, TWO ROLES

INTELLIGENT · BALANCED · HIGH PERFORMANCE
STAGE 1 / 5
REQUEST INPROMPT
POD ENVELOPE
PRESENTED TO THE CUSTOMER AS ONE POD
PREFILL SIDECAR
HIGH ARITHMETIC DENSITY
LOAD
NEAR SUSTAINED TDP
KV CACHE
POD FABRIC
DECODE FLOOR
HIGH BANDWIDTH MEMORY · CONCURRENCY
LOAD
BURSTY
TOKENS OUT
ONE AT A TIME
TWO POWER PROFILES
TWO REFRESH CYCLES
VENDOR CHOSEN PER ROLE
ONE POD, TWO ROLES · TWO POWER PROFILES · TWO REFRESH CYCLES · VENDOR CHOSEN PER ROLE
01 / 07
THE SITUATION

A training pod becomes idle
capital the moment the run ends.

The obvious answer is to serve inference on it. The less obvious problem is that inference is not one workload and a pod configured as a uniform block of identical GPUs is the wrong shape for it.

WHAT HAPPENS AFTER TRAINING?

There are three possible outcomes for a training pod. Click any card to inspect.

TRAINING RUN

Homogeneous GPU block

• Fully used
UTILIZATION100%

RUN ENDS

Idle capital

• 0% cluster utilization
UTILIZATION0%

SERVE INFERENCE?

Wrong shape for
two phases

IMPACTHIGH
02 / 07
THE INSIGHT

Two phases that want opposite things from the hardware.

A homogeneous fleet sized correctly for one of those phases is sized incorrectly for the other. That is true of the silicon and it is equally true of the power and cooling design wrapped around it.

LIVE WORKLOAD PROFILE OSCILLOSCOPE
Phase 01 · PrefillCompute-Bound

Reads the prompt, builds the KV cache

  • Bound ByCompute (Raw FLOPS)
  • Power ProfileNear sustained TDP (~94%)
  • Silicon ArchitectureHigh arithmetic density accelerators
  • Hardware LifecycleFrequent refresh cycle (Fast compute evolution)
TELEMETRY · POWER PROFILE~94% SUSTAINED LOAD
SUSTAINED TDP (PEAK LOAD)PROMPT INGESTION →CONTINUOUS EXECUTION
Phase 02 · DecodeBandwidth-Bound

Emits output tokens one at a time

  • Bound ByMemory Bandwidth (HBM3e/HBM4)
  • Power ProfileBursty (Low arithmetic utilisation)
  • Silicon ArchitectureHigh bandwidth memory, high concurrency
  • Hardware LifecycleLonger amortisation (Stable memory architecture)
TELEMETRY · POWER PROFILEBURSTY AUTOREGRESSIVE SPIKES
SUSTAINED TDP CEILING (UNDERUTILIZED)1 TOKEN EMITTED PER BURST →42 MS / TOKEN
WHY IT MATTERS

The Architectural Mismatch: Sizing a single homogeneous cluster for Decode starves Prefill of compute throughput; sizing for Prefill wastes massive power and cooling envelope during intermittent Decode memory bandwidth stalls. USDC solves this by co-locating both roles with tailored feeds inside one pod.

03 / 07

What USDC Deploys

A decode floor with a prefill sidecar inside the same envelope.

The pod is configured as a decode floor. Alongside it, inside the same pod envelope, a prefill sidecar holds a smaller number of accelerators chosen for raw compute rather than memory bandwidth. The two are joined by the pod fabric and presented to the customer as one pod with two roles.

One pod. Two roles. Power and cooling provisioned per role.

Trace a role

Hover or press a role to see what it owns inside the pod.

POD ENVELOPE · ONE UNIT TO THE CUSTOMERPOD POWER + COOLING FEEDMETERED AND COOLED PER ROLE, NOT AS ONE AVERAGED LOADPREFILL SIDECARSMALLER NUMBER OF ACCELERATORSCHOSEN FOR RAW COMPUTENEAR SUSTAINED TDPOWN REFRESH CYCLEOWN VENDORDECODE FLOORBULK OF THE PODCHOSEN FOR BANDWIDTH AND CONCURRENCYBURSTYKEEPS RUNNING WHILE THE SIDECAR IS OPENEDOWN VENDORPOD FABRICJOINS THE TWO ROLES · KV CACHE MOVES SIDECAR → FLOOR
SIDECAR · COMPUTE · SUSTAINEDFLOOR · BANDWIDTH · BURSTY

04 / 07

How It Works

A production pattern, not a research idea.

The serving stack routes an incoming request to a prefill worker, which computes the KV cache and then transfers that cache to a decode worker which produces the output tokens.

01

Route

The serving stack routes the incoming request to a prefill worker.

02

Prefill computes the KV cache

The prefill engine reads the prompt and builds the cache under sustained compute load.

03

NIXL moves the cache

In NVIDIA Dynamo the transfer is handled by NIXL, directly from the video memory of the prefill engine to the video memory of the decode engine. The transfer is non blocking, so GPU forward passes continue serving other requests while it happens.

04

Decode starts immediately

With the SGLang backend, prefill runs as a background task and decode begins immediately while the transfer proceeds in parallel.

STEP 01 · ROUTELINK ACTIVE · NODE 04 · RT 0.4MSDISAGGREGATED SERVINGSGLANG · NVIDIA DYNAMO ARCHITECTUREREQUEST INPROMPTROUTERDISPATCHPREFILL WORKERGPU MEMORY · COMPUTE BOUNDGPU ONLINE · READY FOR BATCHKV CACHESUSTAINED PREFILLNIXLNON-BLOCKINGDECODE WORKERGPU MEMORY · BANDWIDTH BOUNDGPU ONLINE · IDLEKV CACHEBURSTY DECODETOKENS OUTSTANDBY01Step 01.The router directs incoming prompt requests to an optimal prefill worker selected for maximum arithmetic compute density.NVIDIA DYNAMO ARCHITECTURE · SGLANG DISAGGREGATED INFERENCE RUNTIMEFORWARD PASSES KEEP SERVING OTHER REQUESTS · 100% NON-BLOCKING PIPELINEUSDC REFERENCE SPEC
05 / 07
REFERENCE ENVELOPE

Two roles,
two envelopes.

The published characteristics USDC plans each role against.
No throughput or latency figures are claimed for
this configuration.

01PREFILL SIDECAR

Reads the prompt,
computes the KV cache.

BOUND BYCompute
POWER PROFILENear sustained TDP
SILICON PREFERENCEHigh arithmetic density
ENVELOPESmaller number of accelerators inside the same pod
02DECODE FLOOR

Emits output tokens
one at a time.

BOUND BYMemory bandwidth
POWER PROFILEBursty
SILICON PREFERENCEHigh bandwidth memory, high concurrency
ENVELOPEThe bulk of the pod
THE TAKEAWAY

A pod configured as a uniform block of identical GPUs is the wrong shape for inference across two distinct phases.

06 / 07
What This Means for the Facility

Where the argument stops being a software argument.

This is where it becomes a USDC argument.

01

Power per role

Prefill sits near sustained TDP while decode is bursty. Metering and cooling them as a single averaged load oversizes one and starves the other. A sidecar lets power and cooling be provisioned per role.

02

Refresh per role

Prefill silicon can be replaced on a different schedule than decode silicon and only the sidecar is opened. The decode floor keeps running.

03

Vendor per role

The two roles do not have to come from the same vendor. A high arithmetic density accelerator can serve prefill while current generation GPUs serve decode. This is where GPU agnostic stops being a slogan.

04

Train, then serve

The same pod serves training between contracts and disaggregated inference under them. The reconfiguration is a software and sidecar change, not a rebuild.

07 / 07
ENGINEERING CONSTRAINT

What This Does
Not Solve

Disaggregation is not always faster.

Disaggregation pays off when prompts are long enough that the cache transfer costs less than recomputing prefill.

For short prompts it can cost more than it saves.

Production stacks handle this by deciding per request whether to disaggregate and the pod supports both modes.

THE HONEST VERSION

The workload decides, not the building.

Anyone who says disaggregation is always faster is selling rather than engineering.

TRANSFER COST VS RECOMPUTE COST

BREAK-EVEN ANALYSIS
TRANSFER KV CACHE COST
RECOMPUTE PREFILL COST
COSTHIGHMEDIUMLOWBREAK-EVEN POINTTransfer cost equalsrecompute costSHORT PROMPTSPROMPT LENGTHLONGER PROMPTS
SHORT PROMPTS

Recompute is cheaper.

Disaggregation may cost more.

LONGER PROMPTS

Transfer is cheaper.

Disaggregation pays off.

DISCUSS YOUR DEPLOYMENT

Start with the workload,
then design the
infrastructure around it.

Bring the model, the prompt profile and the concurrency you expect. USDC will size the sidecar and the floor around them.

Ready to discuss
your deployment?

Let's plan your campus
with confidence.

Discuss Your Deployment
SOURCES

NVIDIA Dynamo documentation, disaggregated serving design notes (NIXL, SGLang backend). No throughput, latency or power figures are claimed for this configuration.

KV cache becomes a network service across the USDC footprint.

Agentic workloads send the same long context back to the model over and over. A shared cache tier turns that repetition from a cost into an advantage and it only works if the sites sit on good fiber.

Cache hit rate
1.7→92.2%
Cache hit rate on Codex traces · vLLM + Mooncake, published
Throughput
3.8×
Throughput improvement · same published benchmark
Inter-site round trip
<10MS
Inter-site round-trip target · three diverse paths
CACHE TIER · POD → SITE → FOOTPRINT
STAGE 1 / 5
SITE A· FABRIC SPEEDPOD 01PRIMARYGPU MEMORYHBM3e · SUB-MICROSECONDHOTCPU MEMORYDDR5 HOST RAM · ~100nsLOCAL NVMePCIe Gen5 SSD · ~10µsPREFIX REUSED · 94% HITPOD 02SECONDARYGPU MEMORYHBM3e · READYCPU MEMORYDDR5 HOST RAMLOCAL NVMePCIe Gen5 SSDREADS SAME POOL · ZERO MISSSITE A CACHE POOLLMCACHE CLUSTER · ANY POD READS INSTANTLYSHARED TIER · FABRIC-INTERCONNECTED · MULTI-TB CAPACITYREADYSITE B· REMOTE MIRRORPOD 01 (SITE B)STANDBYGPU MEMORYHBM3e · READYCPU MEMORYDDR5 HOST RAMLOCAL NVMePCIe Gen5 SSDSESSION FOLLOWS CAPACITYSITE B POOLEXTENDED CACHE MIRRORBACKBONE EXTENDED · <10ms3 DIVERSE PATHS · <10ms
A long prompt prefix is cached in GPU memory on Pod 01.
01 / 07
THE SITUATION

The same prefix,
resent on every turn.

Agentic and long context workloads resend the same prompt prefix on every turn. Each repeat that lands on a node without the cache pays the full prefill cost again.

The user sees it as time to first token and the operator sees it as GPU hours spent recomputing something that was already computed an hour ago.

AGENT TURNS
TURN 1PREFIXFULL PREFILL
TURN 2SAME PREFIXRECOMPUTED
TURN 3SAME PREFIXRECOMPUTED
TURN 4SAME PREFIXRECOMPUTED
...EVERY TURN
COST THE USER SEES

Time to first token

COST THE OPERATOR SEES

GPU hours

BOTTOM LINE

Repeated prefix without cache locality = repeated full prefill cost.
Bad for latency. Bad for efficiency.

02 / 07
THE CONSTRAINT

Scaling out makes the
problem worse, not better.

Cache held in GPU memory is local, small and lost when the instance moves. Once a deployment grows past a single node the hit rate falls, because the router cannot reliably send a request back to the machine that holds its prefix.

CACHE IN GPU MEMORY
PER NODE
Local to one node
Small
Lost when the instance moves
Hit rate falls as nodes are added
ROUTER CANNOT FIND THE NODE HOLDING THE PREFIX
Node • hit
Miss
Miss
Miss
...
CACHE AS A SHARED TIER
PER SITE • PER FOOTPRINT
Lives in a sidecar, not the GPU
Layered: GPU → CPU → NVMe → pool
Any pod on the site can read
Extends to other sites over the backbone
ANY NODE READS THE SHARED POOL
Node • hit
Hit
Hit
Hit
...
BOTTOM LINE

Local GPU cache does not scale. A shared, layered cache tier preserves hit rate and reliability as you grow.

03 / 07

What USDC Deploys

A KV cache tier that lives in a sidecar, not in the GPU.

The tier is layered. GPU memory first, then CPU memory, then local NVMe, then a pool that any pod on the site can read. Across the USDC backbone, that pool extends to other sites in the footprint.

Within a site, cache moves at fabric speed. Between sites, at backbone speed.

  1. T0GPU memoryNode
  2. T1CPU memoryNode
  3. T2Local NVMeNode
  4. T3Site pool · any pod can readSite
  5. T4Pool extended over the USDC backboneFootprint

Trace a scope

Hover or press a scope to see which tiers it spans.

SITE AFABRIC SPEED INTERCONNECTPod 01 · GPUHBM3e · Hot TierCPU MemoryDDR5 Host RAMLocal NVMePCIe Gen5 SSDPod 02 · GPUHBM3e · Hot TierCPU MemoryDDR5 Host RAMLocal NVMePCIe Gen5 SSDSite PoolAny pod on the site reads · fabric speedSHAREDEast–west traffic between podsSITE BREMOTE MIRROR · DIVERSE PATHSPod 01 · GPUHBM3e · ReadyCPU MemoryDDR5 Host RAMLocal NVMePCIe Gen5 SSDPod 02 · PlannedSite PoolExtended over the backbone · backbone speedMIRRORUSDC BackboneThree diverse paths · round-trip target under ten millisecondsSession context follows the customer to whichever site has capacity<10ms RTTNODE · SITE · FOOTPRINTFIBER ADJACENCY IS A SITING REQUIREMENT

04 / 07

How It Works

A cluster-wide pool, moved without touching the GPU.

Named software, doing the job today.

01

Mooncake Store · pool

Runs a cluster wide distributed KV cache pool, with a master server holding metadata and clients on each GPU node.

02

GPUDirect RDMA · transfer

Moves cache without consuming GPU streaming multiprocessors and without a CPU staging buffer, on dedicated background threads so GPU kernel launches are not blocked.

03

LMCache · tiers and reuse

Does the same job across GPU memory, CPU memory, local SSD and remote backends and reuses cache across requests, sessions and engine instances.

Step 01 · PoolSoftware layer

Step 01. Mooncake Store keeps a cluster-wide pool: one master server for metadata, a client on every GPU node.

05 / 07
PUBLISHED RESULTS

Why the facility
design matters.

The published results on agentic traces are the
reason this matters. They describe what the
software layer achieves.

1.792.2%

Cache hit rate

Codex traces  •  vLLM +
Mooncake

3.8×

Throughput
improvement

Codex traces  •  vLLM +
Mooncake

46×

Median time to
first token, lower

Codex traces  •  vLLM +
Mooncake

8.6×

End-to-end latency,
lower

Codex traces  •  vLLM +
Mooncake

60GPUs

Near-linear throughput
scaling to 60 GB200 GPUs,
hit rate held above 95%

Codex traces  •  vLLM +
Mooncake

ATTRIBUTION

These are published benchmark figures from the vLLM and Mooncake teams, measured on their traces and their hardware. They describe what the software layer achieves. USDC cites them to explain why the facility design matters. They are not a USDC measured result.

06 / 07
What This Means for the Facility

Fiber adjacency becomes a siting requirement.

A shared cache tier is only useful if the sites holding it are close in network terms. That turns fiber adjacency into a siting requirement rather than a convenience and it is the strongest available argument for how USDC selects land.

01

Inter-site fabric

Three diverse paths with a round trip target under ten milliseconds, as described on the Global Network page, make cross-site cache reuse and session migration practical.

02

East–west provisioning

Cache movement is a bandwidth consumer, not a rounding error. Sites have to be provisioned for east to west traffic between pods, not only for north to south traffic to the internet.

03

Land on fiber routes

Land selection favours parcels on dense fiber routes over the lowest cost acreage. A cheaper parcel that adds latency between sites removes the reason the footprint exists.

04

Capacity as scheduling

A customer can be routed to whichever site has capacity while the session context follows them, which converts a multi site footprint from an operational burden into a scheduling advantage.

07 / 07
ENGINEERING CONSTRAINT

What This Does
Not Solve

The site boundary is where the cache tier changes from a performance feature to a capacity feature.

Cross site reuse works for prefix reuse and for moving a session to where capacity exists. It does not work for a tight prefill and decode loop split across two cities.

Within a site, cache moves at fabric speed. Between sites it moves at backbone speed and the workload has to tolerate that difference.

THE HONEST VERSION

Within a site: performance. Between sites: capacity.

Do not split a prefill–decode loop across two cities.

WHAT CROSSES THE SITE BOUNDARY

LATENCY DOMAIN & CACHE SCOPE
WITHIN A SITE
Fabric speed
Prefix reuse● Supported
Session migration● Supported
Prefill ↔ decode loop● Supported
BETWEEN SITES
Backbone speed
Prefix reuse● Supported
Session migration● Supported
Prefill ↔ decode loopNot across two cities
✕ Excluded
DISCUSS YOUR DEPLOYMENT

Start with the workload,
then design the
infrastructure around it.

Bring the prompt reuse profile and the sites you need to reach. USDC will plan the cache tier and the fiber around both.

Ready to discuss
your deployment?

Let's plan your campus
with confidence.

Discuss Your Deployment
SOURCES

vLLM and Mooncake published benchmark results on Codex traces  •  Mooncake Store and LMCache documentation. Benchmark figures are measured by their authors on their hardware and are not USDC results. Inter-site targets are from the USDC Global Network page.