Systems & Infrastructure
Networking & the Data Path
IntermediateAt scale the network is the computer: how GPUs talk to each other, why bandwidth falls off a cliff past the rack, and how data gets from storage into GPU memory.
From one GPU to a SuperPOD: scale up vs scale out
A frontier model never fits on one GPU. NVIDIA's answer is a hierarchy of ever-larger “single” computers - GPU, node, NVL72 rack, SuperPOD. Up to the rack, GPUs share one domain at 1.8 TB/s per GPU; past it, racks are joined over InfiniBand or Spectrum-X at roughly 50-100 GB/s per GPU. Each step out buys scale and gives up bandwidth.
That hierarchy splits into two moves. Scaling up grows the NVLink domain into one bigger GPU; scaling out stitches racks together over InfiniBand or Ethernet. Zoom from a single node to a SuperPOD and watch where the fat NVLink links give way to the thin network - and where storage bypasses the CPU entirely.
Scale up, then scale out - one zoom from GPU to SuperPOD
Step the camera back. Each level folds into a single tile of the next, and every link is drawn on one bandwidth scale - so you can watch the fat NVLink pipes give way to the thin network.
01 / 04
1 GPU
The unit of compute
A single accelerator with its own HBM. Once a model or its KV cache exceeds this memory, you must scale out.
- NVLink domain
- 1 GPU
- Widest link
- NVLink 1.8 TB/s leaves the package
Grace NVLink-C2C
A coherent ~900 GB/s link fuses each Grace CPU to its Blackwell GPUs, with LPDDR5X acting as a capacity tier behind HBM - CPU and GPU share one memory space without crossing PCIe.
Scale up - grow the NVLink domain, make one bigger GPU. Bandwidth-rich (1.8 TB/s/GPU), but capped at 72 GPUs per domain today.
Scale out - add racks over InfiniBand / Ethernet. Near-unlimited size, but each hop carries ~18–36× less bandwidth than NVLink - so keep tight traffic in-domain and spread the loosely-coupled work out.
Schematic: tile counts and positions are illustrative; link widths share one scale with a 1 px floor.
The interconnect
At scale, the network is the computer. How fast GPUs exchange gradients and activations decides whether a 10,000-GPU cluster behaves like 10,000 GPUs or a fraction of that.
Which link does what - four interconnects on one path
Two nodes and the network between them. Pick an interconnect to light up where it sits and send a packet along the traffic it carries.
NVLink
GPU ↔ GPU, inside a node
Direct high-bandwidth links between GPUs. 900 GB/s on Hopper, 1.8 TB/s per GPU on Blackwell - orders of magnitude faster than PCIe.
Inside the node the links are fat; between nodes they are thin. NVLink and NVSwitch keep GPUs talking at terabytes per second; InfiniBand or Spectrum-X carries everything that has to leave. The bandwidth cliff further down puts numbers on that gap.
Schematic: 4 GPUs per node and 2 spine switches are drawn for clarity.
The bandwidth cliff
Those interconnects aren't interchangeable - they span more than two orders of magnitude in bandwidth. Plotting them on a log scale shows why the topology is built the way it is: stay on NVLink and you move terabytes per second; cross InfiniBand and you fall off a cliff.
The bandwidth cliff - a pipe map from HBM outward
Each pipe is one hop further from the GPU die, drawn as thick as its bandwidth. Pick a payload and tap a pipe: the ruler below shows how long the same transfer takes on every tier (time = size ÷ bandwidth).
1.00 TB over InfiniBand 400G (NDR)
20.00 s
50 GB/s · 160× slower than HBM
Quantum-2 InfiniBand, scale-out per port.
Same payload inside the NVLink domain
555.6 ms
36× longer once it leaves the domain over IB 400G.
Tap any pipe or ruler row to change the tier. Pick GPUDirect Storage to spread it across more DPUs.
Same 1.00 TB transfer, every tier · log time
1.00 TB clears HBM in 125.0 ms but takes 20.00 s over a single 400G InfiniBand link. The biggest step is the NVLink-domain edge - a 18× (800G) to 36× (400G) drop - which is why 72 GPUs share one NVLink domain. When thousands of GPUs stall on checkpoints, data loading, or KV reloads, the bottleneck is the storage and network path - exactly where a high-bandwidth data platform keeps the fleet busy.
Order-of-magnitude figures. Pipe thickness ∝ √bandwidth. Times are size ÷ bandwidth and ignore latency and protocol overhead.
| Tier | Bandwidth | Time |
|---|---|---|
| HBM3e (B200) | 8 TB/s | 125.0 ms |
| NVLink 5 (per GPU) | 1.8 TB/s | 555.6 ms |
| PCIe Gen5 x16 | 128 GB/s | 7.81 s |
| InfiniBand 800G (XDR) | 100 GB/s | 10.00 s |
| InfiniBand 400G (NDR) | 50 GB/s | 20.00 s |
| GPUDirect Storage (per 400G DPU) | 400 GB/s | 2.50 s |
| NVSwitch domain (aggregate) | 130 TB/s | 7.7 ms |
Does the job fit one NVLink domain?
This calculator asks whether a job fits inside one NVLink domain; the pipe map in the bandwidth cliff above shows how long a transfer takes across every tier - and why the storage and network path, not raw compute, usually decides utilization.
NVLink domain playground - does your job fit one domain?
Each square is one GPU of your job. The green box is one NVLink domain (1.8 TB/s per GPU); every GPU outside it spills onto InfiniBand. Grow the job and watch it spill.
Same 64-GPU job, two domain sizes - tap one to select it
Effective all-to-all BW
1.80 TB/s
100% in-domain · 72-GPU domain
Per-step time
0.04 ms
64 MB / effective BW
all traffic stays on NVLink
~6× faster when the job fits one 72-GPU domain
Keeping all-to-all traffic inside the NVL72 domain (1.8 TB/s/GPU) instead of crossing InfiniBand is a ~18× (800G) to ~36× (400G) bandwidth advantage per GPU - the whole reason 72 GPUs are wired into one domain.
Formula is illustrative: effective BW = f·1800 + (1−f)·cross, where f = min(domain, GPUs)/GPUs - the share of GPUs inside the domain box. Real collectives also pay latency and topology overheads.
DPUs and SuperNICs - keeping GPUs fed
Two specialized NICs sit in every Blackwell tray. The ConnectX-8 moves GPU-to-GPU traffic at up to 800 Gb/s; the BlueField-3 offloads networking, storage, and security off the host CPU so its cores aren't stolen from the work that feeds the GPUs. Toggle the node to see the difference. The next generation, BlueField-4, doubles networking to 800 Gb/s via ConnectX-9 and adds a new headline role - offloading KV-cache “context memory” to a petabyte NVMe tier.
DPU vs SuperNIC - who keeps the GPUs fed
Two NVIDIA networking chips do very different jobs on the same node. The DPU takes north-south infrastructure work off the host CPU; the SuperNIC carries east-west GPU-to-GPU traffic between nodes. Toggle the DPU and watch where the work lands.
Networking - packet processing, RoCE
Storage I/O - NVMe-oF, GPUDirect Storage
Security - encryption, isolation
Management - telemetry, control plane
The host CPU burns cycles moving packets, handling storage and security - cores that should be feeding the GPUs.
Rule of thumb: the SuperNIC moves GPU data fast; the DPU takes infrastructure work off the CPU. Both exist so GPUs stay fed instead of idle.
Specs are NVIDIA datasheet figures (ConnectX-8 up to 800 Gb/s; BlueField-3 up to 400 Gb/s, 16 Armv8.2 cores). In GB200/GB300 trays BlueField-3 pairs with ConnectX-8. Host core count and the split between tasks are illustrative.
BlueField-3 to BlueField-4 - the AI-factory jump
6x compute · 4x larger AI factoriesBlueField-4 is one of the six Vera Rubin platform chips. It roughly doubles networking, quadruples the Arm core count, and shifts the DPU from general datacenter offload to a purpose-built AI-factory infrastructure processor. Both chips are drawn to the same scale.
The headline new role: KV-cache context memory
GPU HBM is small and expensive, so the KV cache overflows on long context. BlueField-4 turns a petabyte NVMe pool into a fast “context memory” tier - storing and reusing KV cache across the whole fleet, at line rate, with encryption handled on the DPU.
Decode writes KV blocks into GPU HBM
Paired with NVIDIA Dynamo (NIXL / STX storage), this persists KV cache off the GPU so a reused prefix never has to be recomputed - directly attacking the memory wall that makes inference memory-bound. Block counts are illustrative.
Specs: 800 Gb/s (ConnectX-9), 6x compute / 3x memory bandwidth vs BlueField-3, PCIe Gen6, Grace CPU, ships 2026 with Vera Rubin, KV-cache / inference-context-memory role. Press-reported (pending NVIDIA datasheet): 64 Arm cores (Neoverse V2), 128 GB LPDDR5.
NVIDIA networking: Spectrum-X
Spectrum-X is NVIDIA's Ethernet platform for AI - Spectrum-4 switch silicon paired with BlueField/ConnectX SuperNICs, adding adaptive routing and congestion control that standard Ethernet lacks. Two switches anchor the 400G and 800G tiers:
Same 64 ports, twice the speed
Both switches are Spectrum-4 with 64 ports. What changes is the speed of each port - and so the total the switch can move.
Spectrum SN5600
800G64× 800GbE (OSFP), 2U64 ports × 800 Gb/s = 51.2 Tb/s
Spectrum-4 flagship. Can also break out to 128× 400GbE. The switch behind 800G Spectrum-X AI clusters.
Spectrum SN5400
400G64× 400GbE (QSFP-DD)64 ports × 400 Gb/s = 25.6 Tb/s
The 400G Spectrum-X tier - same Spectrum-4 lineage, sized for 400GbE leaf/spine fabrics.
GPUDirect Storage - skipping the CPU
The last mile of keeping GPUs busy is the path from storage into HBM. lets the NIC DMA data straight into GPU memory, bypassing the CPU bounce buffer - higher bandwidth, lower latency, and CPU cores freed for real work. Flip between the two paths to see why it matters for utilization.
GPUDirect Storage - keep the GPUs fed, not idle
Storage throughput is part of GPU utilization. Follow one chunk of data from storage into GPU memory: the standard path copies it twice through CPU memory, GPUDirect Storage DMAs it straight across the PCIe switch.
Copies via CPU memory
0
CPU on the data path
not yet
The data sits in storage - local NVMe or a remote array
Two copies. Data is copied into CPU system memory first (a "bounce buffer"), then copied again into GPU memory - extra hops, CPU overhead, higher latency.
This is why a high-bandwidth data platform (e.g. VAST) plugs straight into the GPU fabric - keeping GPUs fed, not idle.
GPUDirect Storage uses the cuFile API + O_DIRECT; the NIC/storage controller DMAs directly into pinned GPU memory. Conceptual diagram - timing and CPU load are illustrative.
Where the data platform plugs in
All this compute is worthless if it sits idle waiting for data. Training reads enormous datasets and writes frequent checkpoints; inference streams context in and logs out. Storage throughput is a first-class part of GPU utilization - which is exactly where a high-bandwidth data platform earns its place in the rack.