Systems & Infrastructure

Networking & the Data Path

Intermediate

At scale the network is the computer: how GPUs talk to each other, why bandwidth falls off a cliff past the rack, and how data gets from storage into GPU memory.

From one GPU to a SuperPOD: scale up vs scale out

A frontier model never fits on one GPU. NVIDIA's answer is a hierarchy of ever-larger “single” computers - GPU, node, NVL72 rack, SuperPOD. Up to the rack, GPUs share one domain at 1.8 TB/s per GPU; past it, racks are joined over InfiniBand or Spectrum-X at roughly 50-100 GB/s per GPU. Each step out buys scale and gives up bandwidth.

That hierarchy splits into two moves. Scaling up grows the NVLink domain into one bigger GPU; scaling out stitches racks together over InfiniBand or Ethernet. Zoom from a single node to a SuperPOD and watch where the fat NVLink links give way to the thin network - and where storage bypasses the CPU entirely.

Scale up, then scale out - one zoom from GPU to SuperPOD

Step the camera back. Each level folds into a single tile of the next, and every link is drawn on one bandwidth scale - so you can watch the fat NVLink pipes give way to the thin network.

Scale upScale out
GraceCPUC2CBlackwell GPUHBM either sideNVLink

01 / 04

1 GPU

The unit of compute

A single accelerator with its own HBM. Once a model or its KV cache exceeds this memory, you must scale out.

NVLink domain
1 GPU
Widest link
NVLink 1.8 TB/s leaves the package

Grace NVLink-C2C

A coherent ~900 GB/s link fuses each Grace CPU to its Blackwell GPUs, with LPDDR5X acting as a capacity tier behind HBM - CPU and GPU share one memory space without crossing PCIe.

1/4Still inside one NVLink domain
NVLink 1.8 TB/s per GPU NVLink-C2C ~900 GB/s InfiniBand / Spectrum-X ~50–100 GB/s per GPU GPUDirect Storage path

Scale up - grow the NVLink domain, make one bigger GPU. Bandwidth-rich (1.8 TB/s/GPU), but capped at 72 GPUs per domain today.

Scale out - add racks over InfiniBand / Ethernet. Near-unlimited size, but each hop carries ~18–36× less bandwidth than NVLink - so keep tight traffic in-domain and spread the loosely-coupled work out.

Schematic: tile counts and positions are illustrative; link widths share one scale with a 1 px floor.

The interconnect

At scale, the network is the computer. How fast GPUs exchange gradients and activations decides whether a 10,000-GPU cluster behaves like 10,000 GPUs or a fraction of that.

Which link does what - four interconnects on one path

Two nodes and the network between them. Pick an interconnect to light up where it sits and send a packet along the traffic it carries.

Quantum IBQuantum IBNode ANICNVSwitchNode BNICNVSwitch

NVLink

GPU ↔ GPU, inside a node

Direct high-bandwidth links between GPUs. 900 GB/s on Hopper, 1.8 TB/s per GPU on Blackwell - orders of magnitude faster than PCIe.

Inside the node the links are fat; between nodes they are thin. NVLink and NVSwitch keep GPUs talking at terabytes per second; InfiniBand or Spectrum-X carries everything that has to leave. The bandwidth cliff further down puts numbers on that gap.

Schematic: 4 GPUs per node and 2 spine switches are drawn for clarity.

The bandwidth cliff

Those interconnects aren't interchangeable - they span more than two orders of magnitude in bandwidth. Plotting them on a log scale shows why the topology is built the way it is: stay on NVLink and you move terabytes per second; cross InfiniBand and you fall off a cliff.

The bandwidth cliff - a pipe map from HBM outward

Each pipe is one hop further from the GPU die, drawn as thick as its bandwidth. Pick a payload and tap a pipe: the ruler below shows how long the same transfer takes on every tier (time = size ÷ bandwidth).

1.00 TB
1 MB1 GB1 TB
NVLINK DOMAINHOSTSCALE-OUTSTORAGEdomain edge · 18× to 36× dropHBM3e8 TB/sbaselineNVLink 51.8 TB/s4.4× slowerPCIe Gen5128 GB/s63× slowerIB 800G100 GB/sIB 400G50 GB/sGDS ×8 DPUs400 GB/s20× slowerNVSwitch 130 TB/s rack

1.00 TB over InfiniBand 400G (NDR)

20.00 s

50 GB/s · 160× slower than HBM

Quantum-2 InfiniBand, scale-out per port.

Same payload inside the NVLink domain

555.6 ms

36× longer once it leaves the domain over IB 400G.

Tap any pipe or ruler row to change the tier. Pick GPUDirect Storage to spread it across more DPUs.

Same 1.00 TB transfer, every tier · log time

HBM3e125.0 msNVLink 5555.6 msPCIe Gen5 x167.81 sInfiniBand 800G (XDR)10.00 sInfiniBand 400G (NDR)20.00 sGDS · 8 DPUs2.50 sNVSwitch (whole rack)7.7 ms1 ns10 ns100 ns1 µs10 µs100 µs1 ms10 ms100 ms1 s10 s100 s1 minlog time →
NVLink domainhost path (PCIe)scale-out networkstorage (GPUDirect)NVSwitch aggregate (whole rack)

1.00 TB clears HBM in 125.0 ms but takes 20.00 s over a single 400G InfiniBand link. The biggest step is the NVLink-domain edge - a 18× (800G) to 36× (400G) drop - which is why 72 GPUs share one NVLink domain. When thousands of GPUs stall on checkpoints, data loading, or KV reloads, the bottleneck is the storage and network path - exactly where a high-bandwidth data platform keeps the fleet busy.

Order-of-magnitude figures. Pipe thickness ∝ √bandwidth. Times are size ÷ bandwidth and ignore latency and protocol overhead.

Transfer time for 1.00 TB on each tier
TierBandwidthTime
HBM3e (B200)8 TB/s125.0 ms
NVLink 5 (per GPU)1.8 TB/s555.6 ms
PCIe Gen5 x16128 GB/s7.81 s
InfiniBand 800G (XDR)100 GB/s10.00 s
InfiniBand 400G (NDR)50 GB/s20.00 s
GPUDirect Storage (per 400G DPU)400 GB/s2.50 s
NVSwitch domain (aggregate)130 TB/s7.7 ms

DPUs and SuperNICs - keeping GPUs fed

Two specialized NICs sit in every Blackwell tray. The ConnectX-8 moves GPU-to-GPU traffic at up to 800 Gb/s; the BlueField-3 offloads networking, storage, and security off the host CPU so its cores aren't stolen from the work that feeds the GPUs. Toggle the node to see the difference. The next generation, BlueField-4, doubles networking to 800 Gb/s via ConnectX-9 and adds a new headline role - offloading KV-cache “context memory” to a petabyte NVMe tier.

DPU vs SuperNIC - who keeps the GPUs fed

Two NVIDIA networking chips do very different jobs on the same node. The DPU takes north-south infrastructure work off the host CPU; the SuperNIC carries east-west GPU-to-GPU traffic between nodes. Toggle the DPU and watch where the work lands.

north-southPCIe - every packeteast-westNorth-south networkstorage · management · tenantsEast-west fabricGPUs in other nodesPlain NIC (no DPU)passes it all upno cores of its own - the CPU does itConnectX-8 SuperNICup to 800 Gb/sRDMA / RoCE + GPUDirectGPU-to-GPU traffic between nodes -the same with or without a DPUHost CPUsaturated6 of 16 cores feed the GPUs (illustrative)GPUstalk to other nodes via the SuperNIC ↑GPUGPUGPUGPUNetworkingStorage I/OSecurityManagement

Networking - packet processing, RoCE

Storage I/O - NVMe-oF, GPUDirect Storage

Security - encryption, isolation

Management - telemetry, control plane

The host CPU burns cycles moving packets, handling storage and security - cores that should be feeding the GPUs.

Infrastructure work (north-south)Handled on the DPUApp / data-prep feeding GPUsGPU-to-GPU traffic (east-west)

Rule of thumb: the SuperNIC moves GPU data fast; the DPU takes infrastructure work off the CPU. Both exist so GPUs stay fed instead of idle.

Specs are NVIDIA datasheet figures (ConnectX-8 up to 800 Gb/s; BlueField-3 up to 400 Gb/s, 16 Armv8.2 cores). In GB200/GB300 trays BlueField-3 pairs with ConnectX-8. Host core count and the split between tasks are illustrative.

BlueField-3 to BlueField-4 - the AI-factory jump

6x compute · 4x larger AI factories

BlueField-4 is one of the six Vera Rubin platform chips. It roughly doubles networking, quadruples the Arm core count, and shifts the DPU from general datacenter offload to a purpose-built AI-factory infrastructure processor. Both chips are drawn to the same scale.

network400 Gb/sBlueField-3ships 2022cores16 Arm (Cortex-A78)ConnectX-7integrated NICdatacenter offloadPCIe Gen5 host linknetwork800 Gb/sBlueField-4ships 2026 · Vera Rubin64 Arm (Grace)ConnectX-9integrated NICAI-factory infrastructurePCIe Gen6 host link
2x networking400 → 800 Gb/s
4x Arm cores16 → 64
Gen6 host linkPCIe Gen5 → Gen6
CX-9 integrated NICConnectX-7 → ConnectX-9
The headline new role: KV-cache context memory

GPU HBM is small and expensive, so the KV cache overflows on long context. BlueField-4 turns a petabyte NVMe pool into a fast “context memory” tier - storing and reusing KV cache across the whole fleet, at line rate, with encryption handled on the DPU.

offloadreuseGPU HBMsmall, fastBlueField-4800 Gb/s · NVMe-oFidle - encrypts in lineNVMe context memorypetabytes · shared by the fleetABCD
1/7Decode writes KV blocks into GPU HBM

Decode writes KV blocks into GPU HBM

KV blockwaiting - no room in HBMblock a returning request reusesother sessions' context

Paired with NVIDIA Dynamo (NIXL / STX storage), this persists KV cache off the GPU so a reused prefix never has to be recomputed - directly attacking the memory wall that makes inference memory-bound. Block counts are illustrative.

Specs: 800 Gb/s (ConnectX-9), 6x compute / 3x memory bandwidth vs BlueField-3, PCIe Gen6, Grace CPU, ships 2026 with Vera Rubin, KV-cache / inference-context-memory role. Press-reported (pending NVIDIA datasheet): 64 Arm cores (Neoverse V2), 128 GB LPDDR5.

NVIDIA networking: Spectrum-X

Spectrum-X is NVIDIA's Ethernet platform for AI - Spectrum-4 switch silicon paired with BlueField/ConnectX SuperNICs, adding adaptive routing and congestion control that standard Ethernet lacks. Two switches anchor the 400G and 800G tiers:

Same 64 ports, twice the speed

Both switches are Spectrum-4 with 64 ports. What changes is the speed of each port - and so the total the switch can move.

Spectrum SN5600

800G64× 800GbE (OSFP), 2U
51.2 Tb/s

64 ports × 800 Gb/s = 51.2 Tb/s

Spectrum-4 flagship. Can also break out to 128× 400GbE. The switch behind 800G Spectrum-X AI clusters.

Spectrum SN5400

400G64× 400GbE (QSFP-DD)
25.6 Tb/s

64 ports × 400 Gb/s = 25.6 Tb/s

The 400G Spectrum-X tier - same Spectrum-4 lineage, sized for 400GbE leaf/spine fabrics.

800G port400G portFaceplates are schematic.

GPUDirect Storage - skipping the CPU

The last mile of keeping GPUs busy is the path from storage into HBM. lets the NIC DMA data straight into GPU memory, bypassing the CPU bounce buffer - higher bandwidth, lower latency, and CPU cores freed for real work. Flip between the two paths to see why it matters for utilization.

GPUDirect Storage - keep the GPUs fed, not idle

Storage throughput is part of GPU utilization. Follow one chunk of data from storage into GPU memory: the standard path copies it twice through CPU memory, GPUDirect Storage DMAs it straight across the PCIe switch.

1copy 1: into DRAM2copy 2: out to GPUCPU + system DRAMidlebounce bufferCPU loadqualitative, not measuredStorage / NVMelocal or remoteNICnetwork adapterPCIe switchGPU, NIC, CPU attachGPU HBMwhere compute happens
Standard pathin storage

Copies via CPU memory

0

CPU on the data path

not yet

1/5The data sits in storage - local NVMe or a remote array

The data sits in storage - local NVMe or a remote array

Two copies. Data is copied into CPU system memory first (a "bounce buffer"), then copied again into GPU memory - extra hops, CPU overhead, higher latency.

Standard path (via bounce buffer)GPUDirect Storage (direct DMA)PCIe / network link

This is why a high-bandwidth data platform (e.g. VAST) plugs straight into the GPU fabric - keeping GPUs fed, not idle.

GPUDirect Storage uses the cuFile API + O_DIRECT; the NIC/storage controller DMAs directly into pinned GPU memory. Conceptual diagram - timing and CPU load are illustrative.

Where the data platform plugs in

All this compute is worthless if it sits idle waiting for data. Training reads enormous datasets and writes frequent checkpoints; inference streams context in and logs out. Storage throughput is a first-class part of GPU utilization - which is exactly where a high-bandwidth data platform earns its place in the rack.