Systems & Infrastructure

GPUs & Racks

Fundamental

The GPUs and racks AI runs on: what each generation adds, how to read the numbers, and why power - not chips - now sets the limit.

The GPU lineup

Three generations are in play today: Hopper (H100, H200), Blackwell (B200, GB200) and Blackwell Ultra (B300, GB300), with Ada Lovelace (L40S) and the RTX PRO 6000 covering inference, graphics, and simulation. Each generation moves the same three levers: more memory, more bandwidth, and faster links between GPUs.

Nine GPUs, three levers, one scale

Switch the lever to see how each generation moves memory, bandwidth, and GPU-to-GPU links. Tap a GPU for its spec card.

Memory capacity: how big a model and its KV cache can live on the chip.

B200

2024–25

Blackwell

Memory
192 GB HBM3e
Bandwidth
~8 TB/s
NVLink/GPU
1.8 TB/s
Per system
8 / HGX node

Dual-die design; adds FP4 for inference throughput.

HopperBlackwellBlackwell UltraVera RubinBlackwell (RTX)Ada Lovelacesplit bar = superchip, two GPUs

GB200 and GB300 are Grace-Blackwell superchips, so their memory and bandwidth cover two GPUs.

Three things to know about Vera Rubin

  • It is an NVL72 rack again. Early material called it NVL144, counting dies; NVIDIA now counts packages. Vera Rubin NVL72 holds 72 Rubin GPUs (two dies each) and 36 Vera CPUs - the same footprint and NVLink domain size as today's GB300 NVL72. It began shipping in limited volume in September 2026.
  • Named after an astronomer. Vera Rubin's work on galaxy rotation gave the first strong evidence for dark matter - the platform pairs a Vera CPU with a Rubin GPU.
  • HBM4 keeps memory the headline. 288 GB per GPU at 19.2 TB/s - 2.4x Blackwell Ultra's bandwidth on the same capacity. Inference stays memory-bound, so bandwidth and capacity (not just FLOPS) decide how much and context a GPU can hold.

The roadmap ahead

NVIDIA now ships a new data-center architecture roughly every year, so sizing a multi-year build means knowing what's next. Three things define the next two years: memory moves to HBM4, the rack becomes the unit of compute, and inference begins to disaggregate its and phases across the fleet. Click an era to see what changes.

NVIDIA datacenter GPU roadmap - where AI infrastructure is heading

The unit of compute keeps growing - from one GPU, to a node, to a rack, to a multi-rack NVLink domain - while memory (HBM4) and disaggregated inference (separate context vs decode silicon) define the next two years. Each dot below is one GPU in an era's NVLink domain.

1/6Hopper · 2022–24 · 8-GPU NVLink domain
2022202320242025202620272028Hopper8 GPUsBlackwell72 GPUsBlackwell Ultra72 GPUsVera Rubin72 GPUsRubin Ultra144 GPUsFeynmanTBD?HBM per GPU80–141 GB192 GB288 GB288 GB--HBM bandwidth per GPU3.35–4.8 TB/s~8 TB/s~8 TB/s19.2 TB/s--NVLink per GPU0.9 TB/s1.8 TB/s1.8 TB/s3 TB/s--Power per NVLink domain10 kW node132 kW142 kW~190-230 kW est.~600 kW-FP4 per NVLink domain-~1.44 EF-~3.6 EF~15 EF-
shipping (solid)roadmap (outlined)announced (dashed)- not stated on this page
Hopper2022–24 · shipping
GPUs

H100, H200

FP8 Transformer Engine; the workhorse of the LLM boom. H200 added 141 GB HBM3e for long context.

NVLink domain

8

8-GPU node

Roadmap dates and figures are NVIDIA-announced and subject to change; future-gen numbers are estimates. Era positions are approximate within their year. Hopper's NVLink domain is an 8-GPU node, so its power is per node.

Reading a spec sheet

Four numbers tell you most of what you need to know about a data-center GPU:

Decode a datasheet - the four lines that matter

A trimmed B200 datasheet. Tap a circled line to see what it tells you and how it compares.

NVIDIA B200datasheet
ArchitectureBlackwell
NVLink / GPU1.8 TB/s
Per system8 / HGX node

* dense; sparse figures run ~2× higher

Line 1 of 4

Memory capacity (HBM)

How big a model - and its KV cache - can live on one GPU. The first wall you hit. 80 GB on H100, 192 GB on B200, 288 GB on B300.

H10080 GB
B200192 GB
B300288 GB

GB of HBM per GPU

Compute throughput

Peak tensor-core throughput per GPU, by precision. Lower precision trades numerical range for speed: each step down (FP16 → → ) roughly doubles the rate, which is why inference has marched toward FP4 on Blackwell.

GPUFP4FP8 / INT8BF16 / FP16FP64
H100-1.0 PF0.5 PF67 TF
H200-1.0 PF0.5 PF67 TF
B2009 PF4.5 PF2.25 PF40 TF
GB20020 PF10 PF5 PF90 TF
B30013.5 PF4.5 PF2.25 PF1.25 TF
GB30030 PF10 PF5 PF2.8 TF
Rubin (VR200)35 PF17.5 PF4 PF-
RTX PRO 60004 PF2 PF1 PF-
L40S-1.5 PF0.36 PF-

Representative dense peak figures (PF = PFLOPS, TF = TFLOPS); structured sparsity roughly doubles them. These are per-chip numbers - GB200 / GB300 rows are per Grace-Blackwell superchip (two GPUs), and their dies clock higher than the air-cooled HGX boards. A full NVL72 rack multiplies these across 72 GPUs (e.g. ~1.44 EF FP4 sparse for GB200 NVL72), so the per-rack totals in the rack diagram below are far larger and do not contradict these figures. Blackwell Ultra (B300, GB300) raises FP4 but keeps Blackwell's FP8 and BF16 rates, and cuts INT8 and FP64 sharply - so for those two it is a step back from B200, not forward. Confirm against current NVIDIA datasheets before sizing.

Nodes and racks

GPUs ship in standardized building blocks. What matters for sizing is how many GPUs share one fast memory domain - and how much that adds up to, because that pooled memory is what a single large model or KV cache can spread across.

Pooled HBM per NVLink domain

One linear scale from an 8-GPU node to a full rack. Tap a system to see how many GPUs share its memory pool.

05101520TB of HBM

GB200 NVL72

72× B200~13.5 TBtotal HBM

72 × 192 GB as one NVLink domain - a rack that acts as one GPU.

One NVLink domain · 72 GPUs

The step change is the rack: every 8-GPU node here stays at or under ~2.3 TB, while one NVL72 pools ~13.5-20.7 TB that a single large model or KV cache can spread across.

Vera Rubin NVL72 (dashed) began shipping in limited volume in September 2026; its pool matches GB300 NVL72.

The NVL72 rack, up close

The flagship building block deserves a closer look. An NVL72 rack groups its 9 NVLink switch trays in the center, flanked by two banks of compute trays, with power shelves at top and bottom - a copper spine fuses all 72 GPUs into one all-to-all NVLink-5 domain. The GB200 and GB300 builds share this rack; GB300 raises HBM, FP4, and power.

Inside an NVL72 rack - a rack that acts as one GPU

Left, the physical rack: 18 compute trays in two banks, the 9 NVLink switch trays in the center, power at top and bottom. Right, the same 72 GPUs as one fabric. Tap two GPUs (or a tray) to follow the path.

GB200/300
72 GPUs36 Grace18 NVSwitch chipsone NVLink-5 domain

Physical rack

mgmt switch · 2× Spectrum Ethernetcopper spine · >5,000 cables · 130 TB/s1234567891011121314151617183 power shelves33 kW · 6× 5.5 kW PSUCompute trays 1-92× Grace + 4 Blackwell9 switch trays2× NVSwitch eachCompute trays 10-182× Grace + 4 Blackwell3 power shelves33 kW each

Logical fabric

123456789101112131415161718NVSwitch9 trays · 18 chipsGPU 1 (tray 1)GPU 2 (tray 1)GPU 3 (tray 1)GPU 4 (tray 1)GPU 5 (tray 2)GPU 6 (tray 2)GPU 7 (tray 2)GPU 8 (tray 2)GPU 9 (tray 3)GPU 10 (tray 3)GPU 11 (tray 3)GPU 12 (tray 3)GPU 13 (tray 4)GPU 14 (tray 4)GPU 15 (tray 4)GPU 16 (tray 4)GPU 17 (tray 5)GPU 18 (tray 5)GPU 19 (tray 5)GPU 20 (tray 5)GPU 21 (tray 6)GPU 22 (tray 6)GPU 23 (tray 6)GPU 24 (tray 6)GPU 25 (tray 7)GPU 26 (tray 7)GPU 27 (tray 7)GPU 28 (tray 7)GPU 29 (tray 8)GPU 30 (tray 8)GPU 31 (tray 8)GPU 32 (tray 8)GPU 33 (tray 9)GPU 34 (tray 9)GPU 35 (tray 9)GPU 36 (tray 9)GPU 37 (tray 10)GPU 38 (tray 10)GPU 39 (tray 10)GPU 40 (tray 10)GPU 41 (tray 11)GPU 42 (tray 11)GPU 43 (tray 11)GPU 44 (tray 11)GPU 45 (tray 12)GPU 46 (tray 12)GPU 47 (tray 12)GPU 48 (tray 12)GPU 49 (tray 13)GPU 50 (tray 13)GPU 51 (tray 13)GPU 52 (tray 13)GPU 53 (tray 14)GPU 54 (tray 14)GPU 55 (tray 14)GPU 56 (tray 14)GPU 57 (tray 15)GPU 58 (tray 15)GPU 59 (tray 15)GPU 60 (tray 15)GPU 61 (tray 16)GPU 62 (tray 16)GPU 63 (tray 16)GPU 64 (tray 16)GPU 65 (tray 17)GPU 66 (tray 17)GPU 67 (tray 17)GPU 68 (tray 17)GPU 69 (tray 18)GPU 70 (tray 18)GPU 71 (tray 18)GPU 72 (tray 18)AB

GPU 6 (tray 2) → GPU 63 (tray 16): through the NVSwitch plane, at full NVLink speed.

Blackwell GPUGrace CPUNVLink switch tray (2× NVSwitch)power shelfcopper spine

Why the switches sit in the middle: centering them keeps every compute tray within a short, equal copper run of the NVLink fabric. No in-rack optics - the spine is all copper - and the packet doesn't care which tray it lands in.

Pooled HBM

~13.5–20 TB

FP4 compute

~1.1–1.44 EF FP4

Aggregate HBM BW

~570 TB/s

Power (liquid-cooled)

~120–142 kW

GB200/300 NVL72 (Blackwell (B200) → Blackwell Ultra (B300)). Both share the same rack - GB300 raises HBM and FP4. Figures span both as per-rack totals, not the per-chip numbers in the compute table above. Tray drawing is schematic.

How you actually buy it

The same GPUs reach customers through very different consumption models - from a bare baseboard an OEM builds a server around, to a turnkey validated cluster, to rented capacity in a neocloud. The choice is usually about who owns the integration risk and the CapEx, not the silicon - and a reference architecture is the thread that de-risks every path.

Deployment model explorer - how enterprises actually buy NVIDIA AI

Same silicon, very different go-to-market. Each model sits on two questions: do you own the hardware or rent it, and who owns the integration? Pick one to see who it's for, how you buy it, and the tradeoff you're signing up for.

you integrate  →  turnkey
~$34 vs ~$98/hr
← own it (CapEx)rent it (OpEx) →
small halo: server / systemlarge halo: cluster-scalepre-validated band
HGX
The 8-GPU baseboard
Own (CapEx)You or your OEM integrateserver / baseboard
Who it's for
OEMs/ODMs, neoclouds, and most enterprises.
How you buy
NVIDIA sells the 8-GPU baseboard; partners (Supermicro, Dell, etc.) build the server around it.
Tradeoff
Maximum choice and best price - but you (or your vendor) own the integration.

Reference architecture: a pre-validated blueprint spanning compute, network, storage and software - tested for thermals, power, mechanicals and signal integrity. It removes integration risk so an AI factory goes live in weeks, not months - the single most deployment-relevant concept for a buyer.

Placements are qualitative, read from each model's tradeoffs. Pricing figures are illustrative market estimates (neocloud vs hyperscaler H100-equivalent hourly).

The power wall - and why racks went liquid

A single NVL72 rack draws 120–142 kW - five to six times what an air-cooled rack can dissipate. That's why Blackwell racks ship direct-to-chip liquid-cooled by default, and why per-GPU power keeps climbing every generation. Toggle the views to see both walls - and where the roadmap (600 kW, then ~1 MW racks) is heading.

The power wall - why cooling now sets the design

Rack power has blown past the air-cooling ceiling, and per-GPU TDP keeps climbing. Direct-to-chip liquid cooling stopped being optional somewhere around 25 kW per rack.

1 block = 25 kW, what air can cool in one rackpast the ceiling - liquid requiredfar over the limitair ceiling
rackcold platescoolwarmCDUin-rowcoolwarmfacilitywatertechnology loopclean coolantfacility loopsite water

Heat to remove · GB200 NVL72

132 kW

5.3× what air can remove from one rack - so the heat leaves in liquid.

5–6× past the air ceiling - direct-to-chip cold plates + CDU mandatory.

Direct-to-chip cold plates sit on each GPU; an in-row CDU (coolant distribution unit) isolates the clean technology loop on the chips from the facility water loop.

Rack/TDP figures are approximate, from NVIDIA and vendor reference designs; Vera Rubin is an estimate range; the Rubin Ultra Kyber figure (~600 kW, 800 VDC, heading toward 1 MW racks) is NVIDIA's 2025 roadmap direction. The cooling schematic is simplified.

Once power - not chips - caps how much you can deploy, the metric that matters shifts from FLOPS to tokens per watt and cost per token. On those, each generation is a step change: Blackwell delivers far more tokens per megawatt and a fraction of Hopper's cost per token. Drag the slider to see how that reshapes the bill at scale.

Inference economics - power is the new constraint

At scale the KPIs aren't FLOPS - they're cost-per-token and tokens-per-watt. Slide your monthly token volume: both bills move along a log axis, and the ~35× gap between them stays the same width at every scale. The same megawatt still serves 50× the tokens.

Blackwell vs Hopper
~35× lower cost per token · >50× tokens per megawatt · Blackwell vs Hopper
1.0T
1B1T100T

Monthly inference bill

fixed log axis - each gridline is 10×

$100$1K$10K$100K$1M$10M$100M$1BHopperBlackwell$4.20M$120.0K35× lesssame spend on Blackwell: 35T tokens

Hopper $4.20M/mo at $4.20 per 1M tokens

Blackwell $120.0K/mo at $0.12 per 1M tokens

Flip it around: Hopper's $4.20M would buy 35T tokens a month on Blackwell - 35× the volume for the same spend.

Same megawatt - how many token streams?

relative units: 1 stream = Hopper's output per MW

Hopper (H100)1 MW

1 stream(baseline)

Blackwell (GB300 NVL72)1 MW

50 streams(50×)

Serving 1.0T tokens a month on Hopper takes 50× the power Blackwell needs - so on a fixed power budget, Blackwell serves 50× the tokens.

Why this matters: at scale, inference revenue is gated by your megawatt budget, not GPU sticker price - so the metric that matters is tokens per watt and per dollar, not FLOPS.

Figures are illustrative round numbers (Blackwell vs Hopper, reasoning-model inference): ~$0.12 vs ~$4.20 per 1M tokens, >50× tokens/MW. Real results vary by model, precision (FP4), and workload.