Systems & Infrastructure
GPUs & Racks
FundamentalThe GPUs and racks AI runs on: what each generation adds, how to read the numbers, and why power - not chips - now sets the limit.
The GPU lineup
Three generations are in play today: Hopper (H100, H200), Blackwell (B200, GB200) and Blackwell Ultra (B300, GB300), with Ada Lovelace (L40S) and the RTX PRO 6000 covering inference, graphics, and simulation. Each generation moves the same three levers: more memory, more bandwidth, and faster links between GPUs.
Nine GPUs, three levers, one scale
Switch the lever to see how each generation moves memory, bandwidth, and GPU-to-GPU links. Tap a GPU for its spec card.
Memory capacity: how big a model and its KV cache can live on the chip.
B200
2024–25Blackwell
- Memory
- 192 GB HBM3e
- Bandwidth
- ~8 TB/s
- NVLink/GPU
- 1.8 TB/s
- Per system
- 8 / HGX node
Dual-die design; adds FP4 for inference throughput.
GB200 and GB300 are Grace-Blackwell superchips, so their memory and bandwidth cover two GPUs.
Three things to know about Vera Rubin
- It is an NVL72 rack again. Early material called it NVL144, counting dies; NVIDIA now counts packages. Vera Rubin NVL72 holds 72 Rubin GPUs (two dies each) and 36 Vera CPUs - the same footprint and NVLink domain size as today's GB300 NVL72. It began shipping in limited volume in September 2026.
- Named after an astronomer. Vera Rubin's work on galaxy rotation gave the first strong evidence for dark matter - the platform pairs a Vera CPU with a Rubin GPU.
- HBM4 keeps memory the headline. 288 GB per GPU at 19.2 TB/s - 2.4x Blackwell Ultra's bandwidth on the same capacity. Inference stays memory-bound, so bandwidth and capacity (not just FLOPS) decide how much and context a GPU can hold.
The roadmap ahead
NVIDIA now ships a new data-center architecture roughly every year, so sizing a multi-year build means knowing what's next. Three things define the next two years: memory moves to HBM4, the rack becomes the unit of compute, and inference begins to disaggregate its and phases across the fleet. Click an era to see what changes.
NVIDIA datacenter GPU roadmap - where AI infrastructure is heading
The unit of compute keeps growing - from one GPU, to a node, to a rack, to a multi-rack NVLink domain - while memory (HBM4) and disaggregated inference (separate context vs decode silicon) define the next two years. Each dot below is one GPU in an era's NVLink domain.
H100, H200
FP8 Transformer Engine; the workhorse of the LLM boom. H200 added 141 GB HBM3e for long context.
8
8-GPU node
Roadmap dates and figures are NVIDIA-announced and subject to change; future-gen numbers are estimates. Era positions are approximate within their year. Hopper's NVLink domain is an 8-GPU node, so its power is per node.
Reading a spec sheet
Four numbers tell you most of what you need to know about a data-center GPU:
Decode a datasheet - the four lines that matter
A trimmed B200 datasheet. Tap a circled line to see what it tells you and how it compares.
* dense; sparse figures run ~2× higher
Line 1 of 4
Memory capacity (HBM)
How big a model - and its KV cache - can live on one GPU. The first wall you hit. 80 GB on H100, 192 GB on B200, 288 GB on B300.
GB of HBM per GPU
Compute throughput
Peak tensor-core throughput per GPU, by precision. Lower precision trades numerical range for speed: each step down (FP16 → → ) roughly doubles the rate, which is why inference has marched toward FP4 on Blackwell.
| GPU | FP4 | FP8 / INT8 | BF16 / FP16 | FP64 |
|---|---|---|---|---|
| H100 | - | 1.0 PF | 0.5 PF | 67 TF |
| H200 | - | 1.0 PF | 0.5 PF | 67 TF |
| B200 | 9 PF | 4.5 PF | 2.25 PF | 40 TF |
| GB200 | 20 PF | 10 PF | 5 PF | 90 TF |
| B300 | 13.5 PF | 4.5 PF | 2.25 PF | 1.25 TF |
| GB300 | 30 PF | 10 PF | 5 PF | 2.8 TF |
| Rubin (VR200) | 35 PF | 17.5 PF | 4 PF | - |
| RTX PRO 6000 | 4 PF | 2 PF | 1 PF | - |
| L40S | - | 1.5 PF | 0.36 PF | - |
Representative dense peak figures (PF = PFLOPS, TF = TFLOPS); structured sparsity roughly doubles them. These are per-chip numbers - GB200 / GB300 rows are per Grace-Blackwell superchip (two GPUs), and their dies clock higher than the air-cooled HGX boards. A full NVL72 rack multiplies these across 72 GPUs (e.g. ~1.44 EF FP4 sparse for GB200 NVL72), so the per-rack totals in the rack diagram below are far larger and do not contradict these figures. Blackwell Ultra (B300, GB300) raises FP4 but keeps Blackwell's FP8 and BF16 rates, and cuts INT8 and FP64 sharply - so for those two it is a step back from B200, not forward. Confirm against current NVIDIA datasheets before sizing.
Nodes and racks
GPUs ship in standardized building blocks. What matters for sizing is how many GPUs share one fast memory domain - and how much that adds up to, because that pooled memory is what a single large model or KV cache can spread across.
Pooled HBM per NVLink domain
One linear scale from an 8-GPU node to a full rack. Tap a system to see how many GPUs share its memory pool.
GB200 NVL72
72× B200~13.5 TBtotal HBM
72 × 192 GB as one NVLink domain - a rack that acts as one GPU.
One NVLink domain · 72 GPUs
The step change is the rack: every 8-GPU node here stays at or under ~2.3 TB, while one NVL72 pools ~13.5-20.7 TB that a single large model or KV cache can spread across.
Vera Rubin NVL72 (dashed) began shipping in limited volume in September 2026; its pool matches GB300 NVL72.
The NVL72 rack, up close
The flagship building block deserves a closer look. An NVL72 rack groups its 9 NVLink switch trays in the center, flanked by two banks of compute trays, with power shelves at top and bottom - a copper spine fuses all 72 GPUs into one all-to-all NVLink-5 domain. The GB200 and GB300 builds share this rack; GB300 raises HBM, FP4, and power.
Inside an NVL72 rack - a rack that acts as one GPU
Left, the physical rack: 18 compute trays in two banks, the 9 NVLink switch trays in the center, power at top and bottom. Right, the same 72 GPUs as one fabric. Tap two GPUs (or a tray) to follow the path.
Physical rack
Logical fabric
GPU 6 (tray 2) → GPU 63 (tray 16): through the NVSwitch plane, at full NVLink speed.
Why the switches sit in the middle: centering them keeps every compute tray within a short, equal copper run of the NVLink fabric. No in-rack optics - the spine is all copper - and the packet doesn't care which tray it lands in.
Pooled HBM
~13.5–20 TB
FP4 compute
~1.1–1.44 EF FP4
Aggregate HBM BW
~570 TB/s
Power (liquid-cooled)
~120–142 kW
GB200/300 NVL72 (Blackwell (B200) → Blackwell Ultra (B300)). Both share the same rack - GB300 raises HBM and FP4. Figures span both as per-rack totals, not the per-chip numbers in the compute table above. Tray drawing is schematic.
How you actually buy it
The same GPUs reach customers through very different consumption models - from a bare baseboard an OEM builds a server around, to a turnkey validated cluster, to rented capacity in a neocloud. The choice is usually about who owns the integration risk and the CapEx, not the silicon - and a reference architecture is the thread that de-risks every path.
Deployment model explorer - how enterprises actually buy NVIDIA AI
Same silicon, very different go-to-market. Each model sits on two questions: do you own the hardware or rent it, and who owns the integration? Pick one to see who it's for, how you buy it, and the tradeoff you're signing up for.
HGX
The 8-GPU baseboard- Who it's for
- OEMs/ODMs, neoclouds, and most enterprises.
- How you buy
- NVIDIA sells the 8-GPU baseboard; partners (Supermicro, Dell, etc.) build the server around it.
- Tradeoff
- Maximum choice and best price - but you (or your vendor) own the integration.
Reference architecture: a pre-validated blueprint spanning compute, network, storage and software - tested for thermals, power, mechanicals and signal integrity. It removes integration risk so an AI factory goes live in weeks, not months - the single most deployment-relevant concept for a buyer.
Placements are qualitative, read from each model's tradeoffs. Pricing figures are illustrative market estimates (neocloud vs hyperscaler H100-equivalent hourly).
The power wall - and why racks went liquid
A single NVL72 rack draws 120–142 kW - five to six times what an air-cooled rack can dissipate. That's why Blackwell racks ship direct-to-chip liquid-cooled by default, and why per-GPU power keeps climbing every generation. Toggle the views to see both walls - and where the roadmap (600 kW, then ~1 MW racks) is heading.
The power wall - why cooling now sets the design
Rack power has blown past the air-cooling ceiling, and per-GPU TDP keeps climbing. Direct-to-chip liquid cooling stopped being optional somewhere around 25 kW per rack.
Heat to remove · GB200 NVL72
132 kW
5.3× what air can remove from one rack - so the heat leaves in liquid.
5–6× past the air ceiling - direct-to-chip cold plates + CDU mandatory.
Direct-to-chip cold plates sit on each GPU; an in-row CDU (coolant distribution unit) isolates the clean technology loop on the chips from the facility water loop.
Rack/TDP figures are approximate, from NVIDIA and vendor reference designs; Vera Rubin is an estimate range; the Rubin Ultra Kyber figure (~600 kW, 800 VDC, heading toward 1 MW racks) is NVIDIA's 2025 roadmap direction. The cooling schematic is simplified.
Once power - not chips - caps how much you can deploy, the metric that matters shifts from FLOPS to tokens per watt and cost per token. On those, each generation is a step change: Blackwell delivers far more tokens per megawatt and a fraction of Hopper's cost per token. Drag the slider to see how that reshapes the bill at scale.
Inference economics - power is the new constraint
At scale the KPIs aren't FLOPS - they're cost-per-token and tokens-per-watt. Slide your monthly token volume: both bills move along a log axis, and the ~35× gap between them stays the same width at every scale. The same megawatt still serves 50× the tokens.
Monthly inference bill
fixed log axis - each gridline is 10×
Hopper $4.20M/mo at $4.20 per 1M tokens
Blackwell $120.0K/mo at $0.12 per 1M tokens
Flip it around: Hopper's $4.20M would buy 35T tokens a month on Blackwell - 35× the volume for the same spend.
Same megawatt - how many token streams?
relative units: 1 stream = Hopper's output per MW
1 stream(baseline)
50 streams(50×)
Serving 1.0T tokens a month on Hopper takes 50× the power Blackwell needs - so on a fixed power budget, Blackwell serves 50× the tokens.
Why this matters: at scale, inference revenue is gated by your megawatt budget, not GPU sticker price - so the metric that matters is tokens per watt and per dollar, not FLOPS.
Figures are illustrative round numbers (Blackwell vs Hopper, reasoning-model inference): ~$0.12 vs ~$4.20 per 1M tokens, >50× tokens/MW. Real results vary by model, precision (FP4), and workload.