Systems & Infrastructure
NVIDIA Infrastructure
FundamentalThe GPUs, racks, and networks that AI actually runs on - and how to match them to a workload.
Many jobs, one fabric
Almost everything in AI infrastructure exists to keep expensive GPUs busy. But “keeping them busy” means different things for different workloads - and the right hardware choice follows the workload, not the other way around. These are the major classes you will be sizing for:
Seven workloads, five kinds of hardware
Each dot marks a representative sweet spot. Columns grow from one PCIe card to a rack that acts as one GPU. Tap a workload to read it, or a column to see which workloads use it.
Training
Pretraining and large fine-tunes: thousands of GPUs in lockstep, bound by interconnect and memory.
Representative picks: GB200/GB300 NVL72, H100/H200 clusters
Read down the columns: the PCIe cards and Hopper nodes turn up across the most workloads, while the NVL72 rack is the pick for jobs that need one very large NVLink domain.
The hardware shown is a representative sweet spot, not a limit - any modern data-center GPU can serve inference; the right one depends on model size, context length, and your latency and cost budget.
Go deeper
The hardware story splits in two: the GPUs and racks that do the compute, and the networks and data path that keep them fed. Each has its own page.
Six paradigms driving the build-out
Above the technical workloads sits a market-level view: the kinds of AI demand pulling infrastructure into data centers. The first two are two faces of the same plant - an AI factory manufactures intelligence (raw materials: data + electricity; product: a model), and the token factory turns that model into tokens, where the currency is tokens per watt. (NVIDIA often uses “AI factory” for the whole operation, training and inference alike - the split below is about emphasis.)
One plant, four kinds of demand
The AI factory and the token factory are two stages of one plant. The other four paradigms are demand pulling on it. Tap any tile to read it.
The plant
AI Factory
Training
Building the model. NVIDIA reframes the data center as a plant that manufactures intelligence - raw materials are data and electricity, the product is a trained model. Thousands of GPUs in lockstep, interconnect-bound.
Hardware: GB200/GB300 NVL72, H100/H200 clusters
Training & fine-tuning →The flagship hardware story across these: GB200 and GB300 NVL72 racks, where 72 GPUs act as one accelerator and FP4 maximizes tokens per watt for large-scale reasoning inference.
The real moat: CUDA and the full stack
A competitor can match a chip far sooner than it can match twenty years of CUDA. The durable lock-in is the software stack - the CUDA platform, the sprawling CUDA-X libraries (cuDNN, cuBLAS, cuDF, cuML, cuVS, cuGraph, NCCL and hundreds more), and the frameworks and microservices built on top. Explore the layers to see why the ecosystem, not the silicon, is the hard thing to leave.
The NVIDIA stack - the real moat isn't the chip
Competitors can match a chip; matching ~20 years of CUDA-X libraries and every framework optimizing for them is far harder. Trace a workload to see how far down its dependencies run - or tap any layer or chip to see what it does.
3
CUDA-X libraries on this path
5
layers the call crosses
1
platform under all of it: CUDA
Call path
Train an LLM
NeMo builds on PyTorch, which calls cuDNN and cuBLAS for the math and NCCL to sync gradients across GPUs.
Why it's sticky: every lit box is compiled for, and tuned on, CUDA. Moving a workload to other silicon means an equally fast equivalent for each of those libraries - multiplied across every workload a company runs. The chip is replaceable; the call paths are much harder to replace.
CUDA-X spans 400+ libraries; shown here is a representative slice, and the call paths are representative, not exhaustive. Names are NVIDIA products/libraries plus the open frameworks that run on them.
The serving and training layers (NIM, NeMo, Dynamo, TensorRT-LLM) get their own treatment on the inference & serving and training & fine-tuning pages.