Memory & Efficiency
Model Sizing & Parallelism
IntermediateA model's parameter count is only the start. This page breaks down where GPU memory actually goes, then shows how the six parallelism strategies split a model across hundreds or thousands of GPUs - with the math, the interconnect limits, and the real recipes frontier labs use.
Where the memory goes
Inference pays for weights and the KV cache. Training pays for far more: every weight needs a and an , and Adam's states alone are ~6× the weights. A 70B model that serves in 140 GB needs over 1.1 TB of state to train - which is the whole reason parallelism exists. Training & Fine-Tuning breaks this bill down interactively; here the point is what gets sharded and why.
The bill for a 70B model - training vs serving
Same model, same BF16 weights. Tap a segment to see where each term comes from.
Training
> 1.1 TB
Inference
140 GB + KV cache
Six ways to split a model
When a model - or its training state - is too big for one GPU, you split it across a node of 8 GPUs (and beyond). Each strategy cuts along a different axis and pays a different communication cost. Pick one to see exactly what lands on each of the 8 GPUs, how they talk to each other, and when you'd reach for it.
Parallelism explorer - cut the work, then watch the GPUs talk
Each strategy cuts a different object across the same 8 GPUs. Pick one: the cut lines show what gets split, each GPU's color shows which piece it owns, and the dots show the collective it pays for.
The work - and where it's cut
8 GPUs - what each one holds
Ring all-reduce (activations) · twice per layer, every forward & backwardTP - Vertical cuts through every layer: each GPU owns 1/8 of every weight matrix and sees the same batch.
- What gets split
- Individual weight matrices. Each layer's matmul is sliced column/row-wise across GPUs and recombined.
- Collective
- Ring all-reduce (activations)
- How often
- Twice per layer, every forward & backward
- Interconnect
- Bandwidth-hungry - must stay inside one NVLink domain. Degree is usually capped at 8 (one node).
- When to use TP
- When a single layer is too big for one GPU, or to cut latency. The innermost axis.
In the wild: Megatron-LM standard; Llama 3 405B used TP=8 within each node. Layer, expert and token counts in the picture are illustrative (an 80-layer model, 256 experts, a 128K sequence).
Make it concrete - split real model weights
Before any clever strategy, parallelism's first job is simply holding the weights. Pick a real model and spread it across 8 GPUs (or 1, 2, 4) to see how much each device carries - and how many GPUs you need just to fit the parameters, before the KV cache and activations are added.
Weights-split playground - fit a real model across GPUs
Pick a model, a GPU count, and an HBM size. Tensor / pipeline / FSDP parallelism all do the same first job: divide the weights so each GPU holds only its share.
Model
GPUs
GPU model
Llama 3.3-70B is 140 GB of weights (BF16 · 70B params). Split 8 ways, each H200 holds 17.5 GB - 12% of its 141 GB, leaving 124 GB per GPU (988 GB in total) for the KV cache and activations.
8 × H200 · each bar is one GPU's 141 GB of HBM
free
GPU 0
free
GPU 1
free
GPU 2
free
GPU 3
free
GPU 4
free
GPU 5
free
GPU 6
free
GPU 7
It fits. Splitting 140 GB across 8 GPUs leaves 17.5 GB per device - under the 141 GB budget, with headroom for activations and the KV cache. Spreading weights this way is exactly what tensor, pipeline, and FSDP parallelism do.
The bandwidth wall: why TP stays in one node
is the most communication-hungry strategy: it all-reduces activations twice per layer, on every forward and backward pass. That only works when GPUs share an domain. The moment you cross to or Ethernet - 18–144× slower - the GPUs stall waiting on each other. Watch the matmul split, then see the bandwidth cliff.
Split the matmul - a tensor-parallel MLP in three frames
Megatron's trick: cut W1 by columns and W2 by rows, so the only communication in the whole MLP block is one all-reduce at the end.
Why TP stays inside one node - the same all-reduce over each link
NVLink 5 (Blackwell)
1800 GB/s · intra-node
0% · needs 1×
NVLink 4 (Hopper)
900 GB/s · intra-node
0% · needs 2×
InfiniBand NDR
50 GB/s · inter-node
0% · needs 36×
100G Ethernet
12.5 GB/s · inter-node
0% · needs 144×
clock: 0.0× the NVLink 5 transfer time
Pipe thickness grows with bandwidth (log scale). Times are ratios of the bandwidths only - latency and overlap are ignored.
TP's all-reduce runs twice per layer - for an 80-layer model that's 160 collectives per forward pass. At ~18-36× less bandwidth than NVLink, inter-node links would stall the GPUs waiting on each other. That's why TP degree is almost always capped at 8 (one NVLink node), while pipeline and data parallelism - which communicate far less - span nodes.
Real recipes combine them
Frontier runs never use one strategy - they nest several into a 3D (or 4D) mesh: TP innermost in the node, PP across nodes, DP as the outer loop, plus EP or CP when the model calls for it. Here's what that looks like in practice.
Hybrid parallelism - how the axes nest
Frontier runs nest several strategies into one mesh, from the GPUs inside a node out to replicas across the cluster. Pick a real recipe, then press play to build it outward one axis at a time.
Llama 3.1 405B · 128K long-context stage · 16,384 H100
To reach 128K context the same 16,384 GPUs are regrouped: CP=16 splits each sequence across 16 nodes, so DP drops from 128 to 8. The axes nest from the GPU outward: TP, CP, PP, DP.
- one node: 8 GPUs in an NVLink domain - TP 8 lives here
- context group: 16 nodes, each holding 1/16 of the 128K-token sequence (the stripe)
- pipeline: 16 stages, each a slice of the layers; arrows = activations handed on
- data parallel: 8 replicas of everything inside, each on its own data
- every link outside a green box crosses InfiniBand between nodes
Diagrams show a representative slice - the first few nodes and stages - not every GPU. Degrees are from published tech reports (Llama 3 paper, DeepSeek-V3 report); both Llama stages multiply to 16,384 GPUs.
Four recipes on one fleet ruler
Each job's parallelism factors, and the fleet they fill. The ruler is logarithmic - every gridline is 8× more GPUs.
TP inside each NVLink node, PP across nodes, DP as the outer loop. For the 128K long-context stage, CP=16 splits each sequence and DP drops to 8 - both layouts multiply to 16,384 GPUs.
GPUs (log scale)
Takeaway: small jobs get by with one strategy; only frontier runs stack three or four, and their factors multiply out to the whole fleet. Tap a recipe for why it is built that way.
Fleet sizes and recipes are from published tech reports and papers; GPT-4's exact layout is unconfirmed. Numbers illustrate the scaling pattern, not exact reproductions.
Which ones you actually use
Not every strategy shows up in every run. Two are nearly universal, two are standard once you outgrow a single GPU, and two are reserved for specific model shapes.
How often each strategy shows up
From the default in every run to the ones a workload has to ask for. Tap a tier.
Data Parallelism · ZeRO / FSDP
Cheap to communicate (DP) or essential for memory (ZeRO). The default outer loop of nearly every training run.
Memory is the bottleneck - everywhere
Parallelism splits the model across GPUs; quantization shrinks what each GPU has to hold; the KV cache is the memory that grows with every token served. They're three views of the same constraint - HBM is scarce and bandwidth is finite.