Memory & Efficiency

Model Sizing & Parallelism

Intermediate

A model's parameter count is only the start. This page breaks down where GPU memory actually goes, then shows how the six parallelism strategies split a model across hundreds or thousands of GPUs - with the math, the interconnect limits, and the real recipes frontier labs use.

Where the memory goes

Inference pays for weights and the KV cache. Training pays for far more: every weight needs a and an , and Adam's states alone are ~6× the weights. A 70B model that serves in 140 GB needs over 1.1 TB of state to train - which is the whole reason parallelism exists. Training & Fine-Tuning breaks this bill down interactively; here the point is what gets sharded and why.

The bill for a 70B model - training vs serving

Same model, same BF16 weights. Tap a segment to see where each term comes from.

Training

> 1.1 TB

Inference

140 GB + KV cache

KV
no gradients, no optimizer state
dashed grid = 80 GB, one H100's HBMhatched = grows with batch, context or users (not to scale)

Six ways to split a model

When a model - or its training state - is too big for one GPU, you split it across a node of 8 GPUs (and beyond). Each strategy cuts along a different axis and pays a different communication cost. Pick one to see exactly what lands on each of the 8 GPUs, how they talk to each other, and when you'd reach for it.

Parallelism explorer - cut the work, then watch the GPUs talk

Each strategy cuts a different object across the same 8 GPUs. Pick one: the cut lines show what gets split, each GPU's color shows which piece it owns, and the dots show the collective it pays for.

The work - and where it's cut

Model · 80 layers × widthBatchL1L80same on all 8Sequence · 128K tokens

8 GPUs - what each one holds

Ring all-reduce (activations) · twice per layer, every forward & backward
GPU 0width 1/8GPU 1width 2/8GPU 2width 3/8GPU 3width 4/8GPU 4width 5/8GPU 5width 6/8GPU 6width 7/8GPU 7width 8/8
1/14Reduce-scatter partial activations - pass 1 of 7
color = the GPU that owns the piecegrey = full copy on every GPUcut linedot = data in flight

TP - Vertical cuts through every layer: each GPU owns 1/8 of every weight matrix and sees the same batch.

What gets split
Individual weight matrices. Each layer's matmul is sliced column/row-wise across GPUs and recombined.
Collective
Ring all-reduce (activations)
How often
Twice per layer, every forward & backward
Interconnect
Bandwidth-hungry - must stay inside one NVLink domain. Degree is usually capped at 8 (one node).
When to use TP
When a single layer is too big for one GPU, or to cut latency. The innermost axis.

In the wild: Megatron-LM standard; Llama 3 405B used TP=8 within each node. Layer, expert and token counts in the picture are illustrative (an 80-layer model, 256 experts, a 128K sequence).

Make it concrete - split real model weights

Before any clever strategy, parallelism's first job is simply holding the weights. Pick a real model and spread it across 8 GPUs (or 1, 2, 4) to see how much each device carries - and how many GPUs you need just to fit the parameters, before the KV cache and activations are added.

Weights-split playground - fit a real model across GPUs

Pick a model, a GPU count, and an HBM size. Tensor / pipeline / FSDP parallelism all do the same first job: divide the weights so each GPU holds only its share.

Model

GPUs

GPU model

Llama 3.3-70B is 140 GB of weights (BF16 · 70B params). Split 8 ways, each H200 holds 17.5 GB - 12% of its 141 GB, leaving 124 GB per GPU (988 GB in total) for the KV cache and activations.

8 × H200 · each bar is one GPU's 141 GB of HBM

124 GB
free

GPU 0

124 GB
free

GPU 1

124 GB
free

GPU 2

124 GB
free

GPU 3

124 GB
free

GPU 4

124 GB
free

GPU 5

124 GB
free

GPU 6

124 GB
free

GPU 7

weightsfree for KV cache + activations

It fits. Splitting 140 GB across 8 GPUs leaves 17.5 GB per device - under the 141 GB budget, with headroom for activations and the KV cache. Spreading weights this way is exactly what tensor, pipeline, and FSDP parallelism do.

The bandwidth wall: why TP stays in one node

is the most communication-hungry strategy: it all-reduces activations twice per layer, on every forward and backward pass. That only works when GPUs share an domain. The moment you cross to or Ethernet - 18–144× slower - the GPUs stall waiting on each other. Watch the matmul split, then see the bandwidth cliff.

Split the matmul - a tensor-parallel MLP in three frames

Megatron's trick: cut W1 by columns and W2 by rows, so the only communication in the whole MLP block is one all-reduce at the end.

TP degreeW1 slice/GPU: 8,192 × 7,168 = 59M params
1/3Column-parallel: X · W1

Why TP stays inside one node - the same all-reduce over each link

NVLink 5 (Blackwell)

1800 GB/s · intra-node

0% · needs 1×

NVLink 4 (Hopper)

900 GB/s · intra-node

0% · needs 2×

InfiniBand NDR

50 GB/s · inter-node

0% · needs 36×

100G Ethernet

12.5 GB/s · inter-node

0% · needs 144×

clock: 0.0× the NVLink 5 transfer time

Pipe thickness grows with bandwidth (log scale). Times are ratios of the bandwidths only - latency and overlap are ignored.

TP's all-reduce runs twice per layer - for an 80-layer model that's 160 collectives per forward pass. At ~18-36× less bandwidth than NVLink, inter-node links would stall the GPUs waiting on each other. That's why TP degree is almost always capped at 8 (one NVLink node), while pipeline and data parallelism - which communicate far less - span nodes.

Real recipes combine them

Frontier runs never use one strategy - they nest several into a 3D (or 4D) mesh: TP innermost in the node, PP across nodes, DP as the outer loop, plus EP or CP when the model calls for it. Here's what that looks like in practice.

Hybrid parallelism - how the axes nest

Frontier runs nest several strategies into one mesh, from the GPUs inside a node out to replicas across the cluster. Pick a real recipe, then press play to build it outward one axis at a time.

×
×
×
=
16,384 GPUsfull job · 16,384 H100
DP × 8 replicasPP 16 · pipelineS1… 16S2… 16S3… 16⋮ 16 stages
5/5× DP 8 → 16,384 GPUs (replicas · outer loop)

Llama 3.1 405B · 128K long-context stage · 16,384 H100

To reach 128K context the same 16,384 GPUs are regrouped: CP=16 splits each sequence across 16 nodes, so DP drops from 128 to 8. The axes nest from the GPU outward: TP, CP, PP, DP.

  • one node: 8 GPUs in an NVLink domain - TP 8 lives here
  • context group: 16 nodes, each holding 1/16 of the 128K-token sequence (the stripe)
  • pipeline: 16 stages, each a slice of the layers; arrows = activations handed on
  • data parallel: 8 replicas of everything inside, each on its own data
  • every link outside a green box crosses InfiniBand between nodes

Diagrams show a representative slice - the first few nodes and stages - not every GPU. Degrees are from published tech reports (Llama 3 paper, DeepSeek-V3 report); both Llama stages multiply to 16,384 GPUs.

Four recipes on one fleet ruler

Each job's parallelism factors, and the fleet they fill. The ruler is logarithmic - every gridline is 8× more GPUs.

8K stageTP=8×PP=16×DP=128= 16,384
128K stageTP=8×CP=16×PP=16×DP=8= 16,384

TP inside each NVLink node, PP across nodes, DP as the outer loop. For the 128K long-context stage, CP=16 splits each sequence and DP drops to 8 - both layouts multiply to 16,384 GPUs.

PP=16·EP=64·ZeRO-1·DualPipe
FSDP / ZeRO-3
DP only

GPUs (log scale)

Takeaway: small jobs get by with one strategy; only frontier runs stack three or four, and their factors multiply out to the whole fleet. Tap a recipe for why it is built that way.

Fleet sizes and recipes are from published tech reports and papers; GPT-4's exact layout is unconfirmed. Numbers illustrate the scaling pattern, not exact reproductions.

Which ones you actually use

Not every strategy shows up in every run. Two are nearly universal, two are standard once you outgrow a single GPU, and two are reserved for specific model shapes.

How often each strategy shows up

From the default in every run to the ones a workload has to ask for. Tap a tier.

Data Parallelism · ZeRO / FSDP

Cheap to communicate (DP) or essential for memory (ZeRO). The default outer loop of nearly every training run.

Memory is the bottleneck - everywhere

Parallelism splits the model across GPUs; quantization shrinks what each GPU has to hold; the KV cache is the memory that grows with every token served. They're three views of the same constraint - HBM is scarce and bandwidth is finite.