Real World Deployment
Running Models on GB300 & B300
AdvancedSpec sheets say what a GPU can do; deployments show what teams actually do with it. This page follows real 2026 recipes and benchmarks on NVIDIA Blackwell Ultra - which models fit where, how the big ones are served, where each system pays off, and what trips people up.
Two shapes of Blackwell Ultra
Blackwell Ultra ships in two very different shapes. puts 8 GPUs in one air-cooled server, all joined by 5 at 1.8 TB/s per GPU, with x86 host CPUs attached over PCIe Gen6. joins 72 of the same GPUs and 36 Grace CPUs into a single liquid-cooled rack with one 130 TB/s NVLink domain. Those differences decide how large a model one NVLink domain can hold, how widely experts can be spread, and what a data center has to provide in power and cooling.
Two shapes of Blackwell Ultra - an 8-GPU HGX B300 box vs a 72-GPU GB300 NVL72 rack
Same GPU generation, two very different systems. Tap a fact to light up the parts it is about and see why it matters.
HGX B300
8-GPU box, air-cooled
2,304 GB
HBM total, spec sheet
GB300 NVL72
72-GPU rack, liquid-cooled
20.7 TB
HBM total, spec sheet
Scale-up vs scale-out, per GPU (both directions)
NVLink carries ~9x the bandwidth of a network port. Expert parallelism sends all-to-all traffic between GPUs at every MoE layer, so spreading experts beyond 8 GPUs is cheap only inside an NVL72 rack - on HGX B300 that traffic has to leave the box over the network.
Does it fit? Weights, KV and users
A Blackwell Ultra GPU carries 288 GB of , but only about 268 GB on HGX B300 and about 278 GB on GB300 is usable, and the serving runtime keeps part of that for activations and buffers. Whatever the weights leave behind becomes , and the KV cache decides how many users fit at a given context length. Models under ~150 GB run one replica per GPU, 400B-1T MoE models need 2-4 GPUs just to hold their weights, and 1.6-2.8T models all but fill an 8-GPU box - which is why they move to larger NVLink domains.
Does it fit? Weights, KV and users on Blackwell Ultra and Vera Rubin
Pick a model, precision and system. The planner puts the weights on the fewest GPUs that still leave room for KV (one replica), fills the rest of each GPU's usable HBM with KV cache, and counts how many users fit at your context length.
69 linear-attention layers keep a fixed state; only 24 MLA layers grow KV. MXFP4 experts with BF16 attention average ~0.56 B/param.
288 GB HBM3e per GPU, ~268 GB usable on HGX B300 · 8-GPU NVLink domain
16 users use 8% of the KV headroom; up to 199 fit at 128K.
- Weights 196 GB
- KV in use 3.6 GB
- Free KV room 41.6 GB
- Runtime reserve (10%) 26.8 GB
1 replica of 8 GPUs. Users are spread evenly across replicas.
- < 150 GB of weights1 GPU per replica - many replicas side by side
- 400B - 1T MoE2-4 GPUs hold the weights; an 8-GPU box adds room for KV
- 1.6T - 2.8T MoEan 8-GPU box barely holds the weights - NVL72 or multi-node for users
Assumptions: usable HBM ~268 GB/GPU (HGX B300) and ~278 GB/GPU (GB300); NVFP4 ~0.5625 and MXFP4 ~0.53 bytes/param including scales, FP8 1. KV is spread evenly across a replica's GPUs (head-sharded or context-parallel); a latent cache copied to every GPU holds fewer users. Linear-attention state, prefix caching and KV offload are not modeled. The runtime reserve is an illustrative default.
KV bytes per token come from the KV Cache planner's model list; the attention design behind each number is on Attention Mechanisms.
The serving playbook on Blackwell Ultra
Across MLPerf submissions, engine recipes and published deployments, the same six techniques show up again and again. None is specific to one vendor or engine - they are how the hardware gets used well.
4-bit weights, 8-bit KV
NVFP4 (or native MXFP4) weights with an FP8 KV cache is the default in MLPerf submissions and engine recipes. Blackwell Ultra's dense FP4 rate is 1.5× B200's, while its FP8 rate did not rise - and INT8 was cut sharply, so W8A8 INT8 is off the table.
Split prefill from decode
Large deployments run prompt processing and token generation on separate GPU groups with their own parallelism, handing the KV cache between them over NIXL or Mooncake - orchestrated by NVIDIA Dynamo, SGLang or vLLM.
Wide expert parallelism
Decode for big MoE models spreads experts across 16-72 GPUs inside one NVLink domain, with data-parallel attention and expert load balancing (EPLB) to even out hot experts.
Multi-token prediction
Speculative decoding with MTP heads, EAGLE3 or DSpark raises per-user speed 1.9-3× at the same per-GPU throughput, which moves a deployment along the interactivity curve almost for free.
KV tiering and routing
KV spills from HBM to host memory, local NVMe and shared storage (Dynamo KVBM, SGLang HiCache, LMCache), and KV-aware routers send each request to the replica that already holds its prefix.
Kernels and graphs
FlashInfer and TRT-LLM MoE kernels, DeepGEMM, FlashMLA and full CUDA graphs on decode. GB300 runs fused attention 1.35× faster than GB200 thanks to its doubled exponent units.
Why big MoE models want 72 GPUs
DeepSeek V3 and R1 spread each layer's 256 routed experts across GPUs, and every token visits only a few of them. sets how many GPUs share those experts: the wider the group, the fewer experts each GPU holds, which frees HBM and cuts the weight bytes each decode step reads. The cost is the all-to-all exchange that sends tokens to their experts and back. On 8-GPU HGX B300 boxes that traffic leaves NVLink as soon as EP goes past 8, while a GB300 NVL72 keeps all 72 GPUs in one NVLink domain.
Wide expert parallelism - why big MoE models want a 72-GPU NVLink domain
DeepSeek V3 and R1 (671B parameters, 37B active) have 256 routed experts in every MoE layer. Spread them over more GPUs and each GPU holds fewer - but every token now has to travel to wherever its experts live.
Dispatch: every copy of the token rides NVLink to the GPU holding its expert.
Experts per GPU
8
256 / 32 per layer
Expert weights / GPU
~12 GB
of ~370 GB in NVFP4
HBM left for KV
~241 GB
of ~278 GB usable per GPU
All-to-all off NVLink
0%
all 72 GPUs share NVLink
One GPU's HBM (illustrative)
Reading its experts once at 8 TB/s: ~1.4 ms per decode step
One 72-GPU NVLink domain (130 TB/s). Every EP size up to EP72 keeps the all-to-all on NVLink, so the rack can trade experts per GPU for KV cache and bigger decode batches without touching the scale-out network. At EP32 the other 40 GPUs run more replicas.
1.8x
output tokens/s per GPU for EP32 vs EP8 on an NVL72 rack at 100 tok/s per user, in NVIDIA benchmarks (Oct 2025).
4,021 → 12,587
decode tok/s per GPU for Kimi K2.5 on GB200 NVL72 with wide EP16, in InferenceX benchmarks.
Load balancing matters more as EP widens. Real routing is uneven - some experts are hot - and the all-to-all waits for the busiest GPU. Expert-parallel load balancing (EPLB) places extra copies of hot experts and reshuffles placement so no GPU becomes the straggler. At EP72, 256 experts don't divide evenly (3 or 4 per GPU), which leaves room for those copies.
Illustrative: ~370 GB of routed-expert weights (671B x ~0.56 bytes in NVFP4), ~25 GB per GPU kept for attention and shared weights plus activations, one token routed to 8 experts, evenly spread routing. Scale-out packets are drawn 3x slower than NVLink; the per-GPU bandwidth gap is ~9x.
One rack, two jobs: disaggregated serving
Prefill and decode stress a GPU in opposite ways, so large deployments give each its own GPUs and its own parallelism. The InferenceX GB300 recipe for DeepSeek V4 Pro fills exactly one NVL72 rack: ten 4-GPU prefill workers and one 32-GPU decode worker, with the KV cache handed over NIXL inside the rack. Step through one request, then see how software alone raised this setup's throughput about 5× between April and June 2026.
One rack, two jobs - disaggregated serving of DeepSeek V4 Pro on GB300 NVL72
The 72 GPUs below are the rack's 18 compute trays of 4. The recipe disagg-gb300-10p1d-dep4-dep32 splits them into 10 prefill workers and 1 decode worker. Follow one request through.
Request
prompt tokens
Prefill worker P3
4 GPUs · DEP4 · compute-bound
Decode worker
32 GPUs · DEP32 · bandwidth-bound
Tokens
streamed to the user
Where the 72 GPUs go
10 prefill workers x 4 GPUs = 40 + 1 decode worker x 32 GPUs = 32 = 72 GPUs, one NVL72 rack
Prefill is compute-bound
A prompt's tokens are processed all at once, so prefill is a big batch of matrix math. Ten small DEP4 workers each take their own prompts in parallel, which keeps time to first token low as requests pile up.
Decode is bandwidth-bound
Each new token re-reads weights and KV cache for a tiny amount of math. One wide DEP32 worker spreads the experts thin and pools many streams into big batches, so each byte read from HBM serves more tokens.
Same rack, better software
5x
Output tok/s per GPU for DeepSeek V4 Pro with SGLang disaggregated serving on GB300 at ~50 tok/s per user, in benchmarks - about 5x in two months from software alone.
For how prefill/decode disaggregation works on any cluster - and the KV transfer it costs - see Inference serving.
Tray placement is illustrative (one prefill worker per tray); the worker counts and GPU split are the recipe's. DEP = data-parallel attention with expert parallelism across the worker's GPUs.
The prefill/decode split itself is explained on Inference & Serving.
Where NVL72 pays off
Every serving deployment picks a point on one curve: tokens per second per GPU, which sets cost, against tokens per second per user, which sets how fast each reply streams. Batching more users onto a GPU raises throughput but slows each user down. GB300 NVL72's 72-GPU NVLink domain helps most in the middle of that curve, roughly 30-160 tokens/s per user, where large MoE models need expert parallelism across more than 8 GPUs; at the low and high extremes it lands near parity with an 8-GPU HGX B300. The dots below are DeepSeek R1 results in benchmarks.
Where NVL72 pays off - throughput vs interactivity
Pick a matchup and a metric. Dots are DeepSeek R1 benchmark results; the dashed lines only sketch the usual shape of the curve between them.
Only per-dollar results are published for this matchup.
GB300 NVL72 vs HGX B300
On a 1K-in / 1K-out workload, the 72-GPU NVL72 rack is per dollar about 2x cheaper than an 8-GPU HGX B300 box at 88 tok/s per user and about 3.5x cheaper at 162, where a big MoE model's expert all-to-all traffic rides NVLink across the whole rack. By 235 tok/s per user the two cost about the same.
Long context: 128K in / 8K out
On DeepSeek R1, GB300 NVL72 does 226 vs 148 tokens/s per GPU for GB200 (1.53x), with a max decode batch of 40 vs 24 per GPU. No interactivity level is given, so it isn't plotted (LMSYS, Feb 2026).
- HGX B300, 88 tok/s/user: $0.129 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
- HGX B300, 162 tok/s/user: $1.159 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
- GB300 NVL72, 88 tok/s/user: $0.064 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
- GB300 NVL72, 162 tok/s/user: $0.326 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
- 235 tok/s/user: cost about equal - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing; no exact value, so the ring's height is illustrative
Benchmark caveats: InferenceX runs use random data with prefix caching off; disaggregated and aggregated per-GPU numbers aren't directly comparable; MLPerf latency bounds are loose, which hides much of the rack-scale advantage. InferenceX counts input and output tokens together, so an 8K-in / 1K-out test (about 89% input) looks 2-5x cheaper per token than a 1K / 1K test at the same speed. Each matchup uses its own workload, so don't compare values across matchups.
Picking a system
GB300 NVL72
Choose it for
Large MoE models (roughly 600B and up), long contexts, and latency targets in the middle of the curve.
Why
Expert all-to-all stays on the 130 TB/s NVLink domain, so decode can spread experts across 16-72 GPUs; fewer experts per GPU leaves more HBM for KV and bigger batches; a whole prefill + decode layout fits in one rack.
Tradeoff
Up to 142 kW liquid-cooled racks, and 72 GPUs share one NVLink failure domain.
HGX B300
Choose it for
Models whose weights and KV fit in ~2.1 TB: gpt-oss, 70B dense, Qwen3.5-397B, GLM-5.x, Kimi K2.x, MiniMax.
Why
Air-cooled 14.5 kW servers that fit existing data centers, an 8-GPU failure domain, similar per-GPU pricing to GB300 and broad cloud availability. In benchmarks it matches or beats GB300 on cost at low interactivity for some models (MiniMax-M3).
Tradeoff
Scale-up stops at 8 GPUs, so expert parallelism beyond 8 crosses a network about 9× slower than NVLink.
GB200 NVL72 (previous generation)
Choose it for
Workloads that don't need Blackwell Ultra's extra memory or attention speed.
Why
GB300 wins clearly only where 1.5× the HBM or 2× attention unlocks a configuration GB200 can't run - DeepSeek V4 (2.8× in benchmarks) or 128K contexts (1.5×). Elsewhere GB200 is often cheaper per token.
Tradeoff
Less HBM per GPU (about 186 GB) caps batch size and context sooner.
Operators increasingly measure all of this as : power, not GPU count, is usually the binding constraint on how much a site can serve.
What changes with Vera Rubin
Vera Rubin NVL72 began shipping in limited volume in September 2026. It keeps the GB300 NVL72 shape - 72 GPUs in one NVLink domain, 288 GB of HBM each - so the planner above gives the same answer for both: the same models fit with the same number of users. What changes is speed. HBM4 runs at 19.2 TB/s per GPU against 8 TB/s on Blackwell Ultra, NVLink 6 reaches 3 TB/s per GPU, and dense FP4 rises to 35 PF per GPU. The same deployment produces more tokens per second, at roughly 190-230 kW per rack.
One NVLink domain, five generations
Log scale. Pick a measure; the dashed rung is planned, not shipping.
HGX H100
8 GPUs · 2023
GB200 NVL72
72 GPUs · 2025
GB300 NVL72
72 GPUs · 2025-26
Vera Rubin NVL72
72 GPUs · shipping 2H 2026
Rubin Ultra Kyber NVL144
144 GPUs · planned 2027
GB300 to Vera Rubin: the same 20.7 TB - every model that fits one rack fits the other.
Dense (not sparse) FP4. Vera Rubin: 72 GPUs at 35 PF dense FP4 each. Rubin Ultra follows NVIDIA's 2025 roadmap; mid-2026 reports describe a redesign. Server and rack power are maximum or rated figures.
Vera Rubin NVL72 (formerly NVL144)
NVIDIA first announced the Vera Rubin rack as Vera Rubin NVL144, counting the two dies inside each Rubin GPU. It now counts whole GPUs, so the same rack is called Vera Rubin NVL72: 72 Rubin GPUs and 36 Vera CPUs, the same footprint as GB300 NVL72.
Much more host memory
Vera CPUs carry up to 54 TB of LPDDR5X per rack, about 2.6x the rack's HBM (GB300's Grace memory is a bit less than its HBM). Offloading KV cache to host memory inside the rack becomes far more useful.
Rubin CPX is off the roadmap
NVIDIA announced Rubin CPX in 2025 as a GDDR7 chip for long-context prefill, then dropped it at GTC 2026. The 2026 split pairs Rubin GPUs for prefill with Groq 3 LPX racks for decode, under NVIDIA Dynamo.
Rubin Ultra is the next step up
Planned for 2027: Kyber NVL144, 144 four-die GPUs with up to 1 TB of HBM4E each in one 800 VDC rack, and NVL576 linking eight racks. Mid-2026 reports described a two-die redesign and a slip to 2028; NVIDIA says the roadmap is intact.
Software moves the frontier as much as silicon
The same GB300 hardware posted very different numbers a few months apart. Between MLPerf Inference v5.1 (Sep 2025) and v6.0 (Apr 2026), DeepSeek R1 throughput per GPU rose 1.7× offline and 2.8× in the server scenario, and SGLang lifted DeepSeek V4 Pro about 5× between day 0 and June 2026. Turning on alone cut DeepSeek R1's about 21× at 150 tokens/s per user.
Software moves the frontier as much as silicon
Every jump below is on the same GB300 hardware - only the software changed. Tap a result to see where it sits in time.
Highlighted: the dates this result compares.
Compare hardware on the same software date. The same GB300 rack got 1.7-5x faster in a few months, so an older chip on newer software can outscore a faster chip measured on older software.
Pitfalls practitioners hit
288 GB isn't what you get
Usable HBM is about 268-275 GB per GPU on HGX B300 and about 278 GB on GB300. Size against the usable number, then reserve room for activations and buffers.
INT8 and FP64 were cut
Blackwell Ultra traded INT8 and FP64 throughput for FP4. Use FP8, NVFP4 or MXFP4; W4A16 INT4 still works.
HGX B300 is not a smaller GB300
Dense FP4 is 13.5 vs 15 PF per GPU and scale-up stops at 8 GPUs - tuned DeepSeek V4 on B300 ran better at EP4 than EP8.
Day-0 software is rough
New models hit engine bugs in their first weeks - hardcoded hidden sizes, FP8 scale errors producing NaNs, sliding-window memory bugs under disaggregation. Expect weeks of tuning.
Read benchmarks carefully
InferenceX runs random data with prefix caching off, disaggregated and aggregated per-GPU numbers aren't directly comparable, and MLPerf's loose latency bounds hide the rack-scale advantage.
A KV tier must be big to help
Host memory or flash only lifts hit rates when it is roughly 1.5-3× the HBM KV capacity; in one benchmark 3 TB of DRAM on a B300 node added just 1.36% hit rate. GB300's Grace memory is about 0.85× its HBM.
Big models load slowly
Kimi K3 took about 81 minutes to load on 8× B300 - which slows tuning, scaling and failover.
Hybrid models need a memory split
Models with linear-attention or state-space layers keep a fixed state pool beside the KV cache; the split between them must be tuned per workload.
Figures on this page come from vendor spec pages, engine recipes, MLPerf Inference results and InferenceX / LMSYS benchmarks from late 2025 to September 2026; benchmark results depend on workload, software version and latency target.