Real World Deployment

Running Models on GB300 & B300

Advanced

Spec sheets say what a GPU can do; deployments show what teams actually do with it. This page follows real 2026 recipes and benchmarks on NVIDIA Blackwell Ultra - which models fit where, how the big ones are served, where each system pays off, and what trips people up.

Two shapes of Blackwell Ultra

Blackwell Ultra ships in two very different shapes. puts 8 GPUs in one air-cooled server, all joined by 5 at 1.8 TB/s per GPU, with x86 host CPUs attached over PCIe Gen6. joins 72 of the same GPUs and 36 Grace CPUs into a single liquid-cooled rack with one 130 TB/s NVLink domain. Those differences decide how large a model one NVLink domain can hold, how widely experts can be spread, and what a data center has to provide in power and cooling.

Two shapes of Blackwell Ultra - an 8-GPU HGX B300 box vs a 72-GPU GB300 NVL72 rack

Same GPU generation, two very different systems. Tap a fact to light up the parts it is about and see why it matters.

HBM per GPU

HGX B300

8-GPU box, air-cooled

2,304 GB

HBM total, spec sheet

HGX B300airNVSwitch · NVLink 5GPUGPUGPUGPUGPUGPUGPUGPU8x ConnectX-8 800GXeonXeonDDR5
GPUNVSwitch / HBMNICXeon + DDR5

GB300 NVL72

72-GPU rack, liquid-cooled

20.7 TB

HBM total, spec sheet

compute tray4 GPU + 2 Grace9 NVLinkswitch traysliquid-cooledmanifold (left)
GPUGrace CPUNVLink switch trayLiquid manifold
HGX B300GB300 NVL72

Scale-up vs scale-out, per GPU (both directions)

NVLink 51.8 TB/s
ConnectX-8 port (800 Gb/s)~0.2 TB/s

NVLink carries ~9x the bandwidth of a network port. Expert parallelism sends all-to-all traffic between GPUs at every MoE layer, so spreading experts beyond 8 GPUs is cheap only inside an NVL72 rack - on HGX B300 that traffic has to leave the box over the network.

Does it fit? Weights, KV and users

A Blackwell Ultra GPU carries 288 GB of , but only about 268 GB on HGX B300 and about 278 GB on GB300 is usable, and the serving runtime keeps part of that for activations and buffers. Whatever the weights leave behind becomes , and the KV cache decides how many users fit at a given context length. Models under ~150 GB run one replica per GPU, 400B-1T MoE models need 2-4 GPUs just to hold their weights, and 1.6-2.8T models all but fill an 8-GPU box - which is why they move to larger NVLink domains.

Does it fit? Weights, KV and users on Blackwell Ultra and Vera Rubin

Pick a model, precision and system. The planner puts the weights on the fewest GPUs that still leave room for KV (one replica), fills the rest of each GPU's usable HBM with KV cache, and counts how many users fit at your context length.

Presets

69 linear-attention layers keep a fixed state; only 24 MLA layers grow KV. MXFP4 experts with BF16 attention average ~0.56 B/param.

288 GB HBM3e per GPU, ~268 GB usable on HGX B300 · 8-GPU NVLink domain

Weights
KV cache14 KB / token
Context per usermodel max 1M128K
Runtime reserveactivations, graphs, buffers10% · 27 GB
GPUs per replica
Fits comfortably

16 users use 8% of the KV headroom; up to 199 fit at 128K.

Weights / GPU
196 GB
1.57 TB over 8 GPUs
KV room / GPU
45.2 GB
268 usable - 27 reserve - weights
Max users @ 128K
199
1.8 GB KV per user
Replicas
1 x 8
199 users each
One GPU's usable HBM, to scale268 GB usable of 288 GB
  • Weights 196 GB
  • KV in use 3.6 GB
  • Free KV room 41.6 GB
  • Runtime reserve (10%) 26.8 GB
8x B300: 8 GPUs, each bar to scaleunderline color = replica
0
1
2
3
4
5
6
7

1 replica of 8 GPUs. Users are spread evenly across replicas.

The real deployment. Fixstars ran Kimi K3 on 8x B300 in July 2026: 195.9 GB of weights per GPU, ~26 GB free, TP8 with 8-way decode context parallelism, ~528K tokens of context capacity and ~30 concurrent requests (754 tok/s output, ~81-minute model load). The preset sets the runtime reserve to 17% to land on that ~26 GB. In production, NVIDIA Dynamo's Kimi K3 recipe moves to 16-24 GB300 GPUs (aggregated) or 16 GPUs split into prefill and decode, which cuts weights to ~65-98 GB per GPU and leaves the rest for KV.
The general pattern on 288 GB GPUs
  • < 150 GB of weights1 GPU per replica - many replicas side by side
  • 400B - 1T MoE2-4 GPUs hold the weights; an 8-GPU box adds room for KV
  • 1.6T - 2.8T MoEan 8-GPU box barely holds the weights - NVL72 or multi-node for users

Assumptions: usable HBM ~268 GB/GPU (HGX B300) and ~278 GB/GPU (GB300); NVFP4 ~0.5625 and MXFP4 ~0.53 bytes/param including scales, FP8 1. KV is spread evenly across a replica's GPUs (head-sharded or context-parallel); a latent cache copied to every GPU holds fewer users. Linear-attention state, prefix caching and KV offload are not modeled. The runtime reserve is an illustrative default.

KV bytes per token come from the KV Cache planner's model list; the attention design behind each number is on Attention Mechanisms.

The serving playbook on Blackwell Ultra

Across MLPerf submissions, engine recipes and published deployments, the same six techniques show up again and again. None is specific to one vendor or engine - they are how the hardware gets used well.

4-bit weights, 8-bit KV

NVFP4 (or native MXFP4) weights with an FP8 KV cache is the default in MLPerf submissions and engine recipes. Blackwell Ultra's dense FP4 rate is 1.5× B200's, while its FP8 rate did not rise - and INT8 was cut sharply, so W8A8 INT8 is off the table.

Split prefill from decode

Large deployments run prompt processing and token generation on separate GPU groups with their own parallelism, handing the KV cache between them over NIXL or Mooncake - orchestrated by NVIDIA Dynamo, SGLang or vLLM.

Wide expert parallelism

Decode for big MoE models spreads experts across 16-72 GPUs inside one NVLink domain, with data-parallel attention and expert load balancing (EPLB) to even out hot experts.

Multi-token prediction

Speculative decoding with MTP heads, EAGLE3 or DSpark raises per-user speed 1.9-3× at the same per-GPU throughput, which moves a deployment along the interactivity curve almost for free.

KV tiering and routing

KV spills from HBM to host memory, local NVMe and shared storage (Dynamo KVBM, SGLang HiCache, LMCache), and KV-aware routers send each request to the replica that already holds its prefix.

Kernels and graphs

FlashInfer and TRT-LLM MoE kernels, DeepGEMM, FlashMLA and full CUDA graphs on decode. GB300 runs fused attention 1.35× faster than GB200 thanks to its doubled exponent units.

Why big MoE models want 72 GPUs

DeepSeek V3 and R1 spread each layer's 256 routed experts across GPUs, and every token visits only a few of them. sets how many GPUs share those experts: the wider the group, the fewer experts each GPU holds, which frees HBM and cuts the weight bytes each decode step reads. The cost is the all-to-all exchange that sends tokens to their experts and back. On 8-GPU HGX B300 boxes that traffic leaves NVLink as soon as EP goes past 8, while a GB300 NVL72 keeps all 72 GPUs in one NVLink domain.

Wide expert parallelism - why big MoE models want a 72-GPU NVLink domain

DeepSeek V3 and R1 (671B parameters, 37B active) have 256 routed experts in every MoE layer. Spread them over more GPUs and each GPU holds fewer - but every token now has to travel to wherever its experts live.

System
EP size
GB300 NVL72 - one NVLink domain, 72 GPUsEP group: 32 of 72GPU 1: 8 experts per layer (the token starts here)8GPU 2: 8 experts per layer8GPU 3: 8 experts per layer8GPU 4: 8 experts per layer8GPU 5: 8 experts per layer8GPU 6: 8 experts per layer8GPU 7: 8 experts per layer8GPU 8: 8 experts per layer8GPU 9: 8 experts per layer8GPU 10: 8 experts per layer8GPU 11: 8 experts per layer8GPU 12: 8 experts per layer8GPU 13: 8 experts per layer8GPU 14: 8 experts per layer8GPU 15: 8 experts per layer8GPU 16: 8 experts per layer8GPU 17: 8 experts per layer8GPU 18: 8 experts per layer8GPU 19: 8 experts per layer8GPU 20: 8 experts per layer8GPU 21: 8 experts per layer8GPU 22: 8 experts per layer8GPU 23: 8 experts per layer8GPU 24: 8 experts per layer8GPU 25: 8 experts per layer8GPU 26: 8 experts per layer8GPU 27: 8 experts per layer8GPU 28: 8 experts per layer8GPU 29: 8 experts per layer8GPU 30: 8 experts per layer8GPU 31: 8 experts per layer8GPU 32: 8 experts per layer8GPU 33: not in this EP group (serves other replicas)GPU 34: not in this EP group (serves other replicas)GPU 35: not in this EP group (serves other replicas)GPU 36: not in this EP group (serves other replicas)GPU 37: not in this EP group (serves other replicas)GPU 38: not in this EP group (serves other replicas)GPU 39: not in this EP group (serves other replicas)GPU 40: not in this EP group (serves other replicas)GPU 41: not in this EP group (serves other replicas)GPU 42: not in this EP group (serves other replicas)GPU 43: not in this EP group (serves other replicas)GPU 44: not in this EP group (serves other replicas)GPU 45: not in this EP group (serves other replicas)GPU 46: not in this EP group (serves other replicas)GPU 47: not in this EP group (serves other replicas)GPU 48: not in this EP group (serves other replicas)GPU 49: not in this EP group (serves other replicas)GPU 50: not in this EP group (serves other replicas)GPU 51: not in this EP group (serves other replicas)GPU 52: not in this EP group (serves other replicas)GPU 53: not in this EP group (serves other replicas)GPU 54: not in this EP group (serves other replicas)GPU 55: not in this EP group (serves other replicas)GPU 56: not in this EP group (serves other replicas)GPU 57: not in this EP group (serves other replicas)GPU 58: not in this EP group (serves other replicas)GPU 59: not in this EP group (serves other replicas)GPU 60: not in this EP group (serves other replicas)GPU 61: not in this EP group (serves other replicas)GPU 62: not in this EP group (serves other replicas)GPU 63: not in this EP group (serves other replicas)GPU 64: not in this EP group (serves other replicas)GPU 65: not in this EP group (serves other replicas)GPU 66: not in this EP group (serves other replicas)GPU 67: not in this EP group (serves other replicas)GPU 68: not in this EP group (serves other replicas)GPU 69: not in this EP group (serves other replicas)GPU 70: not in this EP group (serves other replicas)GPU 71: not in this EP group (serves other replicas)GPU 72: not in this EP group (serves other replicas)

Dispatch: every copy of the token rides NVLink to the GPU holding its expert.

Token's GPUEP group (number = experts per layer)Holds one of the token's expertsNVLink (1.8 TB/s per GPU)Scale-out network (ConnectX-8, ~9x slower)

Experts per GPU

8

256 / 32 per layer

Expert weights / GPU

~12 GB

of ~370 GB in NVFP4

HBM left for KV

~241 GB

of ~278 GB usable per GPU

All-to-all off NVLink

0%

all 72 GPUs share NVLink

One GPU's HBM (illustrative)

Reading its experts once at 8 TB/s: ~1.4 ms per decode step

Expert weights ~12 GBOther weights + activations ~25 GBKV cache + batch ~241 GB

One 72-GPU NVLink domain (130 TB/s). Every EP size up to EP72 keeps the all-to-all on NVLink, so the rack can trade experts per GPU for KV cache and bigger decode batches without touching the scale-out network. At EP32 the other 40 GPUs run more replicas.

1.8x

output tokens/s per GPU for EP32 vs EP8 on an NVL72 rack at 100 tok/s per user, in NVIDIA benchmarks (Oct 2025).

4,021 → 12,587

decode tok/s per GPU for Kimi K2.5 on GB200 NVL72 with wide EP16, in InferenceX benchmarks.

Load balancing matters more as EP widens. Real routing is uneven - some experts are hot - and the all-to-all waits for the busiest GPU. Expert-parallel load balancing (EPLB) places extra copies of hot experts and reshuffles placement so no GPU becomes the straggler. At EP72, 256 experts don't divide evenly (3 or 4 per GPU), which leaves room for those copies.

Illustrative: ~370 GB of routed-expert weights (671B x ~0.56 bytes in NVFP4), ~25 GB per GPU kept for attention and shared weights plus activations, one token routed to 8 experts, evenly spread routing. Scale-out packets are drawn 3x slower than NVLink; the per-GPU bandwidth gap is ~9x.

One rack, two jobs: disaggregated serving

Prefill and decode stress a GPU in opposite ways, so large deployments give each its own GPUs and its own parallelism. The InferenceX GB300 recipe for DeepSeek V4 Pro fills exactly one NVL72 rack: ten 4-GPU prefill workers and one 32-GPU decode worker, with the KV cache handed over NIXL inside the rack. Step through one request, then see how software alone raised this setup's throughput about 5× between April and June 2026.

One rack, two jobs - disaggregated serving of DeepSeek V4 Pro on GB300 NVL72

The 72 GPUs below are the rack's 18 compute trays of 4. The recipe disagg-gb300-10p1d-dep4-dep32 splits them into 10 prefill workers and 1 decode worker. Follow one request through.

1/5Request arrives: A new prompt arrives; the router hands it to a free prefill worker (P3).
P1T1
P2T2
P3T3
P4T4
P5T5
P6T6
P7T7
P8T8
P9T9
P10T10
D 1/8T11
D 2/8T12
D 3/8T13
D 4/8T14
D 5/8T15
D 6/8T16
D 7/8T17
D 8/8T18
Prefill worker (4 GPUs, DEP4)Decode worker (32 GPUs, DEP32)KV cache over NIXL

Request

prompt tokens

router

Prefill worker P3

4 GPUs · DEP4 · compute-bound

KV cache over NIXL

Decode worker

32 GPUs · DEP32 · bandwidth-bound

stream

Tokens

streamed to the user

Where the 72 GPUs go

40
32

10 prefill workers x 4 GPUs = 40 + 1 decode worker x 32 GPUs = 32 = 72 GPUs, one NVL72 rack

Prefill is compute-bound

A prompt's tokens are processed all at once, so prefill is a big batch of matrix math. Ten small DEP4 workers each take their own prompts in parallel, which keeps time to first token low as requests pile up.

Decode is bandwidth-bound

Each new token re-reads weights and KV cache for a tiny amount of math. One wide DEP32 worker spreads the experts thin and pools many streams into big batches, so each byte read from HBM serves more tokens.

Same rack, better software

5x

Day 0 (Apr 2026)
~2,200
June 2026
~11,200

Output tok/s per GPU for DeepSeek V4 Pro with SGLang disaggregated serving on GB300 at ~50 tok/s per user, in benchmarks - about 5x in two months from software alone.

For how prefill/decode disaggregation works on any cluster - and the KV transfer it costs - see Inference serving.

Tray placement is illustrative (one prefill worker per tray); the worker counts and GPU split are the recipe's. DEP = data-parallel attention with expert parallelism across the worker's GPUs.

The prefill/decode split itself is explained on Inference & Serving.

Where NVL72 pays off

Every serving deployment picks a point on one curve: tokens per second per GPU, which sets cost, against tokens per second per user, which sets how fast each reply streams. Batching more users onto a GPU raises throughput but slows each user down. GB300 NVL72's 72-GPU NVLink domain helps most in the middle of that curve, roughly 30-160 tokens/s per user, where large MoE models need expert parallelism across more than 8 GPUs; at the low and high extremes it lands near parity with an 8-GPU HGX B300. The dots below are DeepSeek R1 results in benchmarks.

Where NVL72 pays off - throughput vs interactivity

Pick a matchup and a metric. Dots are DeepSeek R1 benchmark results; the dashed lines only sketch the usual shape of the curve between them.

Only per-dollar results are published for this matchup.

middle of the curve, ~30-160 tok/s/user$0.05$0.1$0.2$0.5$1.0$2.0050100150200250tokens/s per user (interactivity) →↑ $ per million tokens (log)2x3.5x235: about equal$0.129$1.159$0.064$0.326
HGX B300GB300 NVL72shape illustrative, points from benchmarks

GB300 NVL72 vs HGX B300

On a 1K-in / 1K-out workload, the 72-GPU NVL72 rack is per dollar about 2x cheaper than an 8-GPU HGX B300 box at 88 tok/s per user and about 3.5x cheaper at 162, where a big MoE model's expert all-to-all traffic rides NVLink across the whole rack. By 235 tok/s per user the two cost about the same.

Long context: 128K in / 8K out

On DeepSeek R1, GB300 NVL72 does 226 vs 148 tokens/s per GPU for GB200 (1.53x), with a max decode batch of 40 vs 24 per GPU. No interactivity level is given, so it isn't plotted (LMSYS, Feb 2026).

  • HGX B300, 88 tok/s/user: $0.129 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
  • HGX B300, 162 tok/s/user: $1.159 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
  • GB300 NVL72, 88 tok/s/user: $0.064 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
  • GB300 NVL72, 162 tok/s/user: $0.326 per M tokens - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing
  • 235 tok/s/user: cost about equal - InferenceX by SemiAnalysis, DeepSeek R1 FP4, 1K in / 1K out, July 2026 pricing; no exact value, so the ring's height is illustrative

Benchmark caveats: InferenceX runs use random data with prefix caching off; disaggregated and aggregated per-GPU numbers aren't directly comparable; MLPerf latency bounds are loose, which hides much of the rack-scale advantage. InferenceX counts input and output tokens together, so an 8K-in / 1K-out test (about 89% input) looks 2-5x cheaper per token than a 1K / 1K test at the same speed. Each matchup uses its own workload, so don't compare values across matchups.

Picking a system

GB300 NVL72

Choose it for

Large MoE models (roughly 600B and up), long contexts, and latency targets in the middle of the curve.

Why

Expert all-to-all stays on the 130 TB/s NVLink domain, so decode can spread experts across 16-72 GPUs; fewer experts per GPU leaves more HBM for KV and bigger batches; a whole prefill + decode layout fits in one rack.

Tradeoff

Up to 142 kW liquid-cooled racks, and 72 GPUs share one NVLink failure domain.

HGX B300

Choose it for

Models whose weights and KV fit in ~2.1 TB: gpt-oss, 70B dense, Qwen3.5-397B, GLM-5.x, Kimi K2.x, MiniMax.

Why

Air-cooled 14.5 kW servers that fit existing data centers, an 8-GPU failure domain, similar per-GPU pricing to GB300 and broad cloud availability. In benchmarks it matches or beats GB300 on cost at low interactivity for some models (MiniMax-M3).

Tradeoff

Scale-up stops at 8 GPUs, so expert parallelism beyond 8 crosses a network about 9× slower than NVLink.

GB200 NVL72 (previous generation)

Choose it for

Workloads that don't need Blackwell Ultra's extra memory or attention speed.

Why

GB300 wins clearly only where 1.5× the HBM or 2× attention unlocks a configuration GB200 can't run - DeepSeek V4 (2.8× in benchmarks) or 128K contexts (1.5×). Elsewhere GB200 is often cheaper per token.

Tradeoff

Less HBM per GPU (about 186 GB) caps batch size and context sooner.

Operators increasingly measure all of this as : power, not GPU count, is usually the binding constraint on how much a site can serve.

What changes with Vera Rubin

Vera Rubin NVL72 began shipping in limited volume in September 2026. It keeps the GB300 NVL72 shape - 72 GPUs in one NVLink domain, 288 GB of HBM each - so the planner above gives the same answer for both: the same models fit with the same number of users. What changes is speed. HBM4 runs at 19.2 TB/s per GPU against 8 TB/s on Blackwell Ultra, NVLink 6 reaches 3 TB/s per GPU, and dense FP4 rises to 35 PF per GPU. The same deployment produces more tokens per second, at roughly 190-230 kW per rack.

One NVLink domain, five generations

Log scale. Pick a measure; the dashed rung is planned, not shipping.

HGX H100

8 GPUs · 2023

0.64 TB

GB200 NVL72

72 GPUs · 2025

13.4 TB

GB300 NVL72

72 GPUs · 2025-26

20.7 TB

Vera Rubin NVL72

72 GPUs · shipping 2H 2026

20.7 TB

Rubin Ultra Kyber NVL144

144 GPUs · planned 2027

~144 TB

GB300 to Vera Rubin: the same 20.7 TB - every model that fits one rack fits the other.

Dense (not sparse) FP4. Vera Rubin: 72 GPUs at 35 PF dense FP4 each. Rubin Ultra follows NVIDIA's 2025 roadmap; mid-2026 reports describe a redesign. Server and rack power are maximum or rated figures.

  • Vera Rubin NVL72 (formerly NVL144)

    NVIDIA first announced the Vera Rubin rack as Vera Rubin NVL144, counting the two dies inside each Rubin GPU. It now counts whole GPUs, so the same rack is called Vera Rubin NVL72: 72 Rubin GPUs and 36 Vera CPUs, the same footprint as GB300 NVL72.

  • Much more host memory

    Vera CPUs carry up to 54 TB of LPDDR5X per rack, about 2.6x the rack's HBM (GB300's Grace memory is a bit less than its HBM). Offloading KV cache to host memory inside the rack becomes far more useful.

  • Rubin CPX is off the roadmap

    NVIDIA announced Rubin CPX in 2025 as a GDDR7 chip for long-context prefill, then dropped it at GTC 2026. The 2026 split pairs Rubin GPUs for prefill with Groq 3 LPX racks for decode, under NVIDIA Dynamo.

  • Rubin Ultra is the next step up

    Planned for 2027: Kyber NVL144, 144 four-die GPUs with up to 1 TB of HBM4E each in one 800 VDC rack, and NVL576 linking eight racks. Mid-2026 reports described a two-die redesign and a slip to 2028; NVIDIA says the roadmap is intact.

Software moves the frontier as much as silicon

The same GB300 hardware posted very different numbers a few months apart. Between MLPerf Inference v5.1 (Sep 2025) and v6.0 (Apr 2026), DeepSeek R1 throughput per GPU rose 1.7× offline and 2.8× in the server scenario, and SGLang lifted DeepSeek V4 Pro about 5× between day 0 and June 2026. Turning on alone cut DeepSeek R1's about 21× at 150 tokens/s per user.

Software moves the frontier as much as silicon

Every jump below is on the same GB300 hardware - only the software changed. Tap a result to see where it sits in time.

Sep 2025MLPerf v5.1
Apr 2026MLPerf v6.0V4 Pro day 0
Jun 2026V4 Pro, tuned

Highlighted: the dates this result compares.

Compare hardware on the same software date. The same GB300 rack got 1.7-5x faster in a few months, so an older chip on newer software can outscore a faster chip measured on older software.

Pitfalls practitioners hit

288 GB isn't what you get

Usable HBM is about 268-275 GB per GPU on HGX B300 and about 278 GB on GB300. Size against the usable number, then reserve room for activations and buffers.

INT8 and FP64 were cut

Blackwell Ultra traded INT8 and FP64 throughput for FP4. Use FP8, NVFP4 or MXFP4; W4A16 INT4 still works.

HGX B300 is not a smaller GB300

Dense FP4 is 13.5 vs 15 PF per GPU and scale-up stops at 8 GPUs - tuned DeepSeek V4 on B300 ran better at EP4 than EP8.

Day-0 software is rough

New models hit engine bugs in their first weeks - hardcoded hidden sizes, FP8 scale errors producing NaNs, sliding-window memory bugs under disaggregation. Expect weeks of tuning.

Read benchmarks carefully

InferenceX runs random data with prefix caching off, disaggregated and aggregated per-GPU numbers aren't directly comparable, and MLPerf's loose latency bounds hide the rack-scale advantage.

A KV tier must be big to help

Host memory or flash only lifts hit rates when it is roughly 1.5-3× the HBM KV capacity; in one benchmark 3 TB of DRAM on a B300 node added just 1.36% hit rate. GB300's Grace memory is about 0.85× its HBM.

Big models load slowly

Kimi K3 took about 81 minutes to load on 8× B300 - which slows tuning, scaling and failover.

Hybrid models need a memory split

Models with linear-attention or state-space layers keep a fixed state pool beside the KV cache; the split between them must be tuned per workload.

Figures on this page come from vendor spec pages, engine recipes, MLPerf Inference results and InferenceX / LMSYS benchmarks from late 2025 to September 2026; benchmark results depend on workload, software version and latency target.