Systems & Infrastructure
Inference & Serving
IntermediateServing a model is a different problem from training one. Inference splits into two phases with opposite bottlenecks, and the whole stack - batching, parallelism, KV-cache management, routing, and disaggregation - exists to squeeze the most goodput out of expensive GPUs while holding latency SLOs. This page builds from the two phases up to datacenter-scale disaggregated serving.
What happens at inference time
LLMs generate autoregressively - one token at a time, each conditioned on everything before it. That work splits into two phases with opposite bottlenecks. processes the entire prompt in one parallel pass, populating the ; it's compute-bound and governs time-to-first-token (). then emits one token at a time, each step reading the whole KV cache back from ; it's memory-bandwidth-bound and governs time-per-output-token ( / inter-token latency).
Prefill vs decode on the GPU's roofline
Both phases run the same weights. The roofline caps how fast any kernel can go: left of the ridge, speed is limited by how fast HBM streams bytes (the slanted roof); right of it, by the Tensor Cores (the flat roof). Prefill does lots of math per byte read and lands on the flat roof. Decode does almost none and sits low on the slope - until batching lifts it.
Prefill16.1k FLOP/B
compute-boundOne read of the weights drives 16,384 prompt tokens of math, well past the 295 ridge. Compute is the wall - this is what sets time-to-first-token.
Decode7.8 FLOP/B
memory-boundEach step re-reads the model and a 2,304-token KV cache to make just 8 tokens, so Tensor Cores sit 97% idle and memory bandwidth is the wall. That sets time-per-output-token.
TTFT ≈ prefill
2.32 s
time to first token
TPOT ≈ decode step
42.7 ms
time per output token
End-to-end latency
24.17 s
prefill + 512 decode steps
Raise the batch and decode climbs the slope toward the ridge - that is why serving engines batch decode aggressively. A longer context adds KV bytes to every step and drags it back down. With this generation length, decode dominates end-to-end time.
Illustrative first-order estimates for a 70B-class model on an H100-class GPU (989 TFLOP/s BF16 dense, 3.35 TB/s HBM) - not a production latency model.
The KV cache, and why it dominates memory
In serving terms the KV cache is each request's working state - the keys and values for every token so far, held in HBM until the last output token - and it is the real limit on batch size. Whatever HBM the weights leave free caps how many sequences can decode at once, decode re-reads the whole cache every step, and when the budget runs out the engine must queue, preempt or offload requests. See KV Cache for the full treatment - sizing, PagedAttention, prefix caching and the tiering planner.
Serving across GPUs & nodes
A single request on one GPU is the easy case. Real deployments fan many concurrent requests across multiple GPUs and nodes, and the parallelism strategy you pick is dictated by how much they have to talk to each other - which interconnect can carry that traffic. How each strategy splits a model is covered on Model Sizing; here is what each one costs in traffic while serving.
One model, three ways to cut it across 4 GPUs
The grid is the model: layers run left to right, each layer's width top to bottom. Color shows which GPU holds each piece. Step a forward pass through and count the GPU-to-GPU traffic each cut creates.
Traffic this pass
0
all-reduces so far
Link it needs
NVLink, inside one node
Splits every layer's matrices across GPUs. Each forward pass needs frequent all-reduce - keep it intra-node over NVLink (~900 GB/s on H100, ~1.8 TB/s on Blackwell).
Schematic: 8 layers, 4 GPUs and top-2 expert routing; event counts are for this drawing, not a real model.
Rule of thumb: stays inside a node on NVLink; PP spans nodes over InfiniBand; appears once you serve MoE. They compose.
Splitting the KV cache by token, not by head
Tensor parallelism splits attention by head. MLA caches a single shared latent per token, so there is no head to split, and every GPU in a TP group keeps a full copy of each request's KV. splits the cache by token range instead: each GPU attends over its own slice and the partial results are merged, so KV memory per GPU falls by the number of GPUs.
Why MLA's KV is duplicated under tensor parallelism - and how decode context parallelism fixes it
One long request on an MLA model (DeepSeek-V3-style layout, about 68.6 KiB of KV per token in BF16), served by a group of GPUs.
Split attention by
GPUs
Context
The request's KV cache - 200,000 tokens, 13.1 GiB
Shown as 4 token ranges of about 50,000 tokens each (T0 to T3).
What each GPU holds for this request
Tensor parallelism splits attention by head. MLA caches one shared compressed latent per token - a single KV "head" - so there is nothing to split, and every GPU keeps the full cache. GQA models with fewer KV heads than GPUs (for example 8 KV heads on 16 GPUs) duplicate the same way.
KV per GPU, this request
13.1 GiB
a full copy on every GPU
KV stored across the group
52.4 GiB
the same bytes stored 4 times
Concurrent requests that fit
3
at 200,000 tokens, with an illustrative 40 GiB of free HBM per GPU for KV
vLLM DCP with Kimi K2.6
In benchmarks, TP throughput stopped growing at concurrency 64, while DCP kept scaling to 512 concurrent requests at about 3.3x the throughput per GPU.
NVIDIA Helix in TensorRT-LLM
Attention runs KV-parallel (token ranges spread across GPUs, as above), then the same GPUs regroup for tensor and expert parallelism in the FFN. In benchmarks with DeepSeek-R1 in FP4 at 1M tokens of context on GB300 NVL72, it served up to 32x more concurrent users with 1.8x better interactivity.
The cost: DCP adds two small collectives per layer on every decode step - gathering one token's query and merging the partial outputs - which is cheap next to reading a long KV cache.
Batching keeps the GPU busy
Because decode is bandwidth-bound, the way to use a GPU well is to run many sequences at once. Batching strategy is what turns idle silicon into throughput - and it has evolved a long way past the naive loop.
Same eight requests, three schedulers
Four batch slots on one GPU. Each request needs a prefill, then one decode step per output token. Request E has a long prompt. Watch where slots sit idle and where tokens wait.
Static batching
0/8 requests doneWait to collect a fixed batch, run it to completion, then start the next. Simple but the whole batch stalls on its slowest sequence - GPU idles.
Continuous / in-flight batching
0/8 requests doneSchedule at the iteration level: finished sequences leave and new ones join every step, keeping the GPU saturated. 3–5× throughput over a naive loop.
Chunked prefill
0/8 requests doneBreak a long prompt's prefill into chunks and interleave them with ongoing decode steps, so a big prompt doesn't freeze everyone else's tokens - smooths TTFT spikes.
Schematic: request sizes and durations are illustrative, chosen to show the scheduling pattern - not measured.
Serving engines
A handful of inference engines dominate production. They trade off ease of use, peak performance, hardware breadth, and how aggressively they manage the KV cache.
| Engine | Maker | License | Best at | Throughput | Latency/TTFT | Ease | HW | KV features | Quant | Multi-node | Disagg-PD |
|---|---|---|---|---|---|---|---|---|---|---|---|
| vLLM | vLLM project / community | Apache 2.0 | general-purpose, widest models, easy start | 3–5× naive | good | high (no compile) | NVIDIA/AMD/others | PagedAttention, prefix cache | FP8/INT4/AWQ/GPTQ | yes (TP/PP) | yes (w/ LMCache/Dynamo) |
| TensorRT-LLM | NVIDIA | Apache 2.0 | max perf on NVIDIA | highest on NVIDIA | best (compiled) | low (engine build) | NVIDIA only | paged KV, in-flight | FP8/INT4/FP4 | yes (TP/PP/EP) | yes |
| SGLang | LMSYS / SGLang | Apache 2.0 | shared-prefix / RAG / agents | up to 6.4× on shared prefixes | best on cached prefixes | medium | NVIDIA/AMD | RadixAttention prefix tree | FP8/INT4 | yes | yes |
| NVIDIA Dynamo | NVIDIA | Apache 2.0 | datacenter-scale orchestration (above vLLM/TRT-LLM/SGLang) | 7×/GPU (reported) | 2× TTFT via KV-routing (reported) | medium | NVIDIA-centric | KVBM offload, KV-aware routing | delegates to engine | yes (core) | yes (native) |
| HF TGI (legacy) | Hugging Face | Apache 2.0 | maintenance mode (Mar 2026) | lower; deprecated | moderate | was high | NVIDIA/others | basic | limited | limited | no |
| LMDeploy | Shanghai AI Lab | Apache 2.0 | quantized + long-context on NVIDIA | high | good | medium | NVIDIA | paged KV, prefix cache | strong (AWQ/INT4/FP8) | yes | partial |
Vendor/benchmark multipliers are reported on specific hardware and workloads, not universal. HF TGI is in archived/maintenance mode as of March 2026 - shown here as legacy.
Signature techniques per engine
Each engine is known for one or two core ideas. Knowing them tells you which workload each is built for.
One big idea per engine - and where each sits
Dynamo orchestrates a fleet from above; the three engines below it each run the model on the GPUs. Tap one to read its signature technique.
Orchestration layer
Inference engines
vLLM
Stores KV in fixed-size blocks like OS pages, eliminating fragmentation and enabling near-100% memory use; continuous batching keeps the GPU busy every iteration.
Disaggregated prefill–decode
Prefill and decode have opposite bottlenecks, so co-locating them on the same GPUs forces a TTFT-vs-TPOT tradeoff and leaves resources idle - a burst of prefill stalls everyone's decode. Disaggregation puts them in separate pools connected by a fast KV-transfer path, so each pool can be scaled and tuned independently for higher under SLOs.
Why a prefill burst stalls decode
Watch the gap between tokens for requests that are already streaming. Co-located, each new prompt's prefill takes over the GPU and their tokens pause. Disaggregated, prefill runs on separate GPUs and only its KV cache crosses over.
Interference. One GPU does both jobs, so each prefill pushes back the next decode step for every request already streaming - the gap between tokens (TPOT) spikes. Tuning for a fast first token (TTFT) and for smooth streaming pull against each other.
Schematic: one tick is one decode step; durations are illustrative, not measured.
The reported wins are large: DistServe up to 7.4× more requests / 12.6× tighter SLO (OSDI'24); Splitwise 2.35× at equal cost; Mooncake +75% real requests. But it isn't free: the KV cache must cross the network between pools. Disaggregation pays off at high QPS with multiple replicas, and is overkill at low QPS / single replica, where the transfer overhead dominates the gain.
Disaggregated goodput calculator
Goodput = requests/sec that meet both their TTFT and TPOT SLOs - not just requests served. Both modes own the same GPUs and the same raw capacity (22/s here); the difference is how much of it stays SLO-compliant. The headline is sustainable goodput: the highest load each mode can carry without breaking an SLO.
Co-located · sustainable
6.8 req/s
prefill & decode share every GPU · interference caps SLO-safe load
Disaggregated · sustainable
12.1 req/s
9.4 prefill / 2.6 decode GPUs · each pool meets its own SLO
SLO-safe throughput gain
1.8×
more SLO-compliant requests on the same GPUs
Co-located latency under load
prefill bursts delay first token → TTFT 1140 ms > 1000 ms SLO
Disaggregated latency under load
Each pool is batched & queued for its own SLO, so prefill bursts never stall decode tokens.
Goodput vs offered load
Tap or click the chart to set λ
at 9 req/s (λ): disaggregated 9.0 · co-located 0.0 req/s goodput
Both modes serve the same requests up to capacity, but goodput only counts the ones inside both SLOs. Co-located goodput collapses at 6.8/s when interference breaks an SLO; disaggregated holds up to 12.1/s on the same GPUs.
Real-world anchors: DistServe reports up to ~7.4× more requests (≈2× under tight SLOs), and Splitwise ~2.35× at equal cost - by giving prefill and decode their own pools.
Illustrative first-order model (70B-class model, H100-class GPU): prefill time ∝ prompt length, decode step ∝ context, with a load-dependent interference tax on the co-located mode. The interference model is illustrative and the cited multipliers (DistServe ~7.4×, Splitwise ~2.35×) are hardware- and workload-specific.
Inference routing
Once you have a pool of workers, where you send each request matters. KV-cache-aware routing sends a request to the worker that already holds its prefix, avoiding a redundant prefill. The router also handles prefill/decode assignment and load-aware balancing to keep pools even.
Inference routing - send work where the prefix already lives
Twenty requests arrive in order. Most share a system prompt (S); some are unique (U). The bar under each request is its prompt: the shared prefix, then its own tokens. A request that lands on a worker already holding the prefix reuses that KV instead of prefilling it again.
Incoming queue · 20 waiting
W1worker 1
0
no KVno prefix cached
W2worker 2
0
no KVno prefix cached
W3worker 3
0
no KVno prefix cached
W4worker 4
0
no KVno prefix cached
Prefix cache hits
0/0
shared-prompt requests routed so far
Redundant prefill avoided
0 tok
shared-prefix tokens not recomputed
Prefill time saved
0 ms
summed over cache hits, vs re-prefilling every time
All three policies on the full 20-request stream
Round-robin and load-aware land on near-identical prefix reuse - both ignore where the KV lives. That near-tie is the lesson: only KV-aware routing turns the shared prompt into cache hits. It uses cost-based routing, not blind pinning - a cached worker absorbs extra load until a fresh worker is cheap enough to be worth a re-prefill, so the prefix replicates across workers. Add workers and the shared work spreads over more cache replicas; raise the shared-prefix ratio and it concentrates on fewer workers to maximize reuse. That hit-rate-vs-load balance is what NVIDIA Dynamo and vLLM tune.
KV-aware routing keeps shared-prefix requests on workers holding their KV, turning repeat system prompts into near-free cache hits - the bigger the shared-prefix ratio, the larger the TTFT win.
Illustrative token/latency model: an 800-token system prompt at 0.02 ms of prefill per token.
Metrics & SLOs
You can't tune what you don't measure - but not every metric matters for every workload. Each one is set by a specific phase, and the right SLO depends on what the workload looks like.
Five metrics on one request - and who cares about which
Pick a workload to change the request's shape, then a metric to see what it measures. Timings are illustrative.
Short prompt, streamed reply, one human waiting.
TTFTset by: Prefill
Time to first token
How long before anything appears. Sets perceived responsiveness; grows with prompt length.
What interactive chat cares about
- TTFTCritical
- TPOT / ITLCritical
- E2E latencyMatters
- ThroughputMatters
- GoodputMatters
TTFT must feel instant and TPOT must beat human reading speed (a few tokens/s perceived). Balanced - both latencies are user-facing.
SLOs drive capacity: you provision enough prefill and decode workers so that, at your peak request rate, goodput stays above target - not just raw throughput. That's exactly what the calculator above estimates.
Going further
The frontier of serving optimization, mapped by what each technique buys you. Every lever trades against the others - push latency and you spend memory or compute; the art is holding goodput while you do.
The tradeoff triangle - nine techniques, three levers
Each technique sits nearest the lever it mainly buys. Tap a dot or a name to read it.
Latency
Optimize Latency
Optimize Throughput
Optimize Memory
2 Prefix / KV caching
Reuse KV for shared prompt prefixes to skip redundant prefill - slashes TTFT.
Primary win: latency - it also eases memory pressure
Many techniques span more than one lever (e.g. prefix caching cuts both latency and memory pressure) - they're placed under their primary win.
Serving ties the whole stack together
Every serving decision traces back to model size, the KV cache, and the GPUs underneath. Size the model, understand the cache, then map it onto the infrastructure.