Systems & Infrastructure

Inference & Serving

Intermediate

Serving a model is a different problem from training one. Inference splits into two phases with opposite bottlenecks, and the whole stack - batching, parallelism, KV-cache management, routing, and disaggregation - exists to squeeze the most goodput out of expensive GPUs while holding latency SLOs. This page builds from the two phases up to datacenter-scale disaggregated serving.

What happens at inference time

LLMs generate autoregressively - one token at a time, each conditioned on everything before it. That work splits into two phases with opposite bottlenecks. processes the entire prompt in one parallel pass, populating the ; it's compute-bound and governs time-to-first-token (). then emits one token at a time, each step reading the whole KV cache back from ; it's memory-bandwidth-bound and governs time-per-output-token ( / inter-token latency).

Prefill vs decode on the GPU's roofline

Both phases run the same weights. The roofline caps how fast any kernel can go: left of the ridge, speed is limited by how fast HBM streams bytes (the slanted roof); right of it, by the Tensor Cores (the flat roof). Prefill does lots of math per byte read and lands on the flat roof. Decode does almost none and sits low on the slope - until batching lifts it.

4/10Batch sweep, B = 8: decode reaches 2.7% of peak compute
1101001k10k100k1M1101001karithmetic intensity, FLOP/byte (log) →↑ TFLOP/smemory-bound: Tensor Cores idlespeed set by HBM bandwidthcompute-boundspeed set by Tensor Coresmemory roof · HBM 3.35 TB/scompute roof · 989 TFLOP/sridge 295Decode at batch 1: 1.0 FLOP/byteDecode at batch 2: 2.0 FLOP/byteDecode at batch 4: 4.0 FLOP/byteDecode at batch 8: 7.8 FLOP/byteDecode at batch 16: 15 FLOP/byteDecode at batch 32: 30 FLOP/byteDecode at batch 64: 55 FLOP/byteDecode at batch 128: 96 FLOP/byteDecode at batch 256: 153 FLOP/byteDecode at batch 512: 218 FLOP/byteprefill 16.1kdecode 7.8 · B=8
Memory roof: HBM bandwidth × intensityCompute roof: peak Tensor Core FLOP/sRidge ≈ 295 FLOP/byteDecode at B = 1, 2, 4 … 512

Prefill16.1k FLOP/B

compute-bound
Tensor Cores (compute)100% · saturated
HBM bandwidth (memory)2% · idle

One read of the weights drives 16,384 prompt tokens of math, well past the 295 ridge. Compute is the wall - this is what sets time-to-first-token.

Decode7.8 FLOP/B

memory-bound
Tensor Cores (compute)3% · idle
HBM bandwidth (memory)100% · saturated

Each step re-reads the model and a 2,304-token KV cache to make just 8 tokens, so Tensor Cores sit 97% idle and memory bandwidth is the wall. That sets time-per-output-token.

TTFT ≈ prefill

2.32 s

time to first token

TPOT ≈ decode step

42.7 ms

time per output token

End-to-end latency

24.17 s

prefill + 512 decode steps

Raise the batch and decode climbs the slope toward the ridge - that is why serving engines batch decode aggressively. A longer context adds KV bytes to every step and drags it back down. With this generation length, decode dominates end-to-end time.

Illustrative first-order estimates for a 70B-class model on an H100-class GPU (989 TFLOP/s BF16 dense, 3.35 TB/s HBM) - not a production latency model.

The KV cache, and why it dominates memory

In serving terms the KV cache is each request's working state - the keys and values for every token so far, held in HBM until the last output token - and it is the real limit on batch size. Whatever HBM the weights leave free caps how many sequences can decode at once, decode re-reads the whole cache every step, and when the budget runs out the engine must queue, preempt or offload requests. See KV Cache for the full treatment - sizing, PagedAttention, prefix caching and the tiering planner.

Serving across GPUs & nodes

A single request on one GPU is the easy case. Real deployments fan many concurrent requests across multiple GPUs and nodes, and the parallelism strategy you pick is dictated by how much they have to talk to each other - which interconnect can carry that traffic. How each strategy splits a model is covered on Model Sizing; here is what each one costs in traffic while serving.

One model, three ways to cut it across 4 GPUs

The grid is the model: layers run left to right, each layer's width top to bottom. Color shows which GPU holds each piece. Step a forward pass through and count the GPU-to-GPU traffic each cut creates.

1/9A token's activations enter layer 1.
L1L2L3L4L5L6L7L8slice 1slice 2slice 3slice 4Layer 1, slice 1: GPU 0Layer 1, slice 2: GPU 1Layer 1, slice 3: GPU 2Layer 1, slice 4: GPU 3Layer 2, slice 1: GPU 0Layer 2, slice 2: GPU 1Layer 2, slice 3: GPU 2Layer 2, slice 4: GPU 3Layer 3, slice 1: GPU 0Layer 3, slice 2: GPU 1Layer 3, slice 3: GPU 2Layer 3, slice 4: GPU 3Layer 4, slice 1: GPU 0Layer 4, slice 2: GPU 1Layer 4, slice 3: GPU 2Layer 4, slice 4: GPU 3Layer 5, slice 1: GPU 0Layer 5, slice 2: GPU 1Layer 5, slice 3: GPU 2Layer 5, slice 4: GPU 3Layer 6, slice 1: GPU 0Layer 6, slice 2: GPU 1Layer 6, slice 3: GPU 2Layer 6, slice 4: GPU 3Layer 7, slice 1: GPU 0Layer 7, slice 2: GPU 1Layer 7, slice 3: GPU 2Layer 7, slice 4: GPU 3Layer 8, slice 1: GPU 0Layer 8, slice 2: GPU 1Layer 8, slice 3: GPU 2Layer 8, slice 4: GPU 3all 4 GPUs sync after every layer - keep them on NVLink in one node
GPU 0GPU 1GPU 2GPU 3GPU-to-GPU traffic

Traffic this pass

0

all-reduces so far

Link it needs

NVLink, inside one node

Splits every layer's matrices across GPUs. Each forward pass needs frequent all-reduce - keep it intra-node over NVLink (~900 GB/s on H100, ~1.8 TB/s on Blackwell).

Schematic: 8 layers, 4 GPUs and top-2 expert routing; event counts are for this drawing, not a real model.

Rule of thumb: stays inside a node on NVLink; PP spans nodes over InfiniBand; appears once you serve MoE. They compose.

Splitting the KV cache by token, not by head

Tensor parallelism splits attention by head. MLA caches a single shared latent per token, so there is no head to split, and every GPU in a TP group keeps a full copy of each request's KV. splits the cache by token range instead: each GPU attends over its own slice and the partial results are merged, so KV memory per GPU falls by the number of GPUs.

Why MLA's KV is duplicated under tensor parallelism - and how decode context parallelism fixes it

One long request on an MLA model (DeepSeek-V3-style layout, about 68.6 KiB of KV per token in BF16), served by a group of GPUs.

Split attention by

GPUs

Context

The request's KV cache - 200,000 tokens, 13.1 GiB

Shown as 4 token ranges of about 50,000 tokens each (T0 to T3).

What each GPU holds for this request

GPU 0
13.1 GiB
GPU 1
13.1 GiB
GPU 2
13.1 GiB
GPU 3
13.1 GiB

Tensor parallelism splits attention by head. MLA caches one shared compressed latent per token - a single KV "head" - so there is nothing to split, and every GPU keeps the full cache. GQA models with fewer KV heads than GPUs (for example 8 KV heads on 16 GPUs) duplicate the same way.

KV per GPU, this request

13.1 GiB

a full copy on every GPU

KV stored across the group

52.4 GiB

the same bytes stored 4 times

Concurrent requests that fit

3

at 200,000 tokens, with an illustrative 40 GiB of free HBM per GPU for KV

vLLM DCP with Kimi K2.6

TP, plateau at concurrency 641,863 tok/s/GPU
DCP, concurrency 5126,091 tok/s/GPU

In benchmarks, TP throughput stopped growing at concurrency 64, while DCP kept scaling to 512 concurrent requests at about 3.3x the throughput per GPU.

NVIDIA Helix in TensorRT-LLM

Attention runs KV-parallel (token ranges spread across GPUs, as above), then the same GPUs regroup for tensor and expert parallelism in the FFN. In benchmarks with DeepSeek-R1 in FP4 at 1M tokens of context on GB300 NVL72, it served up to 32x more concurrent users with 1.8x better interactivity.

The cost: DCP adds two small collectives per layer on every decode step - gathering one token's query and merging the partial outputs - which is cheap next to reading a long KV cache.

Batching keeps the GPU busy

Because decode is bandwidth-bound, the way to use a GPU well is to run many sequences at once. Batching strategy is what turns idle silicon into throughput - and it has evolved a long way past the naive loop.

Same eight requests, three schedulers

Four batch slots on one GPU. Each request needs a prefill, then one decode step per output token. Request E has a long prompt. Watch where slots sit idle and where tokens wait.

1/17
Static batching
0/8 requests done

Wait to collect a fixed batch, run it to completion, then start the next. Simple but the whole batch stalls on its slowest sequence - GPU idles.

S1
A
A
A
A
E
E
E
E
S2
B
B
B
B
B
B
B
B
F
F
F
F
S3
C
C
C
C
C
G
G
G
G
G
S4
D
D
D
D
D
D
H
H
H
Continuous / in-flight batching
0/8 requests done

Schedule at the iteration level: finished sequences leave and new ones join every step, keeping the GPU saturated. 3–5× throughput over a naive loop.

S1
A
A
A
A
E
E
E
E
H
H
H
S2
B
B
B
B
B
B
B
B
S3
C
C
C
C
C
F
F
F
F
S4
D
D
D
D
D
D
G
G
G
G
G
Chunked prefill
0/8 requests done

Break a long prompt's prefill into chunks and interleave them with ongoing decode steps, so a big prompt doesn't freeze everyone else's tokens - smooths TTFT spikes.

S1
A
A
A
A
E1
E2
E3
E4
E
E
E
S2
B
B
B
B
B
B
B
B
H
H
H
S3
C
C
C
C
C
F
F
F
F
S4
D
D
D
D
D
D
G
G
G
G
G
time (illustrative steps) →
Prefill (or prefill chunk)Decode step - one tokenToken waiting on a long prefillIdle slot

Schematic: request sizes and durations are illustrative, chosen to show the scheduling pattern - not measured.

Serving engines

A handful of inference engines dominate production. They trade off ease of use, peak performance, hardware breadth, and how aggressively they manage the KV cache.

EngineMakerLicenseBest atThroughputLatency/TTFTEaseHWKV featuresQuantMulti-nodeDisagg-PD
vLLMvLLM project / communityApache 2.0general-purpose, widest models, easy start3–5× naivegoodhigh (no compile)NVIDIA/AMD/othersPagedAttention, prefix cacheFP8/INT4/AWQ/GPTQyes (TP/PP)yes (w/ LMCache/Dynamo)
TensorRT-LLMNVIDIAApache 2.0max perf on NVIDIAhighest on NVIDIAbest (compiled)low (engine build)NVIDIA onlypaged KV, in-flightFP8/INT4/FP4yes (TP/PP/EP)yes
SGLangLMSYS / SGLangApache 2.0shared-prefix / RAG / agentsup to 6.4× on shared prefixesbest on cached prefixesmediumNVIDIA/AMDRadixAttention prefix treeFP8/INT4yesyes
NVIDIA DynamoNVIDIAApache 2.0datacenter-scale orchestration (above vLLM/TRT-LLM/SGLang)7×/GPU (reported)2× TTFT via KV-routing (reported)mediumNVIDIA-centricKVBM offload, KV-aware routingdelegates to engineyes (core)yes (native)
HF TGI (legacy)Hugging FaceApache 2.0maintenance mode (Mar 2026)lower; deprecatedmoderatewas highNVIDIA/othersbasiclimitedlimitedno
LMDeployShanghai AI LabApache 2.0quantized + long-context on NVIDIAhighgoodmediumNVIDIApaged KV, prefix cachestrong (AWQ/INT4/FP8)yespartial

Vendor/benchmark multipliers are reported on specific hardware and workloads, not universal. HF TGI is in archived/maintenance mode as of March 2026 - shown here as legacy.

Signature techniques per engine

Each engine is known for one or two core ideas. Knowing them tells you which workload each is built for.

One big idea per engine - and where each sits

Dynamo orchestrates a fleet from above; the three engines below it each run the model on the GPUs. Tap one to read its signature technique.

Orchestration layer

Inference engines

GPUs

vLLM

Stores KV in fixed-size blocks like OS pages, eliminating fragmentation and enabling near-100% memory use; continuous batching keeps the GPU busy every iteration.

Disaggregated prefill–decode

Prefill and decode have opposite bottlenecks, so co-locating them on the same GPUs forces a TTFT-vs-TPOT tradeoff and leaves resources idle - a burst of prefill stalls everyone's decode. Disaggregation puts them in separate pools connected by a fast KV-transfer path, so each pool can be scaled and tuned independently for higher under SLOs.

Why a prefill burst stalls decode

Watch the gap between tokens for requests that are already streaming. Co-located, each new prompt's prefill takes over the GPU and their tokens pause. Disaggregated, prefill runs on separate GPUs and only its KV cache crosses over.

Arrivals
New requests
each needs a prefill
B↓C↓D↓E↓
GPU 0
prefill and decode share it
B
C
D
E
Gap between tokens
requests already streaming
stall
stall
stall
stall
time (illustrative steps) →
1/25Decode step - each streaming request gets its next token

Interference. One GPU does both jobs, so each prefill pushes back the next decode step for every request already streaming - the gap between tokens (TPOT) spikes. Tuning for a fast first token (TTFT) and for smooth streaming pull against each other.

Prefill (compute-bound)Decode step (bandwidth-bound)KV-cache transferStalled token gap

Schematic: one tick is one decode step; durations are illustrative, not measured.

The reported wins are large: DistServe up to 7.4× more requests / 12.6× tighter SLO (OSDI'24); Splitwise 2.35× at equal cost; Mooncake +75% real requests. But it isn't free: the KV cache must cross the network between pools. Disaggregation pays off at high QPS with multiple replicas, and is overkill at low QPS / single replica, where the transfer overhead dominates the gain.

Disaggregated goodput calculator

Goodput = requests/sec that meet both their TTFT and TPOT SLOs - not just requests served. Both modes own the same GPUs and the same raw capacity (22/s here); the difference is how much of it stays SLO-compliant. The headline is sustainable goodput: the highest load each mode can carry without breaking an SLO.

Offered load λrequests/sec (drives the latency readout)9/s
Total GPUssame hardware budget for both modes12
Prompt lengthdrives prefill cost / TTFT3,072
Output lengthdrives decode steps / TPOT256
TTFT SLOtime-to-first-token, ms1000 ms
TPOT SLOtime-per-output-token, ms80 ms

Co-located · sustainable

6.8 req/s

prefill & decode share every GPU · interference caps SLO-safe load

Disaggregated · sustainable

12.1 req/s

9.4 prefill / 2.6 decode GPUs · each pool meets its own SLO

SLO-safe throughput gain

1.8×

more SLO-compliant requests on the same GPUs

Co-located latency under load

TTFT1140 ms ✗ (≤ 1000)
TPOT94 ms ✗ (≤ 80)

prefill bursts delay first token → TTFT 1140 ms > 1000 ms SLO

Disaggregated latency under load

TTFT747 ms ✓ (≤ 1000)
TPOT63 ms ✓ (≤ 80)

Each pool is batched & queued for its own SLO, so prefill bursts never stall decode tokens.

Goodput vs offered load

Tap or click the chart to set λ

01020300510152025offered load, req/s →↑ goodput, req/sone modeboth miss an SLO6.8/s12.1/sλ 9/s

at 9 req/s (λ): disaggregated 9.0 · co-located 0.0 req/s goodput

Disaggregated goodputCo-located goodputRequests served (same for both)SLO violated

Both modes serve the same requests up to capacity, but goodput only counts the ones inside both SLOs. Co-located goodput collapses at 6.8/s when interference breaks an SLO; disaggregated holds up to 12.1/s on the same GPUs.

Real-world anchors: DistServe reports up to ~7.4× more requests (≈2× under tight SLOs), and Splitwise ~2.35× at equal cost - by giving prefill and decode their own pools.

Illustrative first-order model (70B-class model, H100-class GPU): prefill time ∝ prompt length, decode step ∝ context, with a load-dependent interference tax on the co-located mode. The interference model is illustrative and the cited multipliers (DistServe ~7.4×, Splitwise ~2.35×) are hardware- and workload-specific.

Inference routing

Once you have a pool of workers, where you send each request matters. KV-cache-aware routing sends a request to the worker that already holds its prefix, avoiding a redundant prefill. The router also handles prefill/decode assignment and load-aware balancing to keep pools even.

Inference routing - send work where the prefix already lives

Twenty requests arrive in order. Most share a system prompt (S); some are unique (U). The bar under each request is its prompt: the shared prefix, then its own tokens. A request that lands on a worker already holding the prefix reuses that KV instead of prefilling it again.

Routing policy
1/21Requests queue at the router.

Incoming queue · 20 waiting

S1S2U3S4S5S6U7S8S9U10S11S12S13U14S15S16U17S18S19U20

W1

0

no KV

W2

0

no KV

W3

0

no KV

W4

0

no KV

Shared prefix, prefilledShared prefix, reused from cacheRequest's own tokens

Prefix cache hits

0/0

shared-prompt requests routed so far

Redundant prefill avoided

0 tok

shared-prefix tokens not recomputed

Prefill time saved

0 ms

summed over cache hits, vs re-prefilling every time

All three policies on the full 20-request stream

Round-robin
10/14 hits · 112 ms prefill saved · busiest worker 5
KV-aware
11/14 hits · 123 ms prefill saved · busiest worker 5
Load-aware
10/14 hits · 112 ms prefill saved · busiest worker 5

Round-robin and load-aware land on near-identical prefix reuse - both ignore where the KV lives. That near-tie is the lesson: only KV-aware routing turns the shared prompt into cache hits. It uses cost-based routing, not blind pinning - a cached worker absorbs extra load until a fresh worker is cheap enough to be worth a re-prefill, so the prefix replicates across workers. Add workers and the shared work spreads over more cache replicas; raise the shared-prefix ratio and it concentrates on fewer workers to maximize reuse. That hit-rate-vs-load balance is what NVIDIA Dynamo and vLLM tune.

KV-aware routing keeps shared-prefix requests on workers holding their KV, turning repeat system prompts into near-free cache hits - the bigger the shared-prefix ratio, the larger the TTFT win.

Illustrative token/latency model: an 800-token system prompt at 0.02 ms of prefill per token.

Metrics & SLOs

You can't tune what you don't measure - but not every metric matters for every workload. Each one is set by a specific phase, and the right SLO depends on what the workload looks like.

Five metrics on one request - and who cares about which

Pick a workload to change the request's shape, then a metric to see what it measures. Timings are illustrative.

Short prompt, streamed reply, one human waiting.

TTFT
request arriveslast token
prefill one output token

TTFTset by: Prefill

Time to first token

How long before anything appears. Sets perceived responsiveness; grows with prompt length.

What interactive chat cares about

  • TTFTCritical
  • TPOT / ITLCritical
  • E2E latencyMatters
  • ThroughputMatters
  • GoodputMatters

TTFT must feel instant and TPOT must beat human reading speed (a few tokens/s perceived). Balanced - both latencies are user-facing.

SLOs drive capacity: you provision enough prefill and decode workers so that, at your peak request rate, goodput stays above target - not just raw throughput. That's exactly what the calculator above estimates.

Going further

The frontier of serving optimization, mapped by what each technique buys you. Every lever trades against the others - push latency and you spend memory or compute; the art is holding goodput while you do.

The tradeoff triangle - nine techniques, three levers

Each technique sits nearest the lever it mainly buys. Tap a dot or a name to read it.

Latency

ThroughputMemory

Optimize Latency

Optimize Throughput

Optimize Memory

2 Prefix / KV caching

Reuse KV for shared prompt prefixes to skip redundant prefill - slashes TTFT.

Primary win: latency - it also eases memory pressure

Many techniques span more than one lever (e.g. prefix caching cuts both latency and memory pressure) - they're placed under their primary win.

Serving ties the whole stack together

Every serving decision traces back to model size, the KV cache, and the GPUs underneath. Size the model, understand the cache, then map it onto the infrastructure.