Memory & Efficiency
KV Cache
FundamentalThe KV cache is what makes token generation fast - and what makes inference memory-bound. This page builds from the core idea to how production systems tier it across HBM, DRAM, NVMe, and pooled storage.
Prerequisite: this page assumes you know what keys and values are. If not, read Attention Mechanisms (and Tokenization) first.
Why a cache at all?
Generating each new token requires attending to every previous token. Without a cache, the model recomputes the keys and values for the entire sequence at every single step - an O(n²) explosion. The KV cache saves those keys and values once and reuses them, turning generation from quadratic to linear. The price is memory: that cache grows with every token, every layer, every concurrent user. Below you can watch the difference, then size it for real models.
Models have no memory
The reason a cache is even possible starts here: a is stateless. It keeps nothing between steps - feed it the same tokens and it does the identical work. To write token N it must attend to all N−1 prior tokens, so without saving anything it would recompute every token's keys and values from scratch, every step. The KV cache is the workaround.
A language model has no memory
Each call is a pure function: the full token sequence goes in, one next token comes out, and nothing is kept for the next call. So every call is fed - and processes - the whole prefix again.
Input this call
model
f(tokens)
kept for next call
nothing
Next token
catThe first call sees one token. Everything it works out about "The" is thrown away when it returns.
Because the model is stateless, the keys and values for past tokens are identical at every step - so there is no reason to recompute them. The KV cache simply stores each token's K/V once and reuses it. That is the entire trick, and the rest of this page is about the price it charges: memory.
What exactly gets cached
Self-attention is why keys and values are the things worth keeping. Each new token forms a query and scores it against every prior token's key, then blends their values. Crucially, a past token's K and V never change - so they can be computed once and reused forever.
Self-attention: score, softmax, blend
The newest token forms a Query, scores it against every token's Key (Q·K / √d), turns the scores into weights with softmax, then mixes their Values by those weights. All numbers are toy 4-d vectors, for illustration only.
| Token | K | Q·K/√d | Softmax weight | V |
|---|---|---|---|---|
Cache cached · frozen | 0.80 | 0.49 score 0.80 | ||
the cached · frozen | −0.41 | 0.14 score −0.41 | ||
keys new · computed now | 0.55 | 0.37 score 0.55 |
Strongest link: "Cache", 2 tokens back - not the neighbor "the". Weights follow what the Q and K vectors encode, not how close the tokens are.
Blend: output = Σ weight × V(bars zoomed in)
The blended vector is what this token carries forward. Only its Query - and its own K and V - were new this step; every other K and V row was read from the cache unchanged.
attn(Q, K, V) = softmax(Q·Kᵀ / √d) · V
A past token's K and V depend only on that token and the weights - never on what comes after. They are frozen the moment the token is processed, so they can be computed once and cached forever. That is precisely what the KV cache stores.
Quadratic work, or linear?
Skipping the recompute turns generation from O(n²) into O(n). Step through token generation and watch the work tally for each mode - by a thousand tokens the no-cache path is doing hundreds of times the attention work.
The work triangle - recompute vs cache
Each row is one decode step; each square is one prefix token whose keys and values that step needs. Without a cache the whole row is recomputed, so the total work is the area of the triangle. With a cache only the new token (the diagonal) is computed and the rest is read back.
Step 1: no cache computed 1 tokens this step, cached computed 1. Totals so far 1 vs 1.
No cache
O(n²)This step: 1 computed
With KV cache
O(n)This step: 1 computed
Cumulative work
fixed axesNo cache has done 1.0× the work
The gap widens with every token. At a 1,000-token context the no-cache path does about 501× the attention work of the cached path - which is why no production decoder runs without a KV cache. The cost moves from compute to memory.
The KV cache doesn't fit - so it tiers
is fast but tiny. At scale the KV cache dwarfs it, so serving stacks spill cold KV down a hierarchy - , local , then pooled network storage. Each step down is ~10–100× more capacity but ~10× slower. The art is keeping hot KV in HBM and demoting the rest rather than throwing it away and recomputing.
KV cache hierarchy - demote, then promote back
Each lettered block is one user's KV context. New requests land in HBM; when a tier is full the least-recently-used block is demoted one tier down instead of deleted. Click any block to make that user return.
The cache is empty.
80–288 GB / GPU · 0/3 used
1–2 TB / node · 0/4 used
~30–100 TB / node · 0/6 used
PB-scale · 0/8 used
Empty hierarchy
Every tier has free slots. Press play to watch traffic arrive, or add a request yourself.
Demotion instead of eviction is the whole reason tiered KV systems exist: a returning user whose context is still in DRAM, NVMe or pooled storage is read back and re-promoted rather than recomputed from scratch. Turn off local NVMe to send cold blocks straight to the pooled network tier.
Size it: the tiering & capacity planner
Size a deployment the way a KV cache builds up: a real 2026 model sets the bytes per token (dense vs vs sparse/ changes it dramatically), the workload mix sets tokens per session, and users plus retention set how many sessions you hold. Then pick a GPU fleet, split it into prefill and decode, and watch the KV fill decode HBM and CPU DRAM before the rest lands on the shared tier.
KV tiering & capacity planner
Five steps in the order a KV cache is made: the model sets bytes per token, the workload sets tokens per session, users and retention set how many sessions you hold, and the fleet decides where it all lands - HBM, then CPU DRAM and finally the shared tier.
1Model
MLA latent (512 + 64 RoPE) per layer - tiny cache for a 671B model.
2Workload
What a session holds - tokens per session set the KV per session.
Chat50%
conversation history
RAG30%
question + retrieved documents
Agentic / coding20%
long tool-use sessions
capped at the model's 160K
Images0%
prompt + images
DeepSeek R1 doesn't take images input
Video0%
prompt + video
DeepSeek R1 doesn't take video input
Audio0%
prompt + speech
DeepSeek R1 doesn't take audio input
Average session: 59K tokens = 2.1 GB of KV on DeepSeek R1. Per session: Chat 589 MB · RAG 2.3 GB · Agentic / coding 5.8 GB
3Users and retention
How many sessions' KV you hold at once.
Real time: a session's KV is dropped when it goes idle, so a returning user's context is recomputed from scratch. Pick a retention window to keep it - idle KV moves down to CPU DRAM and the shared tier, where it can be reloaded instead.
4Hardware
Decode nodes hold the session KV; prefill nodes process prompts and hand the KV over.
Hopper baseline - 80GB HBM3 per GPU, 900 GB/s NVLink.
Aggregated serving: all 4 nodes run prefill and decode and hold session KV.
5Where the KV lands
Model weights
671.0 GB
per copy · 2 decode copies, 2 nodes each
KV per session
2.1 GB
59K tokens on average
Total KV demand
2.08 TB
1,000 live sessions, real time
On the shared tier
-
what's left after HBM and DRAM
Live sessions need 2.08 TB of KV in decode HBM, but only 962.0 GB is free after the weights - about 451 of 1,000 sessions can decode at once. The rest wait, or are swapped out and back in.
Where the KV cache lands
demand 2.08 TB
Each tier across the 4 decode nodes, drawn to its own capacity
50% of 2.00 TB host memory per node
Pooled storage every node can read, so a session's KV can be reloaded on any node. No fixed capacity - drawn open-ended. Light: what spilled; solid: what is stored after VAST's ~1.2:1 data reduction on KV data.
HBM fills - 1.14 TB spills to CPU DRAM
45% of the KV cache fits in HBM after the weights; the other 55% lives in DRAM, one hop away. Nothing reaches the shared tier.
Each decode node group holds one copy of the weights plus ~10% activation headroom; the KV cache gets the rest of HBM. Live sessions are placed first, then idle sessions kept for retention, then burst headroom. The shared tier is pooled network storage (VAST, NVIDIA CMX and similar) that every node can read, so a session can be reloaded on any node. Illustrative, not a benchmark.
Attention decides the bill: KV per session by model
KV cache for one session on every model in the planner, each at its usual KV precision. The largest needs 104× the smallest at the same context.
Session length
Log scale. Models with a shorter maximum context are shown at that maximum; pure state-space models keep a fixed-size state instead of a per-token KV cache.
Software that stretches the cache
Before adding hardware, modern serving engines wring far more out of the HBM you already have. These four techniques are the backbone of every production LLM stack.
PagedAttention
2–4× throughputvLLM
Stores KV in fixed-size pages (like OS virtual memory) instead of one contiguous block per request. Eliminates the 60–80% fragmentation waste of pre-allocating max-length buffers, so far more requests fit in HBM at once.
RadixAttention
5–6× on shared prefixesSGLang
Organizes cached prefixes in a radix tree so requests that share a system prompt or few-shot examples reuse the same KV blocks. Huge for chat and agent workloads where every request starts with the same instructions.
Continuous batching
Higher GPU utilizationOrca / vLLM
Instead of waiting for a whole batch to finish, it swaps completed sequences out and new ones in at every decode step. Keeps the GPU saturated when requests have wildly different output lengths.
12 requests finished in 12 steps, 1 idle slot-step
A static batch finishes 7 in the same time (illustrative lengths).
Prefix / KV reuse
TTFT 11s → 1.5sLMCache · Mooncake
“Prefill once, reuse everywhere.” Cache the KV of a long document or shared context and serve it to many requests from a fast tier instead of recomputing prefill each time. In benchmarks, VAST + LMCache cut time-to-first-token ~7× on 128K contexts.
Time to first token, 128K-token context
The shared context's KV is already stored, so the request loads it instead of recomputing - ~7× faster to first token. (from a VAST + LMCache benchmark)
Prefix caching a hybrid model
assumes the cache can be cut at any token. That holds for KV, which is stored in small pages. such as Qwen3.5, Kimi K3 and Nemotron 3 also carry a fixed-size recurrent state per sequence, which can only be resumed from a snapshot taken at a block boundary - so a request that diverges between snapshots recomputes from the last one. One snapshot weighs as much as thousands of tokens of KV, and it is kept in fp32.
Prefix caching a hybrid model - KV pages vs state snapshots
Two requests share a 6,144-token prompt, then the second one diverges. How much of the shared part can it skip?
Hybrid model
69 KDA : 24 MLA layers
State snapshot every
768 tokens is the LMCache Kimi K3 recipe
Full-attention model
reuse 3,696 · recompute 4
KV cached in 16-token pages - reuse runs to the last page before the split
Hybrid - Kimi K3
reuse 3,072 · recompute 628
Recurrent state resumes only from a snapshot - everything after it replays
One state snapshot
414 MiB
per sequence, kept in fp32 - about 15.7K tokens of this model's KV (27 KiB/token)
Attention KV for the prompt
162 MiB
6,144 tokens x 27 KiB, sliceable at any page
Keeping every snapshot
3.23 GiB
8 snapshots x 414 MiB - finer snapshots buy more reuse with more storage
Engine support
- vLLM - prefix caching for hybrids through
mamba_cache_mode="align", which snapshots state on block boundaries; still experimental. - SGLang - MambaRadixCache keeps states and KV pages in separate LRU pools, with extra state slots per request for snapshots.
- LMCache - the Kimi K3 recipe stores the KDA state as an opaque page, snapshotted every 768 tokens.
Takeaway: a tiered KV store serving hybrid models has to hold whole state blobs - hundreds of MiB each, all or nothing - next to the usual token pages. Prompt length and page size are illustrative.
Throughput figures come from each project's paper or benchmarks and depend on the workload. See Inference & Serving for prefill/decode and disaggregated serving.
How engines manage KV cache
Step through the first two ideas above. Paging decides how KV is laid out in HBM; prefix reuse decides how much of it has to be computed at all.
PagedAttention - KV in fixed-size blocks, like OS pages
The same 48 slots of KV memory under two allocators. Reserving a contiguous max-length region per request leaves much of it empty. PagedAttention treats KV as non-contiguous fixed-size blocks allocated on demand, like OS virtual memory pages, so fragmentation disappears and more sequences fit in HBM.
Contiguous, reserved up front
3 of 4 requests fitOne 16-slot region per request, sized for the longest it might get.
- Tokens stored
- 9
- Reserved but empty
- 39 slots
Paged, allocated on demand
3 of 4 requests fit12 blocks of 4 slots in one shared pool; a block table maps each request.
Block tables
- Tokens stored
- 9
- Allocated but empty
- 3 slots
36 slots still free for more requests.
Waste in the paged pool is at most one part-filled block per request, while contiguous reservations sit mostly empty. That reclaimed memory is what lets the engine batch more sequences at once.
Illustrative sizes: 48 slots, 4-token blocks, 16-token reservations.
Prefix caching and RadixAttention - reuse the KV prompts share
Four prompts arrive one after another. All start with the same system prompt, and two pairs also share a document. Compare how much gets prefilled with no reuse, with one cached prefix, and with a radix tree.
A radix tree of cached prefixes, so partially overlapping prompts share whatever KV they have in common automatically.
Prompts, in arrival order
Radix tree of cached prefixes
Prefilled
0
tokens
Reused
0
tokens
Work saved
0%
of prefill
Over all four prompts: no reuse prefills 7,000 tokens, caching the system prompt 5,800, and the radix tree 3,400 - it also catches the shared documents without anyone declaring them.
Illustrative token counts: system prompt 400, each document 1,200, each question 150 (question bars drawn wider than scale so they stay readable).
KV cache acceleration projects
Beyond the engines, a layer of projects pushes KV reuse and offload further - especially for long context and RAG. For how LMCache, KVBM and NIXL fit together on shared storage, see Context Memory.
Where each project moves the KV cache
Three projects stretch the cache down the memory hierarchy; CacheBlend widens which cached chunks can be reused at all. Tap a project to read it.
the small subset of tokens CacheBlend recomputes to reuse a chunk
LMCache
Adds: Multi-tier KV layer: offloads KV to CPU DRAM / disk and shares it across instances.
Helps when: Long contexts and repeated prefixes that overflow HBM - 3–10× latency cut with vLLM.
Multipliers come from each project's paper or benchmarks and depend on the hardware and workload.
Quantizing the cache itself
The KV cache is just numbers - and like weights, it can be stored at lower precision. Dropping KV from (2 bytes) to (1 byte) halves the cache footprint and the bandwidth read per decode step, with minimal quality loss on most workloads. It composes with every attention variant and every tiering trick above.
Same cache, half the bytes
Eight cached K/V values, stored at two precisions.
Bytes in memory
8 values × 1 B = 8 Bdashed cells are freed
per request, every layer
HBM bandwidth spent
Takeaway: halving the bytes per value halves both bars - the cache needs half the memory, and every decode step streams half as much from HBM.
Toggle KV precision in the planner above between BF16 and FP8 and watch the spill shrink: halving the bytes per token halves the total cache, which can pull a scenario back up a whole tier.
FP8 is the production default. Blackwell adds KV - 4-bit values with an FP8 scale per 16 - at about 56% of FP8's footprint. In benchmarks on Blackwell (Sep 2026) it raised decode throughput by 26-30% at the same concurrency and by 37-78% when memory was the limit, with prefill unchanged; accuracy held on Qwen3.5-397B but Qwen3.8-27B dropped 1-1.6 points on reasoning and coding tests, so validate it on your own workload. DeepSeek V4.1-Flash goes further and stores its KV in FP4 natively.
See Quantization & Precision for how FP8/FP4 formats actually work.
GPU HBM at a glance
HBM capacity and bandwidth are the hard ceiling on how much KV stays hot. This is the fleet the planner sizes against.
| Node | HBM / GPU | GPUs | Total HBM | Bandwidth | NVLink |
|---|---|---|---|---|---|
| 8× H100 (DGX/HGX) | 80 GB | 8 | 640 GB | ~3.35 TB/s | 900 GB/s |
| 8× H200 | 141 GB | 8 | 1,128 GB | ~4.8 TB/s | 900 GB/s |
| 8× B200 | 180 GB | 8 | 1,440 GB | ~8 TB/s | 1800 GB/s |
| 8× B300 | 288 GB | 8 | 2,304 GB | ~8 TB/s | 1800 GB/s |
| GB200 NVL72 | 186 GB | 72 | 13,392 GB | ~8 TB/s | 1800 GB/s |
| GB300 NVL72 | 288 GB | 72 | 20,736 GB | ~8 TB/s | 1800 GB/s |
| Vera Rubin NVL72 | 288 GB | 72 | 20,736 GB | ~19.2 TB/s | 3000 GB/s |
Specs are approximate and vendor-published. NVLink figures are per-GPU intra-domain bandwidth; NVL72 systems pool all 72 GPUs into a single domain.
The other way out: have no KV cache at all
Everything on this page tackles a KV cache that grows linearly with context. State-space models (Mamba) sidestep it entirely: instead of caching every past token's keys and values, they carry one fixed-size recurrent state that's updated token-by-token - so memory stays constant no matter how long the context grows. The tradeoff is lossy recall, which is why frontier models go hybrid (mostly SSM layers, a few attention layers).
When the cache outgrows the cluster: pooled, persistent KV
At 100K concurrent users with 64K-token contexts and multi-day retention, the KV cache reaches tens of petabytes - far beyond any local NVMe. A pooled, networked KV tier (VAST, NVIDIA CMX, Mooncake) makes that practical: KV persists across requests and nodes, returning users skip prefill entirely, and data reduction shrinks the footprint further. It turns the KV cache from a per-GPU scratchpad into shared cluster infrastructure.