Memory & Efficiency

KV Cache

Fundamental

The KV cache is what makes token generation fast - and what makes inference memory-bound. This page builds from the core idea to how production systems tier it across HBM, DRAM, NVMe, and pooled storage.

Prerequisite: this page assumes you know what keys and values are. If not, read Attention Mechanisms (and Tokenization) first.

Why a cache at all?

Generating each new token requires attending to every previous token. Without a cache, the model recomputes the keys and values for the entire sequence at every single step - an O(n²) explosion. The KV cache saves those keys and values once and reuses them, turning generation from quadratic to linear. The price is memory: that cache grows with every token, every layer, every concurrent user. Below you can watch the difference, then size it for real models.

Models have no memory

The reason a cache is even possible starts here: a is stateless. It keeps nothing between steps - feed it the same tokens and it does the identical work. To write token N it must attend to all N−1 prior tokens, so without saving anything it would recompute every token's keys and values from scratch, every step. The KV cache is the workaround.

A language model has no memory

Each call is a pure function: the full token sequence goes in, one next token comes out, and nothing is kept for the next call. So every call is fed - and processes - the whole prefix again.

1/5Call 1: f(1 token) → "cat"

Input this call

The×1

model

f(tokens)

kept for next call

nothing

Next token

cat

The first call sees one token. Everything it works out about "The" is thrown away when it returns.

Because the model is stateless, the keys and values for past tokens are identical at every step - so there is no reason to recompute them. The KV cache simply stores each token's K/V once and reuses it. That is the entire trick, and the rest of this page is about the price it charges: memory.

What exactly gets cached

Self-attention is why keys and values are the things worth keeping. Each new token forms a query and scores it against every prior token's key, then blends their values. Crucially, a past token's K and V never change - so they can be computed once and reused forever.

Self-attention: score, softmax, blend

The newest token forms a Query, scores it against every token's Key (Q·K / √d), turns the scores into weights with softmax, then mixes their Values by those weights. All numbers are toy 4-d vectors, for illustration only.

1/3New token "keys" attends over 3 tokens
New tokenkeys
Q[1.2, 0.1, −0.3, −0.5]
toy vectors · d = 4
Attention of the new token keys over each token: key, score, softmax weight and value
TokenKSoftmax weightV

Cache

cached · frozen

0.49

score 0.80

the

cached · frozen

0.14

score −0.41

keys

new · computed now

0.37

score 0.55

softmax weights, stackedsum = 1.00
Cache
the
keys

Strongest link: "Cache", 2 tokens back - not the neighbor "the". Weights follow what the Q and K vectors encode, not how close the tokens are.

Blend: output = Σ weight × V(bars zoomed in)

×0.49
+
×0.14
+
×0.37
=
[0.71, 0.17, 0.01, −0.07]

The blended vector is what this token carries forward. Only its Query - and its own K and V - were new this step; every other K and V row was read from the cache unchanged.

attn(Q, K, V) = softmax(Q·Kᵀ / √d) · V

A past token's K and V depend only on that token and the weights - never on what comes after. They are frozen the moment the token is processed, so they can be computed once and cached forever. That is precisely what the KV cache stores.

Quadratic work, or linear?

Skipping the recompute turns generation from O(n²) into O(n). Step through token generation and watch the work tally for each mode - by a thousand tokens the no-cache path is doing hundreds of times the attention work.

The work triangle - recompute vs cache

Each row is one decode step; each square is one prefix token whose keys and values that step needs. Without a cache the whole row is recomputed, so the total work is the area of the triangle. With a cache only the new token (the diagonal) is computed and the rest is read back.

1/16Step 1: 1 vs 1 computed

Step 1: no cache computed 1 tokens this step, cached computed 1. Totals so far 1 vs 1.

No cache

O(n²)

This step: 1 computed

prefix token →1481216computed so far1at step 16: 136

With KV cache

O(n)

This step: 1 computed

prefix token →1481216computed so far1+ 0 readat step 16: 16

Cumulative work

fixed axes

No cache has done 1.0× the work

total work (units)040801200481216tokens generated →11no cachewith cache
computed (recompute)computed (new token only)read from cachenot generated yet

The gap widens with every token. At a 1,000-token context the no-cache path does about 501× the attention work of the cached path - which is why no production decoder runs without a KV cache. The cost moves from compute to memory.

The KV cache doesn't fit - so it tiers

is fast but tiny. At scale the KV cache dwarfs it, so serving stacks spill cold KV down a hierarchy - , local , then pooled network storage. Each step down is ~10–100× more capacity but ~10× slower. The art is keeping hot KV in HBM and demoting the rest rather than throwing it away and recomputing.

KV cache hierarchy - demote, then promote back

Each lettered block is one user's KV context. New requests land in HBM; when a tier is full the least-recently-used block is demoted one tier down instead of deleted. Click any block to make that user return.

1/9Empty cache - press play, or add a request

The cache is empty.

GPU HBMHOTsub-µs0/3
CPU DRAMWARM1–10 µs0/4
Local NVMeCOLD100 µs–1 ms0/6
NetworkPERSISTENTms0/8
Evictednothing yet - every context is still in some tier

Empty hierarchy

Every tier has free slots. Press play to watch traffic arrive, or add a request yourself.

hits 0recomputes 0
one block per user - it keeps its letter and color as it moves empty slot hit: read back, promoted miss: prefill recomputed

Demotion instead of eviction is the whole reason tiered KV systems exist: a returning user whose context is still in DRAM, NVMe or pooled storage is read back and re-promoted rather than recomputed from scratch. Turn off local NVMe to send cold blocks straight to the pooled network tier.

Size it: the tiering & capacity planner

Size a deployment the way a KV cache builds up: a real 2026 model sets the bytes per token (dense vs vs sparse/ changes it dramatically), the workload mix sets tokens per session, and users plus retention set how many sessions you hold. Then pick a GPU fleet, split it into prefill and decode, and watch the KV fill decode HBM and CPU DRAM before the rest lands on the shared tier.

KV tiering & capacity planner

Five steps in the order a KV cache is made: the model sets bytes per token, the workload sets tokens per session, users and retention set how many sessions you hold, and the fleet decides where it all lands - HBM, then CPU DRAM and finally the shared tier.

1Model

KV bytes / token34.3 KB

MLA latent (512 + 64 RoPE) per layer - tiny cache for a 671B model.

KV precision (defaults to how the model is usually served)

2Workload

What a session holds - tokens per session set the KV per session.

Chat50%

conversation history

RAG30%

question + retrieved documents

Agentic / coding20%

long tool-use sessions

capped at the model's 160K

Images0%

prompt + images

DeepSeek R1 doesn't take images input

Video0%

prompt + video

DeepSeek R1 doesn't take video input

Audio0%

prompt + speech

DeepSeek R1 doesn't take audio input

Average session: 59K tokens = 2.1 GB of KV on DeepSeek R1. Per session: Chat 589 MB · RAG 2.3 GB · Agentic / coding 5.8 GB

3Users and retention

How many sessions' KV you hold at once.

Live concurrent sessionssessions being served right now - their KV must be in GPU memory
Keep idle sessions' KV for (retention)how long a returning user's context can be reloaded instead of recomputedReal time

Real time: a session's KV is dropped when it goes idle, so a returning user's context is recomputed from scratch. Pick a retention window to keep it - idle KV moves down to CPU DRAM and the shared tier, where it can be reloaded instead.

4Hardware

Decode nodes hold the session KV; prefill nodes process prompts and hand the KV over.

Hopper baseline - 80GB HBM3 per GPU, 900 GB/s NVLink.

GPU nodesservers, or racks for NVL72 systems
Prefill nodesshare of nodes that only process prompts (disaggregated serving)none

Aggregated serving: all 4 nodes run prefill and decode and hold session KV.

CPU DRAM reserved for KVof 2.00 TB host memory per node%

5Where the KV lands

Model weights

671.0 GB

per copy · 2 decode copies, 2 nodes each

KV per session

2.1 GB

59K tokens on average

Total KV demand

2.08 TB

1,000 live sessions, real time

On the shared tier

-

what's left after HBM and DRAM

Live sessions need 2.08 TB of KV in decode HBM, but only 962.0 GB is free after the weights - about 451 of 1,000 sessions can decode at once. The rest wait, or are swapped out and back in.

Where the KV cache lands

demand 2.08 TB

GPU HBM 962.0 GB · 45%CPU DRAM 1.14 TB · 55%Shared tier -

Each tier across the 4 decode nodes, drawn to its own capacity

GPU HBMFULL2.50 TB / 2.50 TB
weights, one copy per 2 nodes (1.31 TB) ~10% activation headroom KV gets the rest (962.0 GB)
CPU DRAM1.14 TB / 4.00 TB · 29%

50% of 2.00 TB host memory per node

Shared tier (pooled network storage)nothing lands here

Pooled storage every node can read, so a session's KV can be reloaded on any node. No fixed capacity - drawn open-ended. Light: what spilled; solid: what is stored after VAST's ~1.2:1 data reduction on KV data.

HBM fills - 1.14 TB spills to CPU DRAM

45% of the KV cache fits in HBM after the weights; the other 55% lives in DRAM, one hop away. Nothing reaches the shared tier.

Each decode node group holds one copy of the weights plus ~10% activation headroom; the KV cache gets the rest of HBM. Live sessions are placed first, then idle sessions kept for retention, then burst headroom. The shared tier is pooled network storage (VAST, NVIDIA CMX and similar) that every node can read, so a session can be reloaded on any node. Illustrative, not a benchmark.

Attention decides the bill: KV per session by model

KV cache for one session on every model in the planner, each at its usual KV precision. The largest needs 104× the smallest at the same context.

Session length

Mistral Large47.2 GB
Mistral Medium 3.547.2 GB
Llama 3.3-70B42.9 GB
Seed-OSS-36B34.4 GB
Qwen3-Coder-480B33.3 GB
Qwen3-VL-235B25.2 GB
MiniMax-M318.0 GB
GLM5.1-744B14.4 GB
Qwen3-32B40K max10.7 GB
Gemma 4 31B10.7 GB
Mistral Large 39.2 GB
Kimi K2 (1T)9.2 GB
Llama 4 Maverick6.4 GB
GLM5.3-744B6.2 GB
Qwen3.8-2.4T6.2 GB
DeepSeek R14.6 GB
Qwen3.8-27B4.3 GB
GPT-OSS-120B2.4 GB
Kimi K3 (2.8T)1.8 GB
Nemotron 3 Ultra1.6 GB
GLM5.3-Flash1.6 GB
DeepSeek V4-Pro655 MB
Nemotron 3 Super550 MB
DeepSeek V4-Flash463 MB
Falcon Mambafixed state

Log scale. Models with a shorter maximum context are shown at that maximum; pure state-space models keep a fixed-size state instead of a per-token KV cache.

Software that stretches the cache

Before adding hardware, modern serving engines wring far more out of the HBM you already have. These four techniques are the backbone of every production LLM stack.

PagedAttention

2–4× throughput

vLLM

Stores KV in fixed-size pages (like OS virtual memory) instead of one contiguous block per request. Eliminates the 60–80% fragmentation waste of pre-allocating max-length buffers, so far more requests fit in HBM at once.

RadixAttention

5–6× on shared prefixes

SGLang

Organizes cached prefixes in a radix tree so requests that share a system prompt or few-shot examples reuse the same KV blocks. Huge for chat and agent workloads where every request starts with the same instructions.

Continuous batching

Higher GPU utilization

Orca / vLLM

Instead of waiting for a whole batch to finish, it swaps completed sequences out and new ones in at every decode step. Keeps the GPU saturated when requests have wildly different output lengths.

12 requests finished in 12 steps, 1 idle slot-step

A static batch finishes 7 in the same time (illustrative lengths).

Prefix / KV reuse

TTFT 11s → 1.5s

LMCache · Mooncake

“Prefill once, reuse everywhere.” Cache the KV of a long document or shared context and serve it to many requests from a fast tier instead of recomputing prefill each time. In benchmarks, VAST + LMCache cut time-to-first-token ~7× on 128K contexts.

Time to first token, 128K-token context

Recompute prefill11 s
Load cached KV from a fast tier1.5 s

The shared context's KV is already stored, so the request loads it instead of recomputing - ~7× faster to first token. (from a VAST + LMCache benchmark)

Prefix caching a hybrid model

assumes the cache can be cut at any token. That holds for KV, which is stored in small pages. such as Qwen3.5, Kimi K3 and Nemotron 3 also carry a fixed-size recurrent state per sequence, which can only be resumed from a snapshot taken at a block boundary - so a request that diverges between snapshots recomputes from the last one. One snapshot weighs as much as thousands of tokens of KV, and it is kept in fp32.

Prefix caching a hybrid model - KV pages vs state snapshots

Two requests share a 6,144-token prompt, then the second one diverges. How much of the shared part can it skip?

Hybrid model

69 KDA : 24 MLA layers

State snapshot every

768 tokens is the LMCache Kimi K3 recipe

Full-attention model

reuse 3,696 · recompute 4

KV cached in 16-token pages - reuse runs to the last page before the split

Hybrid - Kimi K3

reuse 3,072 · recompute 628

Recurrent state resumes only from a snapshot - everything after it replays

reused from cacheshared, but recomputedstate snapshotnew tokens (computed either way)

One state snapshot

414 MiB

per sequence, kept in fp32 - about 15.7K tokens of this model's KV (27 KiB/token)

Attention KV for the prompt

162 MiB

6,144 tokens x 27 KiB, sliceable at any page

Keeping every snapshot

3.23 GiB

8 snapshots x 414 MiB - finer snapshots buy more reuse with more storage

Engine support

  • vLLM - prefix caching for hybrids through mamba_cache_mode="align", which snapshots state on block boundaries; still experimental.
  • SGLang - MambaRadixCache keeps states and KV pages in separate LRU pools, with extra state slots per request for snapshots.
  • LMCache - the Kimi K3 recipe stores the KDA state as an opaque page, snapshotted every 768 tokens.

Takeaway: a tiered KV store serving hybrid models has to hold whole state blobs - hundreds of MiB each, all or nothing - next to the usual token pages. Prompt length and page size are illustrative.

Throughput figures come from each project's paper or benchmarks and depend on the workload. See Inference & Serving for prefill/decode and disaggregated serving.

How engines manage KV cache

Step through the first two ideas above. Paging decides how KV is laid out in HBM; prefix reuse decides how much of it has to be computed at all.

PagedAttention - KV in fixed-size blocks, like OS pages

The same 48 slots of KV memory under two allocators. Reserving a contiguous max-length region per request leaves much of it empty. PagedAttention treats KV as non-contiguous fixed-size blocks allocated on demand, like OS virtual memory pages, so fragmentation disappears and more sequences fit in HBM.

1/9Three requests start. Contiguous reserves 16 slots each; paged takes one block per 4 tokens.
Contiguous, reserved up front
3 of 4 requests fit

One 16-slot region per request, sized for the longest it might get.

R1
R2
R3
R4not arrived yet
Tokens stored
9
Reserved but empty
39 slots
Paged, allocated on demand
3 of 4 requests fit

12 blocks of 4 slots in one shared pool; a block table maps each request.

#0R1
#1
#2R3
#3
#4
#5R2
#6
#7
#8
#9
#10
#11

Block tables

R1→#0
R2→#5
R3→#2
R4→none yet
Tokens stored
9
Allocated but empty
3 slots

36 slots still free for more requests.

KV of a stored token (color = request)Reserved or allocated, still empty

Waste in the paged pool is at most one part-filled block per request, while contiguous reservations sit mostly empty. That reclaimed memory is what lets the engine batch more sequences at once.

Illustrative sizes: 48 slots, 4-token blocks, 16-token reservations.

Prefix caching and RadixAttention - reuse the KV prompts share

Four prompts arrive one after another. All start with the same system prompt, and two pairs also share a document. Compare how much gets prefilled with no reuse, with one cached prefix, and with a radix tree.

A radix tree of cached prefixes, so partially overlapping prompts share whatever KV they have in common automatically.

1/5Four prompts share a system prompt; two pairs also share a document.

Prompts, in arrival order

R1
systemdoc Aq1
R2
systemdoc Aq2
R3
systemdoc Bq3
R4
systemdoc Bq4
Prefilled (color = content)Reused from cache

Radix tree of cached prefixes

system: 400 tokens - not cachedsystem400doc A: 1200 tokens - not cacheddoc A1.2kdoc B: 1200 tokens - not cacheddoc B1.2kq1: 150 tokens - not cachedq1150q2: 150 tokens - not cachedq2150q3: 150 tokens - not cachedq3150q4: 150 tokens - not cachedq4150
Hit for this promptJust addedNot cached

Prefilled

0

tokens

Reused

0

tokens

Work saved

0%

of prefill

Over all four prompts: no reuse prefills 7,000 tokens, caching the system prompt 5,800, and the radix tree 3,400 - it also catches the shared documents without anyone declaring them.

Illustrative token counts: system prompt 400, each document 1,200, each question 150 (question bars drawn wider than scale so they stay readable).

KV cache acceleration projects

Beyond the engines, a layer of projects pushes KV reuse and offload further - especially for long context and RAG. For how LMCache, KVBM and NIXL fit together on shared storage, see Context Memory.

Where each project moves the KV cache

Three projects stretch the cache down the memory hierarchy; CacheBlend widens which cached chunks can be reused at all. Tap a project to read it.

GPU HBMCPU DRAMSSD / diskRemote pool

the small subset of tokens CacheBlend recomputes to reuse a chunk

LMCache

Adds: Multi-tier KV layer: offloads KV to CPU DRAM / disk and shares it across instances.

Helps when: Long contexts and repeated prefixes that overflow HBM - 3–10× latency cut with vLLM.

Multipliers come from each project's paper or benchmarks and depend on the hardware and workload.

Quantizing the cache itself

The KV cache is just numbers - and like weights, it can be stored at lower precision. Dropping KV from (2 bytes) to (1 byte) halves the cache footprint and the bandwidth read per decode step, with minimal quality loss on most workloads. It composes with every attention variant and every tiering trick above.

Same cache, half the bytes

Eight cached K/V values, stored at two precisions.

Bytes in memory

8 values × 1 B = 8 Bdashed cells are freed

Total KV cache footprint0.5×

per request, every layer

Bytes read per decode step0.5×

HBM bandwidth spent

Takeaway: halving the bytes per value halves both bars - the cache needs half the memory, and every decode step streams half as much from HBM.

Toggle KV precision in the planner above between BF16 and FP8 and watch the spill shrink: halving the bytes per token halves the total cache, which can pull a scenario back up a whole tier.

FP8 is the production default. Blackwell adds KV - 4-bit values with an FP8 scale per 16 - at about 56% of FP8's footprint. In benchmarks on Blackwell (Sep 2026) it raised decode throughput by 26-30% at the same concurrency and by 37-78% when memory was the limit, with prefill unchanged; accuracy held on Qwen3.5-397B but Qwen3.8-27B dropped 1-1.6 points on reasoning and coding tests, so validate it on your own workload. DeepSeek V4.1-Flash goes further and stores its KV in FP4 natively.

See Quantization & Precision for how FP8/FP4 formats actually work.

GPU HBM at a glance

HBM capacity and bandwidth are the hard ceiling on how much KV stays hot. This is the fleet the planner sizes against.

NodeHBM / GPUGPUsTotal HBMBandwidthNVLink
8× H100 (DGX/HGX)80 GB8640 GB~3.35 TB/s900 GB/s
8× H200141 GB81,128 GB~4.8 TB/s900 GB/s
8× B200180 GB81,440 GB~8 TB/s1800 GB/s
8× B300288 GB82,304 GB~8 TB/s1800 GB/s
GB200 NVL72186 GB7213,392 GB~8 TB/s1800 GB/s
GB300 NVL72288 GB7220,736 GB~8 TB/s1800 GB/s
Vera Rubin NVL72288 GB7220,736 GB~19.2 TB/s3000 GB/s

Specs are approximate and vendor-published. NVLink figures are per-GPU intra-domain bandwidth; NVL72 systems pool all 72 GPUs into a single domain.

The other way out: have no KV cache at all

Everything on this page tackles a KV cache that grows linearly with context. State-space models (Mamba) sidestep it entirely: instead of caching every past token's keys and values, they carry one fixed-size recurrent state that's updated token-by-token - so memory stays constant no matter how long the context grows. The tradeoff is lossy recall, which is why frontier models go hybrid (mostly SSM layers, a few attention layers).

When the cache outgrows the cluster: pooled, persistent KV

At 100K concurrent users with 64K-token contexts and multi-day retention, the KV cache reaches tens of petabytes - far beyond any local NVMe. A pooled, networked KV tier (VAST, NVIDIA CMX, Mooncake) makes that practical: KV persists across requests and nodes, returning users skip prefill entirely, and data reduction shrinks the footprint further. It turns the KV cache from a per-GPU scratchpad into shared cluster infrastructure.