VAST Data

Context Memory: KV Cache on VAST

Intermediate

Agents and long sessions re-read the same context over and over. Keeping its KV cache in a shared tier - instead of recomputing it - is now a storage decision with a price tag.

Why context became memory

Every token a model reads leaves behind its - for Llama 3.3 70B in FP8, about 160 KB per token, so a 128K-token session holds roughly 21 GB. Agents and coding assistants return to the same session dozens of times, each turn adding a little and re-reading everything. If that KV is gone when the session comes back, the GPUs run over the whole history again before the first new token appears.

~100 : 1

input to output tokens

in a production agent averaging about 50 tool calls per task (Manus, 2025). Almost all of the work is reading context, not writing answers.

>96%

prefix-cache hit rate

on agentic coding traffic, with a median of 43 turns and about 142K input tokens per turn (vLLM AgentX, 2026).

56%

of input tokens served from cache

across DeepSeek's published day of serving, 608B input tokens, with the cache kept on disk rather than in GPU memory.

The sessions themselves have grown. A year ago 8K-30K tokens counted as long context; agentic and coding sessions now mostly run from 100K to 500K tokens - 10 GB or more of KV each. They also run five to ten times as many turns as a chat, and a coding agent often fans out into five to ten subagents at once, all re-reading the same history. A miss on any of those turns means recomputing hundreds of thousands of tokens.

GPU memory can't hold thousands of 20 GB sessions, so the KV has to live somewhere else between turns. The question this page answers: is reading it back cheaper than recomputing it, and by how much? For how an agent's context grows turn by turn, see Agentic AI; for the KV math itself, the KV Cache page.

Recompute or reload: the crossover

Recomputing a prompt costs GPU time that grows faster than the prompt, because every token attends to all the tokens before it. Reloading costs storage bandwidth that grows exactly with the prompt: tokens x KV bytes per token / link speed. So there is always a length past which reloading wins - and with a fast link it wins almost everywhere.

Recompute or reload?

Time to rebuild a prompt's KV cache on the GPUs, against time to read it back from storage. Slow the link to 1-2 GB/s and watch where the lines cross.

Model

Storage link per replica25 GB/s
1 ms10 ms100 ms1 s10 s100 s1000 s1K4K16K64K256K1Mprompt length (tokens)
recompute on 1 GPU reload at 25 GB/s (160 KB of KV per token) reload wins

Reload wins from

every length

Llama 3.3 70B at 25 GB/s

128K prompt, recompute

16 s

GPU prefill

128K prompt, reload

859 ms

19.0x faster than recompute

Illustrative: prefill at 50% of dense tensor peak; reload limited only by the link. Compact KV caches (DeepSeek R1's latent attention) need far less bandwidth to win.

From hit rate to $ per million tokens

The same arithmetic as the token-cost pages - $ per GPU-hour divided by tokens per GPU-second - applied to prefill. A cache hit removes most of the GPU seconds a long prompt would take, so the cost per prompt token falls almost in step with the hit rate.

Cache-hit to $/token bridge

One long prompt arrives. Some share of it was seen before and its KV cache sits in a context-memory tier. Reload that share, compute only the new tail, and compare the cost with recomputing the whole prompt.

Model

1 GPU per replica · Llama 3.3 70B weights in FP8, FP8 KV cache

GPU cost$ per GPU-hour$2.31
Prompt lengthtokens the model must have in context32K
Cache hit rateshare of the prompt already in context memory90%
Storage link per replicasustained GB/s from the context-memory tier25 GB/s

Loading

Prefill time before the first token

Recompute the whole prompt2.4 s
Reload 29K cached, compute 3K new484 ms
GPU prefillKV reload (193 ms for 4.8 GB)

$ / M prompt tokens

$0.009

vs $0.047 recomputed

GPU time saved

80%

1.9 s per request

KV per session

4.5 GB

5.4 GB before 1.2:1 reduction

Time to first token

484 ms

vs 2.4 s recomputed

$ per million prompt tokens vs hit rate

$0.000$0.020$0.040$0.0600%25%50%75%100%share of the prompt found in context memoryrecompute every time

Illustrative model, not a benchmark: prefill at 50% of dense tensor peak (the same assumption as the Tokenomics pages), the storage link as the only limit on reload speed, and GPU cost only - storage, network and the cost of writing the KV the first time are left out. Data reduction of 1.2:1 on this FP8 KV data shrinks the capacity a tier must hold, not the bytes moved.

Public APIs pass this saving on: cached input is billed at a small fraction of fresh input. Writing to the cache can cost extra (1.25x the input price for a 5-minute cache on Anthropic and OpenAI's newest models, for example), which a single re-read usually repays.

API, $ per M tokensInputCached inputShare
OpenAI gpt-6-sol$2.00$0.2010%
Anthropic Claude Sonnet 5.5$2.00$0.2010%
Anthropic Claude Opus 5.5$4.00$0.205%
Google Gemini 3.1 Pro$2.00$0.2010%
DeepSeek V4 Pro$1.32$0.0443%

List prices as of October 2026; DeepSeek at peak hours.

Where KV lives, and what a shared tier adds

NVIDIA numbers the places KV can sit from G1 to G4. Each step down holds far more and answers more slowly. The step that changes the economics is the first one that is shared: below it, a session can only hit on the node that served it last.

G1GPU HBMhot, latency-critical KV for the tokens being generated nowshared by one GPU
G2Host memorystaging and buffering next to the GPUshared by one node
G3Local NVMewarm KV - but tied to one node, so it is harder to manageshared by one node
G3.5Context memory (CMX)Ethernet-attached flash built for KV cacheshared by a pod
G4Shared storagedurable, shared capacity: history, artifacts, cold KVshared by the data center

Explore the capacities and latencies of each tier on the NVIDIA STX & CMX page. Below, the same four sessions are served with node-local KV and with a shared tier.

Where the KV lives decides whether it can be reused

Four sessions keep coming back. The router tries to send each one home, but busy nodes, upgrades and load balancing get in the way.

1/9Each node holds the KV cache of the session it served last.

Node-local KV

each server's own DRAM and NVMe

Node 1GPUs
A
Node 2GPUs
B
Node 3GPUs
C
Node 4GPUs
D

Waiting for the first request

Shared context memory

one pool every node can read

Node 1GPUs
Node 2GPUs
Node 3GPUs
Node 4GPUs
Shared tierABCD

Waiting for the first request

Node-local hits

0 / 0

returning turns

Shared-tier hits

0 / 0

returning turns

Prefill time, local

-

summed over turns

Prefill time, shared

-

summed over turns

Illustrative routing. Each turn is a 64K-token session plus 1K new tokens for Llama 3.3 70B on one GB300 GPU: about 5.9 s to recompute, or 550 ms to reload at 25 GB/s and compute the new tokens.

A shared tier also keeps KV through node restarts, upgrades and long idle gaps while an agent waits on a tool, and lets disaggregated serving write a prompt's KV once and read it from any decode node.

How big should that tier be? Sizing comes down to five questions: KV bytes per token (model and precision), tokens per session (the mix of chat, RAG, agentic and media work), live sessions, how long idle sessions are kept and how fast new ones arrive, and headroom for bursts - then one check, the network speed that sets the reload time. The KV tiering & capacity planner answers all of them; this link opens it on a long-context chat scenario that keeps idle sessions for an hour.

How VAST plugs in

VAST sits at the bottom of the open KV stack rather than replacing any of it. The serving engine keeps its own cache in GPU memory; or 's decide what to offload and prefetch; and move the blocks; and VAST holds them in one namespace every node can read.

1. Serving enginevLLM · SGLang · TensorRT-LLMruns prefill and decode, owns the GPU-resident KV cache
2. KV managerLMCache · NVIDIA Dynamo KVBMdecides which KV blocks to keep, evict, offload and prefetch; commercial platforms such as Tensormesh build on LMCache
3. Data moverNIXL · GPUDirect Storagemoves KV blocks between GPU memory and the tiers, skipping host copies
4. ProtocolNFS over RDMA or TCP · GPUDirect Storage · S3 over TCP · S3 over RDMAfile or object access; S3 over RDMA arrives in the next VAST AI OS release
5. VAST AI OSshared, persistent KV tierone namespace every node reads, with data reduction on KV data (1.4:1 BF16, 1.2:1 FP8)
  • VAST's open-source vLLM connector, VAST Undivided Attention, has been folded into LMCache's GPUDirect Storage backend, so it works with stock LMCache.
  • Dynamo 1.0 reaches VAST through KVBM and NIXL over NFS (RDMA or TCP), GPUDirect Storage or S3 over TCP, with S3 over RDMA arriving in the next VAST AI OS release.
  • KV data reduces about 1.4:1 on VAST for a BF16 cache and 1.2:1 for FP8 (an FP4 cache is already dense), which shrinks the capacity a tier must hold for a given number of sessions.
  • The VAST AI OS also runs on BlueField-4 as part of NVIDIA's context-memory platform (G3.5), with partner systems due in the second half of 2026.

What the benchmarks show

Three published VAST benchmarks, oldest first. The pattern is the same each time: drops by roughly the share of the prompt that no longer has to be recomputed, and the reload runs as fast as the network allows.

Jul 2025

One H100, a 130K-token context

Qwen3-32B on vLLM + LMCache, 32 GB of KV over a 400 Gb/s link with GPUDirect Storage.

Time to first token fell from over 11 s to about 1.5 s; a single H100 read KV at 35 GB/s.

Dec 2025

NVIDIA Dynamo, a 127K-token prompt

Llama 3.1 405B (4-bit) on 8x H100 with Dynamo, vLLM and NIXL, over NFS multipath on 2x 100 GbE.

Time to first token fell from 62 s to 3 s - about 90% less GPU time - with the link running at 181 Gb/s, over 90% of line rate.

Jun 2026

Multi-turn coding sessions

Mistral Medium 3.5 (128B) on 8x H100, ten 140K-token repository contexts revisited over five turns each.

Warm turns sped up from about 22 s to 6.6 s and the whole run finished 1.96x sooner. Decode speed did not change, and the cold first turn was about 4 s slower because it also writes its KV out.

The formula on this page predicts the December result: 127K tokens of 405B-model KV in BF16 is about 66 GB, and 66 GB over a link moving about 22 GB/s takes about 3 s. In other words, that run was limited by its two 100 GbE ports - a faster link would have cut the time further.

Running it in production

A KV cache looks like a simple read-and-write cache, but once it is shared by a whole cluster it becomes a data-management problem. Storage is about 1,000x slower than HBM and about 1,000x cheaper per GB, so the system has to decide carefully what to keep, where, and for how long.

Admission: what to write at all

Not every KV block is worth storing. Admission policies skip short or one-off requests, which saves capacity and flash endurance - every write counts against the drives' rated writes per day.

Placement across tiers

Hot KV stays in HBM or host memory, warmer and larger KV moves down to flash. Popular prefixes - a shared system prompt, a team's code base - can be pinned so they are never evicted.

Many nodes, one copy

When two nodes produce the same prefix, the tier should keep one copy, avoid write conflicts, and expire old sessions with one global eviction policy rather than per-node guesses.

Tenants

A shared cluster serves many teams or customers. Each needs its KV isolated and encrypted, with its own capacity budget and evictions that never touch another tenant.

Restarts and failures

KV kept outside the serving engine survives a GPU failure, an engine restart or a rolling upgrade, so sessions resume by reloading instead of recomputing from scratch.

Retention and deletion

KV is derived from user prompts, so it can hold personal data. Keep it only as long as sessions need it, and delete it on request - the same rules that apply to the conversation itself.

Tenant isolation, quotas and encryption on a shared VAST cluster are covered on the Secure Multitenancy page.

When it pays, and when it doesn't

Context memory pays off for long, revisited contexts: agents, coding assistants, multi-turn chat, and long shared documents read by many requests. It does little in four cases:

Short prompts

Below a few thousand tokens, recompute takes milliseconds and the KV usually still sits in HBM or host memory anyway. KV managers decide per request: very short prompts are simply recomputed, and KV that is unlikely to be reused is never written out.

Little reuse

In a public Kimi serving trace, over half of cached blocks were never read again. A tier only pays for the hits it adds beyond HBM and DRAM.

A tier smaller than the working set

KV that is evicted before the session returns is just extra writes. Size for sessions x context x retention time.

The first turn

Nothing is cached yet, and writing the new KV out adds a little time. The payoff starts on the second turn.

Related: InsightEngine for the wider VAST + NVIDIA inference stack, NVIDIA STX & CMX for the G3.5 tier, and Why Output Costs More for why input tokens are cheap to begin with.

Key takeaways

In one line

Agents and long sessions re-read the same context over and over; keeping its KV cache in a shared tier and reloading it is faster and cheaper than recomputing it on the GPUs.

Key points

  • Llama 3.3 70B in FP8 leaves about 160 KB of KV cache per token - roughly 21 GB for a 128K-token session - far more than GPU memory can hold for thousands of sessions.
  • Recompute time grows faster than the prompt while reload time grows with it, so past some length reloading always wins; with a fast link it wins almost everywhere.
  • A shared tier lets any node reuse a session's KV, so busy nodes, upgrades and load balancing no longer turn returning turns into recomputes.
  • With NVIDIA Dynamo, time to first token for a 127K-token prompt fell from 62 s to 3 s in benchmarks (Dec 2025), limited by its two 100 GbE ports.

Questions to explore

  1. 01What share of your input tokens repeat across turns or requests today?
  2. 02How long do your agent sessions sit idle between turns, and does their KV survive the gap?
  3. 03What storage bandwidth reaches each serving replica, and at what prompt length does reloading beat recomputing?

Common questions

How does VAST connect to vLLM, SGLang or Dynamo?
Through the open KV stack: LMCache or Dynamo's KVBM decide what to offload and prefetch, NIXL and GPUDirect Storage move the blocks, and VAST holds them in one namespace every node can read.
How much capacity does a context-memory tier need?
Sessions x context x KV bytes per token, kept for as long as sessions come back. KV data reduces about 1.4:1 on VAST for a BF16 cache and 1.2:1 for FP8, which shrinks the capacity needed.
When does context memory not help?
Short prompts, contexts that are rarely reused, a tier smaller than the working set, and the first turn of a session, which has nothing to reuse yet.