VAST Data
Context Memory: KV Cache on VAST
IntermediateAgents and long sessions re-read the same context over and over. Keeping its KV cache in a shared tier - instead of recomputing it - is now a storage decision with a price tag.
Why context became memory
Every token a model reads leaves behind its - for Llama 3.3 70B in FP8, about 160 KB per token, so a 128K-token session holds roughly 21 GB. Agents and coding assistants return to the same session dozens of times, each turn adding a little and re-reading everything. If that KV is gone when the session comes back, the GPUs run over the whole history again before the first new token appears.
~100 : 1
input to output tokens
in a production agent averaging about 50 tool calls per task (Manus, 2025). Almost all of the work is reading context, not writing answers.
>96%
prefix-cache hit rate
on agentic coding traffic, with a median of 43 turns and about 142K input tokens per turn (vLLM AgentX, 2026).
56%
of input tokens served from cache
across DeepSeek's published day of serving, 608B input tokens, with the cache kept on disk rather than in GPU memory.
The sessions themselves have grown. A year ago 8K-30K tokens counted as long context; agentic and coding sessions now mostly run from 100K to 500K tokens - 10 GB or more of KV each. They also run five to ten times as many turns as a chat, and a coding agent often fans out into five to ten subagents at once, all re-reading the same history. A miss on any of those turns means recomputing hundreds of thousands of tokens.
GPU memory can't hold thousands of 20 GB sessions, so the KV has to live somewhere else between turns. The question this page answers: is reading it back cheaper than recomputing it, and by how much? For how an agent's context grows turn by turn, see Agentic AI; for the KV math itself, the KV Cache page.
Recompute or reload: the crossover
Recomputing a prompt costs GPU time that grows faster than the prompt, because every token attends to all the tokens before it. Reloading costs storage bandwidth that grows exactly with the prompt: tokens x KV bytes per token / link speed. So there is always a length past which reloading wins - and with a fast link it wins almost everywhere.
Recompute or reload?
Time to rebuild a prompt's KV cache on the GPUs, against time to read it back from storage. Slow the link to 1-2 GB/s and watch where the lines cross.
Model
Reload wins from
every length
Llama 3.3 70B at 25 GB/s
128K prompt, recompute
16 s
GPU prefill
128K prompt, reload
859 ms
19.0x faster than recompute
Illustrative: prefill at 50% of dense tensor peak; reload limited only by the link. Compact KV caches (DeepSeek R1's latent attention) need far less bandwidth to win.
From hit rate to $ per million tokens
The same arithmetic as the token-cost pages - $ per GPU-hour divided by tokens per GPU-second - applied to prefill. A cache hit removes most of the GPU seconds a long prompt would take, so the cost per prompt token falls almost in step with the hit rate.
Cache-hit to $/token bridge
One long prompt arrives. Some share of it was seen before and its KV cache sits in a context-memory tier. Reload that share, compute only the new tail, and compare the cost with recomputing the whole prompt.
Model
1 GPU per replica · Llama 3.3 70B weights in FP8, FP8 KV cache
Loading
Prefill time before the first token
$ / M prompt tokens
$0.009
vs $0.047 recomputed
GPU time saved
80%
1.9 s per request
KV per session
4.5 GB
5.4 GB before 1.2:1 reduction
Time to first token
484 ms
vs 2.4 s recomputed
$ per million prompt tokens vs hit rate
Illustrative model, not a benchmark: prefill at 50% of dense tensor peak (the same assumption as the Tokenomics pages), the storage link as the only limit on reload speed, and GPU cost only - storage, network and the cost of writing the KV the first time are left out. Data reduction of 1.2:1 on this FP8 KV data shrinks the capacity a tier must hold, not the bytes moved.
Public APIs pass this saving on: cached input is billed at a small fraction of fresh input. Writing to the cache can cost extra (1.25x the input price for a 5-minute cache on Anthropic and OpenAI's newest models, for example), which a single re-read usually repays.
| API, $ per M tokens | Input | Cached input | Share |
|---|---|---|---|
| OpenAI gpt-6-sol | $2.00 | $0.20 | 10% |
| Anthropic Claude Sonnet 5.5 | $2.00 | $0.20 | 10% |
| Anthropic Claude Opus 5.5 | $4.00 | $0.20 | 5% |
| Google Gemini 3.1 Pro | $2.00 | $0.20 | 10% |
| DeepSeek V4 Pro | $1.32 | $0.044 | 3% |
List prices as of October 2026; DeepSeek at peak hours.
How VAST plugs in
VAST sits at the bottom of the open KV stack rather than replacing any of it. The serving engine keeps its own cache in GPU memory; or 's decide what to offload and prefetch; and move the blocks; and VAST holds them in one namespace every node can read.
- VAST's open-source vLLM connector, VAST Undivided Attention, has been folded into LMCache's GPUDirect Storage backend, so it works with stock LMCache.
- Dynamo 1.0 reaches VAST through KVBM and NIXL over NFS (RDMA or TCP), GPUDirect Storage or S3 over TCP, with S3 over RDMA arriving in the next VAST AI OS release.
- KV data reduces about 1.4:1 on VAST for a BF16 cache and 1.2:1 for FP8 (an FP4 cache is already dense), which shrinks the capacity a tier must hold for a given number of sessions.
- The VAST AI OS also runs on BlueField-4 as part of NVIDIA's context-memory platform (G3.5), with partner systems due in the second half of 2026.
What the benchmarks show
Three published VAST benchmarks, oldest first. The pattern is the same each time: drops by roughly the share of the prompt that no longer has to be recomputed, and the reload runs as fast as the network allows.
Jul 2025
One H100, a 130K-token context
Qwen3-32B on vLLM + LMCache, 32 GB of KV over a 400 Gb/s link with GPUDirect Storage.
Time to first token fell from over 11 s to about 1.5 s; a single H100 read KV at 35 GB/s.
Dec 2025
NVIDIA Dynamo, a 127K-token prompt
Llama 3.1 405B (4-bit) on 8x H100 with Dynamo, vLLM and NIXL, over NFS multipath on 2x 100 GbE.
Time to first token fell from 62 s to 3 s - about 90% less GPU time - with the link running at 181 Gb/s, over 90% of line rate.
Jun 2026
Multi-turn coding sessions
Mistral Medium 3.5 (128B) on 8x H100, ten 140K-token repository contexts revisited over five turns each.
Warm turns sped up from about 22 s to 6.6 s and the whole run finished 1.96x sooner. Decode speed did not change, and the cold first turn was about 4 s slower because it also writes its KV out.
The formula on this page predicts the December result: 127K tokens of 405B-model KV in BF16 is about 66 GB, and 66 GB over a link moving about 22 GB/s takes about 3 s. In other words, that run was limited by its two 100 GbE ports - a faster link would have cut the time further.
Running it in production
A KV cache looks like a simple read-and-write cache, but once it is shared by a whole cluster it becomes a data-management problem. Storage is about 1,000x slower than HBM and about 1,000x cheaper per GB, so the system has to decide carefully what to keep, where, and for how long.
Admission: what to write at all
Not every KV block is worth storing. Admission policies skip short or one-off requests, which saves capacity and flash endurance - every write counts against the drives' rated writes per day.
Placement across tiers
Hot KV stays in HBM or host memory, warmer and larger KV moves down to flash. Popular prefixes - a shared system prompt, a team's code base - can be pinned so they are never evicted.
Many nodes, one copy
When two nodes produce the same prefix, the tier should keep one copy, avoid write conflicts, and expire old sessions with one global eviction policy rather than per-node guesses.
Tenants
A shared cluster serves many teams or customers. Each needs its KV isolated and encrypted, with its own capacity budget and evictions that never touch another tenant.
Restarts and failures
KV kept outside the serving engine survives a GPU failure, an engine restart or a rolling upgrade, so sessions resume by reloading instead of recomputing from scratch.
Retention and deletion
KV is derived from user prompts, so it can hold personal data. Keep it only as long as sessions need it, and delete it on request - the same rules that apply to the conversation itself.
Tenant isolation, quotas and encryption on a shared VAST cluster are covered on the Secure Multitenancy page.
When it pays, and when it doesn't
Context memory pays off for long, revisited contexts: agents, coding assistants, multi-turn chat, and long shared documents read by many requests. It does little in four cases:
Short prompts
Below a few thousand tokens, recompute takes milliseconds and the KV usually still sits in HBM or host memory anyway. KV managers decide per request: very short prompts are simply recomputed, and KV that is unlikely to be reused is never written out.
Little reuse
In a public Kimi serving trace, over half of cached blocks were never read again. A tier only pays for the hits it adds beyond HBM and DRAM.
A tier smaller than the working set
KV that is evicted before the session returns is just extra writes. Size for sessions x context x retention time.
The first turn
Nothing is cached yet, and writing the new KV out adds a little time. The payoff starts on the second turn.
Related: InsightEngine for the wider VAST + NVIDIA inference stack, NVIDIA STX & CMX for the G3.5 tier, and Why Output Costs More for why input tokens are cheap to begin with.
Key takeaways
In one line
Agents and long sessions re-read the same context over and over; keeping its KV cache in a shared tier and reloading it is faster and cheaper than recomputing it on the GPUs.
Key points
- Llama 3.3 70B in FP8 leaves about 160 KB of KV cache per token - roughly 21 GB for a 128K-token session - far more than GPU memory can hold for thousands of sessions.
- Recompute time grows faster than the prompt while reload time grows with it, so past some length reloading always wins; with a fast link it wins almost everywhere.
- A shared tier lets any node reuse a session's KV, so busy nodes, upgrades and load balancing no longer turn returning turns into recomputes.
- With NVIDIA Dynamo, time to first token for a 127K-token prompt fell from 62 s to 3 s in benchmarks (Dec 2025), limited by its two 100 GbE ports.
Questions to explore
- 01What share of your input tokens repeat across turns or requests today?
- 02How long do your agent sessions sit idle between turns, and does their KV survive the gap?
- 03What storage bandwidth reaches each serving replica, and at what prompt length does reloading beat recomputing?
Common questions
- How does VAST connect to vLLM, SGLang or Dynamo?
- Through the open KV stack: LMCache or Dynamo's KVBM decide what to offload and prefetch, NIXL and GPUDirect Storage move the blocks, and VAST holds them in one namespace every node can read.
- How much capacity does a context-memory tier need?
- Sessions x context x KV bytes per token, kept for as long as sessions come back. KV data reduces about 1.4:1 on VAST for a BF16 cache and 1.2:1 for FP8, which shrinks the capacity needed.
- When does context memory not help?
- Short prompts, contexts that are rarely reused, a tier smaller than the working set, and the first turn of a session, which has nothing to reuse yet.