Systems & Infrastructure
NVIDIA STX & Context Memory
AdvancedAt GTC 2026 NVIDIA introduced BlueField-4 STX - a reference architecture for AI-native storage that inserts a dedicated context-memory tier between GPUs and storage. It is the rack-scale sequel to the KV cache story: caching context not inside one GPU, but across a whole pod.
When the KV cache outgrows the GPU
Long-context and agentic inference explode the size of the KV cache - the per-token attention state a model must keep. Once it overflows GPU , the choice is grim: stall, evict and recompute, or fall back to slow storage. STX's bet is to add a dedicated, networked context-memory tier close to the GPUs, accelerated by BlueField-4, so previously computed context is persisted and reused across requests and GPUs instead of recomputed.
STX is a , not a box NVIDIA sells directly - a blueprint its storage partners (VAST, DDN, Dell, HPE, IBM, NetApp and more) build around. (Context Memory Storage) is its first rack-scale implementation, with partner platforms expected in the second half of 2026.
STX, assembled - five building blocks, one context tier
How the pieces fit in a pod. Play the story to follow one KV block from a GPU down to the context tier and back up to a different GPU, or tap any block for what it does.
A Vera Rubin GPU builds KV cache as it serves a long conversation - more than its own memory can keep.
Vera Rubinthe GPU platform
The next-gen compute platform STX is built to feed - so the GPUs spend cycles generating tokens, not waiting on data.
New to the silicon itself? The NVIDIA Infrastructure page covers the BlueField-4 DPU and Spectrum-X switches in depth.
A new tier in the hierarchy: G3.5
NVIDIA frames the memory and storage stack as numbered tiers - G1 (GPU HBM), G2 (Vera/CPU ), G3 (local node ), and G4 (durable enterprise storage). STX inserts a new one between them: G3.5, Ethernet-attached flash optimized for . It is bigger and more shared than any single node's SSD, yet close enough to feed GPUs without falling to the slow durable floor. Drag the working set and toggle the tier to feel why it exists.
The storage hierarchy STX redraws - meet tier G3.5
Each tier is a step: how fast it answers (across) and how much KV cache the tiers hold together (up). The pooled cache fills them from the bottom like a liquid. Play the story, drag the blue level, or switch G3.5 (CMX) off to see the hole it fills.
Lands in G1 GPU HBM (~0.2 µs)
Comfortably held in a fast tier. Raise the level to reach the band where CMX earns its place.
Capacities (per pod) and latencies are illustrative figures for teaching the tiering - not a datasheet. *G4 latency stands in for “fetch from durable storage or recompute the prefill.”
VAST's role: the CNode runs inside the DPU
Most partners attach storage over the network. VAST goes further: it ports its CNode - the stateless compute layer of its architecture - directly onto the BlueField-4 . That lets KV cache move zero-copy from remote flash into GPU memory, bypassing the host CPU and local SSD entirely, while VAST layers data reduction, security, and lifecycle management onto that shared context tier.
How VAST moves KV cache to the GPU - two paths
VAST ports its CNode onto the BlueField-4 DPU, so context flows straight from shared flash into GPU memory. Both paths end in the same GPU server - race one KV block down each and watch the classic path stop for the host CPU and a copy at every station.
One KV block leaves remote flash on each path at the same moment.
Traditional path
in transit- Hops
- 4
- Copies
- 0
- Host CPU
- not yet
VAST on BlueField-4
in transit- Hops
- 3
- Copies
- 0
- Host CPU
- bypassed
Running the CNode on the DPU means the GPU server's CPU never touches the data path - VAST adds data reduction, security, and lifecycle management to that shared context tier, which is why NVIDIA frames VAST's contribution as “context reuse optimization at scale.”
Simplified for teaching. The zero-copy / virtio-fs detail is from analyst reporting (Blocks & Files) on the STX design, consistent with VAST's own context-reuse material.
Why it matters
The payoff of reusing context instead of recomputing it is more useful work from the same GPUs. The headline figures below compare STX with traditional CPU-based storage.
Context reuse vs recompute - the headline impact
One GPU, the same window of time, a stream of returning requests whose context it has seen before. One lane rebuilds that context with a full prefill; the other fetches the saved KV cache from CMX and gets straight to new tokens.
GPU timeline · schematic
Takeaway: reuse turns a long prefill into a short fetch, so the same GPU window fits more requests and spends more of its time producing tokens.
The payoff · one glyph = the traditional baseline
Token throughput
tokens / second
Energy efficiency
tokens / watt
Data ingest
pages / second
Up to 5× token throughput, 4× energy efficiency, and 2× faster ingest for STX vs traditional CPU-based storage architectures, announced at GTC 2026.
The timeline is schematic: block lengths show the pattern, not measured latencies, and are not derived from the multipliers above.
Where this connects
STX is the infrastructure-scale version of everything the KV cache pages teach, and the data services on that tier are exactly what VAST's InsightEngine builds on for real-time RAG. Follow the thread: