Systems & Infrastructure

NVIDIA STX & Context Memory

Advanced

At GTC 2026 NVIDIA introduced BlueField-4 STX - a reference architecture for AI-native storage that inserts a dedicated context-memory tier between GPUs and storage. It is the rack-scale sequel to the KV cache story: caching context not inside one GPU, but across a whole pod.

When the KV cache outgrows the GPU

Long-context and agentic inference explode the size of the KV cache - the per-token attention state a model must keep. Once it overflows GPU , the choice is grim: stall, evict and recompute, or fall back to slow storage. STX's bet is to add a dedicated, networked context-memory tier close to the GPUs, accelerated by BlueField-4, so previously computed context is persisted and reused across requests and GPUs instead of recomputed.

STX is a , not a box NVIDIA sells directly - a blueprint its storage partners (VAST, DDN, Dell, HPE, IBM, NetApp and more) build around. (Context Memory Storage) is its first rack-scale implementation, with partner platforms expected in the second half of 2026.

STX, assembled - five building blocks, one context tier

How the pieces fit in a pod. Play the story to follow one KV block from a GPU down to the context tier and back up to a different GPU, or tap any block for what it does.

1/5Follow one KV block
Vera Rubinthe GPU platformSpectrum-Xthe fabricBlueField-4the processorCMXContext Memory StorageDOCA Memosthe softwareDOCA Memos · stages and reuses context blocksVera Rubin computeGPU 1GPU 2GPU 3GPU 4Spectrum-X Ethernet fabric · RDMA / RoCEBlueField-4Vera CPUConnectX-9BlueField-4Vera CPUConnectX-9BlueField-4Vera CPUConnectX-9CMX flashG3.5CMX flashG3.5CMX flashG3.5KV
1 · Vera Rubin

A Vera Rubin GPU builds KV cache as it serves a long conversation - more than its own memory can keep.

Vera Rubinthe GPU platform

The next-gen compute platform STX is built to feed - so the GPUs spend cycles generating tokens, not waiting on data.

New to the silicon itself? The NVIDIA Infrastructure page covers the BlueField-4 DPU and Spectrum-X switches in depth.

A new tier in the hierarchy: G3.5

NVIDIA frames the memory and storage stack as numbered tiers - G1 (GPU HBM), G2 (Vera/CPU ), G3 (local node ), and G4 (durable enterprise storage). STX inserts a new one between them: G3.5, Ethernet-attached flash optimized for . It is bigger and more shared than any single node's SSD, yet close enough to feed GPUs without falling to the slow durable floor. Drag the working set and toggle the tier to feel why it exists.

The storage hierarchy STX redraws - meet tier G3.5

Each tier is a step: how fast it answers (across) and how much KV cache the tiers hold together (up). The pooled cache fills them from the bottom like a liquid. Play the story, drag the blue level, or switch G3.5 (CMX) off to see the hole it fills.

1/5
1 · G1A small working set fits in GPU HBM (G1) - the fastest tier, and the smallest.
10 GB100 GB1 TB10 TB100 TB1 PB10 PB100 PB0.1 µs1 µs10 µs100 µs1 ms↑ capacity held, cumulative (log)G1 · GPU HBM~14 TB · ~0.2 µsG2 · Vera / CPU DRAM~50 TB · ~0.5 µsG3 · Local node SSD~0.5 PB · ~90 µsG3.5 · CMX - Ethernet flash~5 PB · ~130 µsNEW IN STXG4 · Durable storageEB-scale · ~2 ms*↑ moreworking set · 1.0 TBlands in G1 · ~0.2 µs
access latency (log) - slower →
1.0 TB
10 GB~10 PB

Lands in G1 GPU HBM (~0.2 µs)

Comfortably held in a fast tier. Raise the level to reach the band where CMX earns its place.

Capacities (per pod) and latencies are illustrative figures for teaching the tiering - not a datasheet. *G4 latency stands in for “fetch from durable storage or recompute the prefill.”

VAST's role: the CNode runs inside the DPU

Most partners attach storage over the network. VAST goes further: it ports its CNode - the stateless compute layer of its architecture - directly onto the BlueField-4 . That lets KV cache move zero-copy from remote flash into GPU memory, bypassing the host CPU and local SSD entirely, while VAST layers data reduction, security, and lifecycle management onto that shared context tier.

How VAST moves KV cache to the GPU - two paths

VAST ports its CNode onto the BlueField-4 DPU, so context flows straight from shared flash into GPU memory. Both paths end in the same GPU server - race one KV block down each and watch the classic path stop for the host CPU and a copy at every station.

1/8Press play to race both paths
GPU SERVERTraditional pathVAST on BlueField-4networkPCIezero-copy DMAStorage serverremote NVMeVAST DNodeshared flash (G3.5)Spectrum-XRDMA / RoCEHost CPUkernel + page cacheLocal SSDstaged to G3 firstBlueField-4DPU · virtio-fsruns VAST CNodeGPU HBMKV cache usable here123KVKV
1/8

One KV block leaves remote flash on each path at the same moment.

Traditional path

in transit
Hops
4
Copies
0
Host CPU
not yet

VAST on BlueField-4

in transit
Hops
3
Copies
0
Host CPU
bypassed
traditional path VAST path memory copy, numbered in orderThe race shows arrival order only - no timings.

Running the CNode on the DPU means the GPU server's CPU never touches the data path - VAST adds data reduction, security, and lifecycle management to that shared context tier, which is why NVIDIA frames VAST's contribution as “context reuse optimization at scale.”

Simplified for teaching. The zero-copy / virtio-fs detail is from analyst reporting (Blocks & Files) on the STX design, consistent with VAST's own context-reuse material.

Why it matters

The payoff of reusing context instead of recomputing it is more useful work from the same GPUs. The headline figures below compare STX with traditional CPU-based storage.

Context reuse vs recompute - the headline impact

One GPU, the same window of time, a stream of returning requests whose context it has seen before. One lane rebuilds that context with a full prefill; the other fetches the saved KV cache from CMX and gets straight to new tokens.

1/25Watch both lanes fill the same window

GPU timeline · schematic

Recompute
prefill it all again
prefill
R1
prefill
R2
prefill
R3
prefill
STX + CMX
fetch the saved context
R1
R2
R3
R4
R5
R6
R7
R8
time (schematic) →
Recomputetime on new tokens
STX + CMXtime on new tokens
Re-prefill a context seen beforeFetch saved KV cache from CMXNew tokens (R1, R2… = request)

Takeaway: reuse turns a long prefill into a short fetch, so the same GPU window fits more requests and spends more of its time producing tokens.

The payoff · one glyph = the traditional baseline

Token throughput

tokens / second

Traditional
1×
STX + CMX
5×

Energy efficiency

tokens / watt

Traditional
1×
STX + CMX
4×

Data ingest

pages / second

Traditional
1×
STX + CMX
2×

Up to 5× token throughput, 4× energy efficiency, and 2× faster ingest for STX vs traditional CPU-based storage architectures, announced at GTC 2026.

The timeline is schematic: block lengths show the pattern, not measured latencies, and are not derived from the multipliers above.

Where this connects

STX is the infrastructure-scale version of everything the KV cache pages teach, and the data services on that tier are exactly what VAST's InsightEngine builds on for real-time RAG. Follow the thread: