VAST Data

InsightEngine with NVIDIA

Intermediate

VAST + NVIDIA's full-stack system for real-time RAG - feeding GPUs directly and keeping them serving, not idling.

What InsightEngine is

Unveiled in October 2024 and available on NVIDIA DGX systems since GTC in March 2025, VAST InsightEngine is described as the “industry's first secure full-stack system for real-time data inferencing.” It is pre-integrated with NVIDIA AI Enterprise and ships DGX-validated, expanding to NVIDIA-Certified Systems - turning the VAST platform into a real-time engine rather than a stack of parts you wire together yourself.

One system, not a stack of parts

Compare the pre-integrated stack with wiring the same four layers together yourself, then follow how availability widens from DGX to NVIDIA-Certified Systems.

Secure, full-stack, real-time

Storage, networking, vector search, and inference microservices come pre-integrated under unified security, so retrieval is grounded in fresh, governed data the moment it is written.

GA on DGX from March 2025

  1. Mar 18, 2025

    Announced at GTC

    Described as the industry's first secure full-stack system for real-time data inferencing

  2. March 2025

    GA on DGX systems

    DGX-validated, pre-integrated with NVIDIA AI Enterprise

  3. Over time

    Expanding to NVIDIA-Certified Systems

    Broadening beyond DGX

Implements NVIDIA's AI Data Platform reference design

The system reached general availability on DGX systems starting March 2025, implementing NVIDIA's AI Data Platform reference design and expanding to NVIDIA-Certified Systems over time.

available expanding separately secured part

See the site's RAG & Vector Search, Inference & Serving, KV Cache, and NVIDIA Infrastructure pages for the building blocks InsightEngine assembles into one product.

“Certified” wording varies across announcements; presented as DGX-validated and expanding to NVIDIA-Certified Systems.

The NVIDIA hardware & software stack

InsightEngine is a layered system spanning NVIDIA silicon and software: BlueField-3 DPUs power the NVMe enclosures and run storage services as offloaded containers, Spectrum-X carries the fabric, cuVS does GPU-accelerated vector search, and microservices plus the AI-Q and VSS blueprints sit on top. Follow a write and a question through the layers, or tap a layer to see what it does.

Follow a write, then a question, through the stack

A document is embedded and indexed the moment it lands; a question moments later retrieves it on the GPU. Each layer lights as the data passes - tap any layer for what it does.

1/7
software ↑hardware ↓RDMA fabricAI-Q + VSS blueprintsAgents & video intelligenceNIM microservicesEmbedding, rerank & reasoningNVIDIA cuVSGPU-accelerated vector searchSpectrum-X networkingAI-optimized Ethernet fabricBlueField-3 DPUs + NVMeOffloaded storage serviceswrite inAI-Q agentEmbedding NIMReason NIMcuVS searchDocument
1 · Write

A document lands - over the Spectrum-X RDMA fabric into NVMe enclosures run by BlueField-3 DPUs.

write path (embed on write)question path (retrieve + answer) source data vectors

BlueField-3 DPUs + NVMe

Offloaded storage services
What it does
NVMe enclosures connected to and powered by BlueField-3 DPUs, which run VAST storage services as offloaded containers - keeping host CPUs free.
Why it matters for AI
Storage runs on the DPU, not the GPU host, so the CPU detour is removed from the data path.

Implements NVIDIA's AI Data Platform reference design, pre-integrated with NVIDIA AI Enterprise. Chip placement within each layer is schematic.

GPUDirect Storage: feeding GPUs without the CPU detour

On the legacy path, every read is copied into host system memory before a second copy lands it in GPU memory. DMAs data straight from NVMe into GPU memory over RDMA, bypassing the CPU/system-memory copy entirely via NFS-over-RDMA and NVIDIA Magnum IO (the general mechanism is explained on Networking & the Data Path). Toggle the two paths to see VAST's numbers.

Inside the server: the CPU detour vs a direct path

Data streams from the VAST NVMe pool over RDMA into a GPU server. Pipe width is drawn to scale with the DGX-2 read throughput measured in benchmarks.

GPU server (DGX-2 in the benchmark)RDMAdirect DMACPU + system DRAMoff the data pathbounce buffer~15% CPU utilizationin benchmarksVAST NVMe poolall-flash, NFS-over-RDMANICRDMA network adapterPCIe switchNIC, GPU and CPU attachGPU memoryfed directly

Read throughput to a DGX-2 at 1 MB IO, in benchmarks

same ratio as the pipe widths

CPU-mediated
~33 GB/s
GPUDirect
94+ GB/s
050100

GB/s · the tick on the GPUDirect bar marks peaks of ~98 GB/s

CPU-mediated
15.2 s
GPUDirect
5.3 s

Illustrative arithmetic: dataset size divided by the DGX-2 benchmark rates (about 2.8x apart). Real loads depend on platform, IO size and configuration.

Different platform

162 GiB/s read with GPUDirect on a DGX A100 with 8× HDR 200Gb NICs, in benchmarks - a different system, so it is not drawn on the DGX-2 scale above.

In benchmarks, ~33 GB/s on the legacy CPU-mediated path vs 94+ GB/s sustained (peaks ~98 GB/s) at 1 MB IO with GPUDirect on a DGX-2, at ~15% CPU utilization; and 162 GiB/s read on a DGX A100 with 8× HDR 200Gb NICs. An earlier figure of 92.6 GB/s to a DGX-2 also circulates. All figures come from benchmarks and vary by platform and configuration. Board layout is schematic.

Real-time RAG pipeline: embed-on-write → index → retrieve

Instead of a nightly batch job, InsightEngine turns every write into a retrieval-ready record: NIM embedding agents fire on write to create vectors and graph relationships, the index lives next to the source data, and GPU-accelerated cuVS search serves it back at inference time. For documents that start outside VAST, SyncEngine is the step before embed-on-write: it copies them onto the platform and keeps them in sync, so the index is only as fresh as the last sync.

Nightly batch vs embed-on-write: what can a question find?

Six documents land during one day. Drag the question through the day and compare what each pipeline can retrieve at that moment.

Nightly batch index

separate job, once a night

Embed on write

InsightEngine

? 17:30
00:0006:0012:0018:0024:00

Nightly batch

0/5of today's writes findable

Today's writes wait for tonight's job - answers come from yesterday's index.

Embed on write

5/5of today's writes findable

Each write is retrieval-ready moments after it lands.

  1. 1

    Embed on write

    When data lands, the platform triggers NIM embedding agents that create vectors and graph relationships - no separate batch indexing job to schedule.

  2. 2

    Index in place

    Vectors and graph relationships are stored natively in the platform next to the source records, so the index never drifts from the data it describes.

  3. 3

    Retrieve on the GPU

    GPU-accelerated vector search via NVIDIA cuVS serves retrieval at inference time, feeding fresh, grounded context straight into the model.

last night's index run write, not yet searchable embed + index in place searchable

Write times and the once-a-night schedule are illustrative.

KV-cache offload to fast storage

For long-context and agentic inference, VAST serves as an external KV-cache tier. Inactive KV cache is evacuated from scarce GPU memory through tiers - GPU device memory → host RAM → local NVMe → external VAST pool, reached over NFS over RDMA, GPUDirect Storage, S3, or S3 over RDMA (in the next VAST AI OS release) - and paged back in across turns instead of being recomputed. With NVIDIA Dynamo using VAST as that tier, time to first token fell from 62 s to 3 s in benchmarks (Dec 2025). The Context Memory page works through when reloading beats recomputing and what a cache hit is worth in $/token; NVIDIA STX & CMX covers the BlueField-4 tier. Step through a busy GPU below to watch the KV state move.

A busy GPU with many conversations

Memory pressure pushes an idle chat's KV cache down the tiers; when the user returns, it is fetched back instead of recomputed. Each square is a KV block.

With NVIDIA Dynamo
1/6
Chat AdecodingChat BdecodingChat CdecodingChat Dnot started

GPU device memory

HBM - scarce, fastest · tens of GB

12/12 slotsfull - memory pressure

Host RAM

system DRAM · hundreds of GB

0/16 slots

Local NVMe

node-local flash · TBs

0/20 slots
over the network: NFS over RDMA or TCP, GDS, S3 - S3 over RDMA next

External VAST pool

NFSoRDMA · GDS · S3 · S3oRDMA · PB+

10/32 slotsgrey = other sessions' KV
1 · Three chats

Three conversations are decoding. Their KV cache fills GPU HBM - the scarcest, fastest tier.

When chat A returns

Recompute prefillre-run the model over the whole context
Fetch KVpage the stored blocks back in

Time to the first new token, schematic - bar lengths are illustrative. The saving matters most at the 100k+ token contexts this targets.

VAST serves as a shared KV-cache tier, reached over NFS over RDMA, GPUDirect Storage, S3 and, from the next VAST AI OS release, S3 over RDMA - keeping KV state in the serving loop across turns and paging in and out to cut prefill recompute.

In benchmarks (Dec 2025), NVIDIA Dynamo with VAST as the KV tier cut time to first token from 62 s to 3 s. Block counts and bin widths are illustrative, not to scale.

Why this matters for AI factories

The throughput of an AI factory is increasingly set by the data layer, not the GPU. InsightEngine attacks all three failure modes at once: stale retrieval, idle GPUs, and repeated prefill.

One GPU, six requests: switch on each fix

A schematic serving window. Orange is waste - the GPU idle on a CPU copy, or re-running prefill it already did. Turn on the three capabilities and the same work finishes sooner, with fresh answers.

GPU time
same six requests
Answers
fresh or stale
✓!✓✓!!
serve tokensidle on CPU copyrecompute prefilldirect readfetch KV! stale answer

44%

of busy GPU time serving tokens (illustrative)

3 of 6

stale answers (illustrative)

Schematic: durations are illustrative units, not measurements, so the percentages only show direction.

Together these keep the most expensive hardware in the building serving tokens instead of waiting - the same throughput logic that runs through the NVIDIA Infrastructure, Inference & Serving, KV Cache, and RAG & Vector Search pages, applied end to end.

Key takeaways

In one line

InsightEngine is VAST's NVIDIA-integrated, full-stack system for real-time RAG - GPUDirect Storage feeds GPUs directly and embed-on-write keeps retrieval current.

Key points

  • Unveiled in October 2024, InsightEngine has been available on NVIDIA DGX systems since GTC in March 2025, pre-integrated with NVIDIA AI Enterprise.
  • GPUDirect Storage DMAs data from NVMe into GPU memory over RDMA - 94+ GB/s on a DGX-2 versus ~33 GB/s on the CPU path in benchmarks.
  • Data is embedded on write via NIM microservices and indexed in place, so cuVS retrieval stays grounded in the latest data - no nightly re-index.
  • VAST also serves as an external KV-cache tier below GPU memory, host RAM, and NVMe; with NVIDIA Dynamo, time to first token fell from 62 s to 3 s in benchmarks (Dec 2025).

Questions to explore

  1. 01How much of your GPU time goes to waiting on data instead of serving tokens?
  2. 02How stale is the index your RAG pipeline retrieves from between batch re-index jobs?
  3. 03At what context length does your serving stack start recomputing prefill instead of reusing cached KV state?

Common questions

Is InsightEngine generally available, or is this a future product?
Available on DGX systems since March 2025, pre-integrated with NVIDIA AI Enterprise and expanding to NVIDIA-Certified Systems.
What throughput does GPUDirect Storage reach?
In benchmarks, ~33 GB/s CPU-mediated versus 94+ GB/s on a DGX-2, and 162 GiB/s on a DGX A100 with 8× HDR 200Gb NICs. Results vary by platform and configuration.
How does the KV-cache tier relate to NVIDIA STX and CMX?
With NVIDIA Dynamo, VAST already serves as an external KV tier (62 s to 3 s time to first token in benchmarks, Dec 2025); in Jan 2026 VAST announced the AI OS running on BlueField-4 for NVIDIA's CMX context-memory platform, covered on the STX page.