VAST Data
VAST Foundation Stacks
AdvancedNVIDIA AI Blueprints get you a working demo; the gap to production is the integration tax. VAST Foundation Stacks are open-source implementations that close it - running RAG, AI-Q deep research, and Video Search & Summarization as production pipelines on the VAST AI OS.
From Blueprint to production
An is a reference workflow - a runnable starting point built from microservices. Getting it into reliable, secure, scalable production is the hard part: data pipelines, vector storage, eventing, orchestration, security, ops. VAST Foundation Stacks pre-package that on one platform, so a Blueprint becomes something you git clone and deploy. They are open source at github.com/vast-data/cosmos-labs.
Blueprint → Foundation Stack
One production stack. The NVIDIA Blueprint supplies the top; the dashed layers are the integration tax. Step forward to watch the Foundation Stack fill them.
Blueprint supplies
the AI recipe
Integration tax
4 of 4 layers left to build
Supplied by the NVIDIA Blueprint
The AI recipe
The models, prompts, and workflow for one use case (RAG, AI-Q, VSS…).
Layers you build yourself
4 / 4
Tap any layer to see what it covers.
A Blueprint gets you a working demo; the Foundation Stack pre-packages the layers underneath, so the gap from pilot to production closes to a git clone.
One shape under every stack
However different the jobs look, all three stacks share the same anatomy: a serverless ingest pipeline on DataEngine fills a unified VAST DB store, and a Kubernetes app serves users by calling NVIDIA NIMs. Learn the shape once and all three click into place.
The shared anatomy
Two lanes, one store: ingest writes into VAST DB, the app reads from it, and both call NVIDIA NIM. Tap a block to read what it does.
DataEngine ingest
serverless, event-driven
Triggers + functions fire the instant data lands in S3 - no always-on service. Chunk, embed, segment, write.
VAST DB sits between both lanes - what ingest writes is exactly what the app reads.
The three pipelines, running
Each launch stack maps to a different NVIDIA Blueprint. Step through the real ordered stages - the function and model names come straight from the VAST blueprint repos.
The three Foundation Stacks - one production shape
Pick a stack and watch its real pipeline run through the layers. Each lane is a layer of the VAST AI OS or NVIDIA NIM; the packet shows which one does each stage. Tap a stage to pin it.
The foundational retrieval-augmented generation pipeline: ingest and embed documents, then answer questions grounded on the most relevant chunks with citations.
Based on the NVIDIA Enterprise RAG Blueprint
Documents land in an S3 bucket, or are pulled from Confluence / GDrive / SharePoint by the Sync Engine.
One-shot retrieve → generate. The foundation every other stack builds on.
RAG vs AI-Q: one-shot vs a research loop
The clearest way to understand AI-Q is to see it as with an agent wrapped around it. Plain RAG retrieves once and generates one answer. AI-Q plans the question into sub-questions, calls retrieval several times with different angles (each call is a full RAG retrieve-and-), reflects on what is missing, and only then writes a structured, cited report. The same ingest pipeline and the same store feed both - VAST DB also persists the agent's conversation state, so agent state lives in the same store as the data.
One pass vs a research loop, on the same clock
Each cell is one step of work. Step counts are illustrative - the point is the shape, not the timing.
RAG
one pass
AI-Q
research loop
Embed query
Retrieve
top-30
Rerank
to 10
Generate
one grounded answer
Plan
sub-questions
Retrieve
angle 1
Retrieve
angle 2
Retrieve
angle 3
Reflect
what is missing?
Retrieve
fill the gap
Synthesize
cited report
time runs down ↓time →
RAG retrieval calls
0
AI-Q retrieval calls
0
RAG - one pass
query → retrieve top-30 → rerank to 10 → generate one grounded answer with citations. Fast, predictable, great for direct questions over your docs.
AI-Q - a research loop
Plan → retrieve repeatedly (multi-angle) → reflect on gaps → synthesize a long-form report. Trades latency for depth - built for “research this for me,” not “answer this.”
Go deeper
Foundation Stacks tie together the building blocks taught across the site - the retrieval mechanics, the agent patterns, the serverless engine, the real-time RAG platform they run on, and the on-ramp that brings the documents onto it.