VAST Data

VAST Foundation Stacks

Advanced

NVIDIA AI Blueprints get you a working demo; the gap to production is the integration tax. VAST Foundation Stacks are open-source implementations that close it - running RAG, AI-Q deep research, and Video Search & Summarization as production pipelines on the VAST AI OS.

From Blueprint to production

An is a reference workflow - a runnable starting point built from microservices. Getting it into reliable, secure, scalable production is the hard part: data pipelines, vector storage, eventing, orchestration, security, ops. VAST Foundation Stacks pre-package that on one platform, so a Blueprint becomes something you git clone and deploy. They are open source at github.com/vast-data/cosmos-labs.

Blueprint → Foundation Stack

One production stack. The NVIDIA Blueprint supplies the top; the dashed layers are the integration tax. Step forward to watch the Foundation Stack fill them.

1/5Blueprint only: the AI recipe sits on four empty layers you would have to build and run yourself.

Supplied by the NVIDIA Blueprint

The AI recipe

The models, prompts, and workflow for one use case (RAG, AI-Q, VSS…).

Layers you build yourself

4 / 4

Tap any layer to see what it covers.

NVIDIA Blueprintintegration tax (you assemble)VAST Foundation Stack

A Blueprint gets you a working demo; the Foundation Stack pre-packages the layers underneath, so the gap from pilot to production closes to a git clone.

One shape under every stack

However different the jobs look, all three stacks share the same anatomy: a serverless ingest pipeline on DataEngine fills a unified VAST DB store, and a Kubernetes app serves users by calling NVIDIA NIMs. Learn the shape once and all three click into place.

The shared anatomy

Two lanes, one store: ingest writes into VAST DB, the app reads from it, and both call NVIDIA NIM. Tap a block to read what it does.

1/7A file lands in S3, and the upload fires a DataEngine trigger.
SERVEINGESTtriggerembedwriteretrievererank + generateUserKubernetes appthe serve planeS3 bucketDataEngine ingestserverless, event-drivenVAST DBone unified storeNVIDIA NIMthe AI, on GPUs

DataEngine ingest

serverless, event-driven

Triggers + functions fire the instant data lands in S3 - no always-on service. Chunk, embed, segment, write.

VAST DB sits between both lanes - what ingest writes is exactly what the app reads.

The three pipelines, running

Each launch stack maps to a different NVIDIA Blueprint. Step through the real ordered stages - the function and model names come straight from the VAST blueprint repos.

The three Foundation Stacks - one production shape

Pick a stack and watch its real pipeline run through the layers. Each lane is a layer of the VAST AI OS or NVIDIA NIM; the packet shows which one does each stage. Tap a stage to pin it.

The foundational retrieval-augmented generation pipeline: ingest and embed documents, then answer questions grounded on the most relevant chunks with citations.

Based on the NVIDIA Enterprise RAG Blueprint

1/9
S3 storageS3 bucketsDataEngineserverlessVAST DBunified storeNVIDIA NIMon GPUs1 · INGEST PIPELINE (SERVERLESS)2 · QUERY (READ PATH)Search readyuser queryUpload/ SyncTriggerExtract+ ChunkEmbed(NIM)StoreEmbedqueryRetrieve(ANN)Rerank(NIM)Generate+ cite
S3 storageUpload / SyncIngest pipeline (serverless)

Documents land in an S3 bucket, or are pulled from Confluence / GDrive / SharePoint by the Sync Engine.

One-shot retrieve → generate. The foundation every other stack builds on.

RAG vs AI-Q: one-shot vs a research loop

The clearest way to understand AI-Q is to see it as with an agent wrapped around it. Plain RAG retrieves once and generates one answer. AI-Q plans the question into sub-questions, calls retrieval several times with different angles (each call is a full RAG retrieve-and-), reflects on what is missing, and only then writes a structured, cited report. The same ingest pipeline and the same store feed both - VAST DB also persists the agent's conversation state, so agent state lives in the same store as the data.

One pass vs a research loop, on the same clock

Each cell is one step of work. Step counts are illustrative - the point is the shape, not the timing.

1/8The same question reaches both pipelines.

RAG

one pass

AI-Q

research loop

Embed query

Retrieve

top-30

Rerank

to 10

Generate

one grounded answer

answered - ready for the next question

Plan

sub-questions

Retrieve

angle 1

Retrieve

angle 2

Retrieve

angle 3

Reflect

what is missing?

Retrieve

fill the gap

Synthesize

cited report

time runs down ↓

RAG retrieval calls

0

AI-Q retrieval calls

0

Under both: the same ingest pipeline and the same VAST DB store. VAST DB also persists AI-Q's conversation state, so agent state lives in the same store as the data.

RAG - one pass

query → retrieve top-30 → rerank to 10 → generate one grounded answer with citations. Fast, predictable, great for direct questions over your docs.

AI-Q - a research loop

Plan → retrieve repeatedly (multi-angle) → reflect on gaps → synthesize a long-form report. Trades latency for depth - built for “research this for me,” not “answer this.”

Go deeper

Foundation Stacks tie together the building blocks taught across the site - the retrieval mechanics, the agent patterns, the serverless engine, the real-time RAG platform they run on, and the on-ramp that brings the documents onto it.