VAST Data
VAST AI Operating System
FundamentalData, not the GPU, increasingly sets the throughput of an AI factory - VAST is the platform built for that.
The data problem in the AI era
The most expensive hardware in a data center now spends much of its life waiting. GPUs sit idle while data crawls to them from slow, tiered storage - industry GPU utilization has been cited as low as ~5% (VentureBeat). VAST's framing is blunt: “this isn't a GPU problem, it's a data problem - your GPUs are starving.”
Where the GPUs wait - tiered storage vs a growing cluster
Warm AI data sits on the slow tier and has to be staged before a GPU can use it, so the data layer delivers a fixed rate. Grow the cluster to see where the bottleneck moves. Illustrative units.
1Why traditional storage fails here
Tiered storage and generic cloud object stores were built for cost-per-terabyte, not for feeding racks. Warm AI data lands on slow tiers, and the GPUs stall waiting for it to be staged.
2The bottleneck moved
As clusters scale into thousands of accelerators, the limiting factor stops being FLOPs and becomes how fast the data layer can keep every GPU fed. Throughput is now a storage problem.
~5% GPU utilization figure cited by VentureBeat; actual utilization varies widely by workload and deployment.
AI factory & token factory
NVIDIA frames the modern AI data center as a factory: it takes data as raw material and turns it into intelligence, measured as token throughput. The output of that factory is set by the whole line - and the data layer, not just the GPU, governs how many tokens come off the end.
The factory line - starve any stage and the whole line slows
Tap a stage to starve it. The flow narrows from that stage on, and the token output drops no matter how fast the later stages are. Relative, illustrative units.
Every stage keeps up, so the line runs at full rate.
Raw data → data layer → GPUs → tokens. Starve any stage and the whole line slows. The same logic that makes NVIDIA infrastructure and inference serving throughput-bound applies upstream: the factory only runs as fast as its slowest input.
See the site's NVIDIA Infrastructure and Inference & Serving pages for how the GPU and serving stages turn fed data into tokens.
One platform, not a stack of point solutions
A typical AI data estate is a stack of separate products - a data lake, a warehouse, a vector DB, a streaming bus, and a compute cluster - each with its own copy of the data and its own ETL between them. VAST collapses all of that into one platform on a single namespace spanning files, objects, tables, and vectors, under one permission and security model.
One namespace under every capability
The same data estate drawn two ways. Switch views, and pick a capability to see what it touches. The pipeline layout in the point-solutions view is illustrative.
Data lake + warehouse: highlighted in the diagram.
The payoff is eliminating the copies and ETL that move data between systems - fewer pipelines to break, one source of truth, and no consistency lag between what the warehouse, the vector index, and the agents see.
The VAST AI Operating System
VAST's thesis is that the AI era needs an operating system, not a pile of point products: storage, a transactional and analytical database, Kafka-compatible streaming, serverless compute and governed agents converge on one platform, so data, compute and intelligence live in the same place. Unveiled at a New York launch in May 2025, with more engines announced at the first VAST Forward conference in February 2026, the AI OS is a layer cake: Storage () at the base, the DataBase / VectorStore on top, the adding serverless compute and an Event Broker, and the running agents and at the top - with , the NVIDIA real-time RAG stack, spanning the layers. SyncEngine is the on-ramp rather than a layer: it copies data from outside systems - file shares, object buckets and collaboration apps - onto the storage layer and keeps it in sync. Click a layer to see what it does and why it matters for AI.
The AI OS - one foundation, four layers, one copy of the data
Every layer addresses the same data on the same namespace. Play the data journey to follow one document from an outside system, through SyncEngine onto storage and up to an agent, then switch to separate systems to see the same path with a copy at every hop. Tap any layer for what it does.
SyncEngine picks up a document from an outside system - an old file share, a cloud bucket or a collaboration app - and copies it in, then keeps it in sync.
SyncEngine
Data onboardingGenerally available- What it is
- The on-ramp, not a layer: it copies data from outside systems - S3 and Google Cloud Storage buckets, NFS and POSIX file systems, Confluence, Google Drive and OneDrive - onto VAST S3 buckets or NFS views, and keeps the copy in sync. Available to all VAST customers since August 2025, at no additional cost.
- Problem it solves
- The data AI could use sits on old file shares, in cloud buckets and in collaboration apps, and moving it usually means scripts plus a standalone migration tool.
- Why it matters for AI
- Scheduled resyncs (or, for AWS S3, bucket events) keep the copy current, and each change that lands can fire a DataEngine trigger - so outside data joins the same event-driven pipeline.
What is shipping today
The engines below are at different stages. As you read the deep dives, keep maturity in mind - most of the platform is generally available, the agent layer is newer, and parts of the broader AI OS are still on the roadmap.
Maturity by layer - the foundation ships, the top is newest
The platform stack, colored by status. Newer layers sit higher. Tap a layer to open its deep dive. Status as of September 2026, from VAST announcements.
- On the roadmap
PolicyEngine and TuningEngine (targeted end 2026) and the wider AI OS vision - see the Roadmap page for what is announced versus already shipping.
- In preview
AgentEngine & MCP (announced May 2025, upcoming release) and Sirius GPU SQL (technical preview since aug 2026) - announced, not yet generally available.
- Generally available
Storage (DASE), DataBase, VectorStore, DataSpace, DataEngine & Event Broker, and InsightEngine are shipping today, and SyncEngine - for bringing outside data in - has been available since August 2025.
Maturity reflects VAST's public announcements as of September 2026 and may change; confirm current availability with VAST before deploying.
For dates, the engines still to come and the closed loop they build toward, see Roadmap & Vision.
Removing the storage-vs-GPU tradeoff
The old choice was fast storage you couldn't afford at scale or cheap storage that starved the GPUs. VAST's architecture aims to remove that tradeoff: feed GPUs directly while keeping all warm AI data on flash.
Three mechanisms behind feeding GPUs from flash
Pick a mechanism, then flip its toggle to compare with the traditional setup. Diagrams are schematic.
GPUDirect Storage
Data moves straight from storage into GPU memory, bypassing the host CPU and bounce buffers on the read path.
One hop: the NIC writes straight into GPU memory. The CPU and system memory stay off the data path.
Realized all-flash cost depends on data reduction rates and configuration.
Where to go deeper
Each pillar of the platform has its own deep dive - the architecture underneath, the data and vector layers, getting outside data in, the compute and agent engines, the NVIDIA stack, and the roadmap.
VAST Architecture
IntermediateDASE: disaggregated, shared-everything architecture.
VAST DataBase
IntermediateA columnar table format - an Iceberg/Delta alternative, native to the platform.
VAST VectorStore
AdvancedVector search at scale, native to the data platform.
VAST DataSpace
IntermediateOne global namespace across edge, on-prem, and every cloud.
VAST SyncEngine
FundamentalFind data across file shares, object stores and SaaS apps - and bring it onto VAST.
DataEngine & Event Broker
IntermediateServerless functions, triggers, and Kafka-compatible streaming.
AgentEngine & MCP
AdvancedRunning and governing AI agents next to the data.
InsightEngine (NVIDIA)
IntermediateReal-time RAG and the NVIDIA AI data-platform stack.
Roadmap & Vision
FundamentalThe AI OS: Polaris, GPU SQL, and the agentic future.
Key takeaways
In one line
GPUs increasingly wait on data, not compute - VAST unifies files, objects, tables and vectors in one platform, on one namespace, to keep them fed.
Key points
- Industry GPU utilization has been cited as low as ~5% (VentureBeat) because data can't keep up with compute.
- VAST unifies files, objects, tables and vectors in one namespace under one permission and security model, not four.
- The AI OS layers Storage (DASE), DataBase/VectorStore, DataEngine + Event Broker, and AgentEngine/MCP, with InsightEngine spanning every layer.
- GPUDirect Storage and an NVMe-oF fabric let every compute node reach the full all-flash pool directly, delivering all-flash at HDD-like cost, depending on data reduction.
Questions to explore
- 01How much of your GPU time goes idle waiting for data to be staged from storage?
- 02How many separate copies of the same dataset do your lake, warehouse and vector index each maintain?
- 03How much pipeline work goes into ETL and consistency-checking between separate data systems?
Common questions
- Where does the ~5% GPU utilization figure come from?
- VentureBeat reported it as an industry-wide figure; actual utilization varies widely by workload and deployment.
- Which parts of this AI OS actually ship today versus later?
- Storage, DataBase, VectorStore, DataSpace, DataEngine + Event Broker, and InsightEngine are shipping, and SyncEngine has been available to all VAST customers since August 2025; AgentEngine/MCP and Sirius GPU SQL are in preview, not yet GA; the wider AI OS vision is on the roadmap.
- What does "all-flash at HDD-like cost" depend on?
- Realized cost depends on data reduction rates and configuration, so it varies by deployment.