VAST Data
VAST VectorStore
AdvancedVector search native to the data platform - vectors live next to the source data, so RAG retrieves without a separate vector DB to sync.
Recap: embeddings & vector search
An embedding model turns content into a vector, and retrieval is a nearest-neighbor search over those vectors - see embeddings and vector search fundamentals for how closeness is measured. VAST exposes three distance functions as built-in SQL: , , and the . This page is about what changes when vectors live inside the data platform instead of a bolt-on vector DB.
Vectors as a first-class type, co-located with source data
In VAST, a vector is a first-class column stored in the same table as its metadata and source content. There is no separate vector database to keep in sync, and no two-step “search the index, then fetch the document from object storage” hop. Inserting a row and its vector is a single atomic transaction - strong consistency, versus the eventual consistency you inherit when vectors live in an external store that an ETL job has to catch up to. Access controls on the vector can't drift from the source data's ACLs, because they are the same row. (For the drift and second-hop costs of a standalone store, see the separate vector DB problem.)
Edit one row, then query it
Row #42 changes and a query arrives before any sync has run. With an external vector DB, the row and its vector live in two systems. In VAST, the vector is a column in the same row.
Schematic: the steps are a story, not a time scale, and the rows are illustrative. The two-step retrieval and drift window are inherent to any design that keeps the vector index in a separate store from the source data; co-location and atomic consistency are design facts of storing vectors as a first-class column.
How the index scales: hierarchical clustering
The reason VAST can drop sharding is the shape of the index itself. Instead of one giant in-memory graph, vectors are organized into a recursive hierarchy of distance-clusters, each level summarized by - think of a library organized into sections, shelves, then books, so you can find a title without walking every aisle. Each level holds up to ~1,000 of the level below, so capacity grows ~1,000× per level while search stays logarithmic.
Explore the two halves of the design below: the L0-L5 capacity hierarchy, and the background reclustering lifecycle - new vectors land in a fast non-clustered table and are organized in the background, taking the tree from messy, overlapping clusters to clean, well-separated ones.
The distance-cluster tree, and how it stays dense
Each level holds ~1,000 of the level below. Zoom through the levels, then watch a background recluster tidy messy clusters.
Tap a circle to zoom into one L4 cluster.
You are inside one
L5 cluster
~1 quadrillion vectors · holds up to 1,000 L4 clusters
Vectors per cluster (log scale)
Each step down divides by ~1,000, so a query that follows the closest branches touches a handful of centroids per level and resolves inside one L1 leaf chunk - logarithmic work. Adding a level multiplies capacity ~1,000× with no resharding, index freeze, or global rebalance.
The L0-L5 capacities, ~1,000× per-level fan-out, 10% K-means sample, and reclustering design are from VAST's published vector index, designed for sub-200 ms search at trillion-vector scale. Circles show a sample of 19 children each; the K-means view runs on toy 2-D data.
Because the tree grows by adding levels rather than shards, scaling from billions to trillions needs no resharding, index freeze, or global rebalance, and continuous ingest never blocks queries. The documented ceiling is 256T rows (2⁴⁸) per table - the 50-billion proof point is roughly 0.02% of it.
Inside the walk: beam width & the ratio rule
The walk isn't a fixed top-k at each level. VAST carries a beam of the closest centroids and widens it toward the leaves, bounded by a per-level min/max and driven by a ratio rule: it keeps expanding a level while the distance to the farthest selected centroid stays within a set ratio (default 1.8) of the closest. The whole walk runs on 8-bit (int8) scalar-quantized codes; only the final leaf candidates are fetched at full precision and reranked. Tune the ratio and watch the work-versus-recall trade.
The beam, drawn by distance from the query
Every dot is a candidate centroid, placed by its distance from the query. The beam is whatever falls inside the dashed ring at ratio × d(closest), clamped to the level's min/max. Drag the ring or the slider.
L2: 81 selected - 81 inside the ring, within 20-150.
Beam per level, on a 1,000-cell fan-out
↓ 4 × 1,000 children = 4.0K scored at L3
↓ 4 × 1,000 children = 4.0K scored at L2
↓ 81 × 1,000 children = 81K scored at L1
Quantized scans (int8)
90K
centroids + leaf vectors compared
Full-precision reranks
189
originals fetched at the leaf only
Corpus scanned
9.0e-6%
of 1.0T vectors
A wider beam explores more branches, so a true neighbor just across a cluster boundary is less likely to be pruned - at the cost of more scans. The walk still touches a vanishing fraction of the corpus, and the leaf rerank restores full precision. Admins set the ratio with vsetting; a session can override min_per_level / max_per_level at query time.
Illustrative model of a 1T-vector table with 1,000-way branching. Beam bounds and the 1.8 default ratio are VAST-documented; the dot fields are toy data, so beam counts show the rule, not a measured index.
Benchmarked: 11× faster, ~91% lower cost
The architecture differences show up in numbers. On a 1-billion-vector, 128-dimension benchmark at 99% recall, VAST delivers roughly 11× the throughput of a leading disk-based shared-nothing system, at about 91% lower cost per thousand searches, because the vectors no longer have to sit in expensive RAM. At 50 billion vectors - a scale the comparison system was not tested at - VAST sustains over a thousand queries per second on just 8 CNodes. To see why RAM is the bill at this scale, size an index in the hyperscale memory calculator - a trillion-vector index can live on NVMe but not in DRAM.
VAST vector search, in benchmarks
More than 11× the throughput at 1 billion vectors (128-dim, 99% recall), about 91% lower cost per search, and a result at 50 billion vectors where the comparison system was not tested.
Throughput at 1B vectors
higher is better1,000 ÷ 89 ≈ 11.2 Milvus-sized blocks of queries per second.
Cost per 1,000 searches
lower is betterThe baseline bill is 11 VAST-sized bills: $0.033 ÷ $0.003 = 11, so about 91% lower.
Tested envelope - throughput vs dataset size
benchmark points only · tap a pointVAST at 50B vectors on 8 CNodes, 375 threads: 1,026 QPS at ~317ms.
Points are separate test runs, not a measured curve - nothing is drawn between them. The two 50B results show the concurrency trade: more threads, more total QPS, higher latency per query.
All figures are from benchmarks: 1B vectors at 128-dim and 99% recall, ~1,000 QPS vs Milvus 2.6 disk (AISAQ) ~89 QPS; ~$0.003 vs ~$0.033 per 1,000 searches; 50B vectors on 8 CNodes, 1,026 QPS at ~317ms (375 threads) and 606 QPS at ~114ms (75 threads); ingest ~1M vectors/s via Parquet; background clustering ~290K vectors/s.
Indexing at write, GPU search, and hybrid vector + SQL queries
Indexes are built at write time - zone maps, the vector index, and secondary indexes are all maintained on ingest, so there is no separate “build the index” batch job. Search runs on CPU or GPU algorithms, including NVIDIA CAGRA, without caching the whole index in GPU memory. Because data is laid out in small columnar chunks, the engine can prune irrelevant blocks early instead of scanning everything.
The payoff for co-location is the hybrid query: vector similarity and SQL metadata filters in a single statement - for example, the nearest neighbors to a query embedding, filtered to a brand and a price range, without round-tripping between two systems.
It is genuinely SQL, not a bolt-on API. Vectors are a first-class column (fixed-dimension PyArrow list types), and built-in distance functions - array_distance() for Euclidean and array_cosine_distance() for cosine similarity - drop straight into a query, so you WHERE-filter on metadata, ORDER BY similarity, and LIMIT to the top matches in one statement. Apps reach it over the Arrow stack (ADBC driver, the vastdb Python SDK), which keeps vectors and results in columnar form end to end.
GPUs accelerate more than model inference here - they speed up the whole vector pipeline. Index builds run about 4× faster on GPU (10 hours down to 2.5 hours) using NVIDIA's cuVS library, while ingesting on the order of a million vectors per second. For SQL, Sirius - an open-source GPU SQL engine from the University of Wisconsin-Madison, in technical preview on VAST - ran up to 20× faster than DuckDB on NVIDIA RTX PRO 6000 in benchmarks (the VAST DataBase page puts that next to its other Sirius figures).
GPUs across the whole vector pipeline
Not just inference: GPUs speed up ingest, index build and SQL query. Pick a stage to race GPU against CPU on the same clock.
Index build on the clock
Both builds start together on the same data.
In benchmarks: GPU index build ~4x faster using NVIDIA cuVS (10h to 2.5h); Sirius GPU SQL up to 20x faster than DuckDB on NVIDIA RTX PRO 6000; ~1M vectors/s ingest (1024-dim). The SQL race is drawn at the best case (20x); real speedups vary by query.
One engine, every retrieval technique
Production retrieval usually means bolting extra techniques onto a vector DB - hybrid fusion, filtered (ACORN), quantized rescoring, MMR diversity, GPU index builds. Because VAST stores vectors as a SQL column, most of these collapse into a single query path. Compare them technique by technique.
VAST vs conventional vector-DB techniques
The retrieval tricks teams bolt onto vector DBs, drawn as pipelines. Pick a technique and watch the conventional stack collapse into one query path.
Combine semantic (vector) and lexical (BM25) signals into one ranking.
Run a vector search and a lexical search separately, then fuse the lists in the app with RRF or DBSF (Qdrant, Weaviate).
one SQL query
Both axes run inside one SQL query; weighted fusion is expressible in SQL - no app-level merge. Default top-1000 candidate pool (vs ~100 on AWS S3 Vectors) gives 10× to rerank over.
Index families, head to head
practical scale, log axis · tap a familySolid = practical or proven scale; dashed = VAST's documented design ceiling (2⁴⁸ rows per table).
- Index location
- HNSW: RAM (graph + vectors)VAST: NVMe flash (vectors + index)
- Memory cost
- HNSW: ≈4d + 8M B/vec (~91TB at 20B×1024d)VAST: Near-zero per vector (only centroids in RAM)
- Practical scale
- HNSW: ≈1B / nodeVAST: 50B proven · 256T ceiling
- Ingest cost
- HNSW: Hours-days graph rebuildVAST: Fast unclustered landing table + background reclustering
- Filter / metadata
- HNSW: Post-filter or ACORNVAST: Native SQL push-down (any column)
- Best for
- HNSW: <1B, high recallVAST: Trillion-scale RAG, real-time updates
In this comparison, only hierarchical K-means scales beyond one node without sharding, and only its memory cost stays flat as the corpus grows.
What you can configure
The vector column and its index expose a small, typed set of knobs at schema, index, and query level.
Every knob, pinned to what it controls
One vector table and one hybrid query. Tap a knob to see where it lives - at schema, index, or query level. Tap the distance metric again to cycle its SQL function.
Schema level
Index level
Query level
Table docs
room for up to 31,936 typed metadata columns beside the vector
one embedding per row, up to 4,096 components, fixed when the column is created
Hybrid query (illustrative)
SELECT id, title FROM docs WHERE brand = 'acme' AND price < 50 ORDER BY array_cosine_distance( embedding, :query_vec) LIMIT 10;
Manhattan distance is not currently supported (under consideration). Configuration details are VAST-documented. The table, column names and query are illustrative.
New in VAST AI OS 5.5
VAST AI OS 5.5, generally available since August 6, 2026, ships the index this page describes as the Hyperscale Vector Index: the hierarchical, quantized, DASE-persisted index above, with memory-bounded queries at trillion-vector scale - so latency plateaus as the dataset grows instead of degrading. The release also adds 50+ aggregation functions to the native query engine, and SQL-native partitioning: a partition layout you can declare on a table when it helps, on top of the per-chunk pruning that lets most VAST DataBase tables skip partition design entirely. In benchmarks against Iceberg-based tables, partitioned tables ran highly selective queries more than 60% faster and updates 200× faster, with about 20% lower aggregate runtime in one manufacturing workload.
VAST AI OS 5.5 · generally available Aug 2026
Vector search and SQL analytics, upgraded
Three additions that shipped in 5.5: a hierarchical vector index, a larger query engine, and SQL-native partitioning. Walk a query through the index to see why it reads so little.
Inside the Hyperscale Vector Index
fragments read: 0 / 21At rest: the whole index is persisted on DASE flash - nothing is pinned in RAM.
Native Query Engine
50+new aggregation functions for in-place SQL analytics
one cell per function - 50 shown, the release adds more
SQL-native partitioning
vs Iceberg-based tables, in benchmarks
- >60%faster on highly selective queries
- 200×faster updates
- ~20%lower aggregate runtime in one manufacturing workload
The index tree is a schematic - real levels hold far more fragments, and the 7-of-21 count is illustrative.
End-to-end RAG on one platform
Because vectors are native, the whole retrieval loop runs on a single platform: ingest the source data, embed it with serverless functions or NVIDIA NIM, index it at write, and retrieve with hybrid queries - no ETL drift between the vectors and the documents they came from. That is the difference between a vector store you have to babysit and one that stays consistent by construction.
Key takeaways
In one line
Vectors live as a first-class SQL column next to source data on VAST, replacing bolt-on vector databases with atomic writes and a single hybrid query.
Key points
- Inserting a row and its vector is one atomic transaction - strong consistency, unlike external vector DBs' ETL sync and two-step retrieval.
- A hierarchical, distance-clustered index (not an in-RAM graph) prunes about 1,000x the candidates per level, so search stays logarithmic without sharding.
- On a 1-billion-vector, 128-dimension, 99%-recall benchmark, VAST delivered roughly 11x the throughput of a leading disk-based system and about 91% lower cost per 1,000 searches.
- GPU index builds run about 4x faster via NVIDIA cuVS (10 hours down to 2.5 hours), and the open-source Sirius GPU SQL engine ran up to 20x faster than DuckDB in benchmarks.
Questions to explore
- 01How many separate systems (vector DB, warehouse, ETL job) keep your embeddings in sync with source data?
- 02How large could your vector corpus grow, and would an in-memory graph index like HNSW still fit in RAM at that scale?
- 03Do your retrieval queries need to combine vector similarity with metadata filters (tenant, date, price) in a single request?
Common questions
- How does search stay fast at trillion-vector scale?
- Each index level prunes ~1,000x the candidates and the DASE design needs no sharding; the design target is sub-200 ms search at trillion-vector scale with no index freezes, and in benchmarks 50 billion vectors on 8 CNodes answered in about 114-317 ms.
- What happens to recall when VAST prunes 1,000x the candidates away at each level?
- The index keeps a "beam" of the closest centroids per level, widened by a tunable ratio rule (default 1.8); only final candidates are fetched at full precision and reranked, reaching 96-99% recall at 50 billion vectors.
- Which distance metrics and configuration limits apply?
- Euclidean, cosine distance, and negative inner product are built into SQL; vectors support up to 4,096 dimensions and up to 31,936 metadata columns; Manhattan distance is not currently supported.