VAST Data

VAST VectorStore

Advanced

Vector search native to the data platform - vectors live next to the source data, so RAG retrieves without a separate vector DB to sync.

Vectors as a first-class type, co-located with source data

In VAST, a vector is a first-class column stored in the same table as its metadata and source content. There is no separate vector database to keep in sync, and no two-step “search the index, then fetch the document from object storage” hop. Inserting a row and its vector is a single atomic transaction - strong consistency, versus the eventual consistency you inherit when vectors live in an external store that an ETL job has to catch up to. Access controls on the vector can't drift from the source data's ACLs, because they are the same row. (For the drift and second-hop costs of a standalone store, see the separate vector DB problem.)

Edit one row, then query it

Row #42 changes and a query arrives before any sync has run. With an external vector DB, the row and its vector live in two systems. In VAST, the vector is a column in the same row.

hops per retrieval 2copies to keep in sync 2consistency eventual
① search → ids② fetch rowsApp / RAGExternal vector DBe.g. Pinecone, Milvus, PGVectoridvector#41v1#42v1#43v1Object store / lakehousesource rows + metadataidtextmetadata#41doc v1acl·tags#42doc v1acl·tags#43doc v1acl·tagsETL job (scheduled)
1/6Steady state: row #42 in the object store and its vector in the vector DB are both v1.
Row #42
object store
v1
v2
Vector #42
external vector DB
v1
stale
v2
ETL job
on a schedule
run
Query
two hops
①
② stale
①
②
Consistency
row vs vector
drift window
story steps →
Fresh versionStale vectorETL runQuery hopDrift windowVAST atomic row

Schematic: the steps are a story, not a time scale, and the rows are illustrative. The two-step retrieval and drift window are inherent to any design that keeps the vector index in a separate store from the source data; co-location and atomic consistency are design facts of storing vectors as a first-class column.

How the index scales: hierarchical clustering

The reason VAST can drop sharding is the shape of the index itself. Instead of one giant in-memory graph, vectors are organized into a recursive hierarchy of distance-clusters, each level summarized by - think of a library organized into sections, shelves, then books, so you can find a title without walking every aisle. Each level holds up to ~1,000 of the level below, so capacity grows ~1,000× per level while search stays logarithmic.

Explore the two halves of the design below: the L0-L5 capacity hierarchy, and the background reclustering lifecycle - new vectors land in a fast non-clustered table and are organized in the background, taking the tree from messy, overlapping clusters to clean, well-separated ones.

The distance-cluster tree, and how it stays dense

Each level holds ~1,000 of the level below. Zoom through the levels, then watch a background recluster tidy messy clusters.

19 of up to 1,000 L4

Tap a circle to zoom into one L4 cluster.

You are inside one

L5 cluster

~1 quadrillion vectors · holds up to 1,000 L4 clusters

Vectors per cluster (log scale)

L5
1015
L4
1012
L3
109
L2
106
L1
103
L0
1

Each step down divides by ~1,000, so a query that follows the closest branches touches a handful of centroids per level and resolves inside one L1 leaf chunk - logarithmic work. Adding a level multiplies capacity ~1,000× with no resharding, index freeze, or global rebalance.

The L0-L5 capacities, ~1,000× per-level fan-out, 10% K-means sample, and reclustering design are from VAST's published vector index, designed for sub-200 ms search at trillion-vector scale. Circles show a sample of 19 children each; the K-means view runs on toy 2-D data.

Because the tree grows by adding levels rather than shards, scaling from billions to trillions needs no resharding, index freeze, or global rebalance, and continuous ingest never blocks queries. The documented ceiling is 256T rows (2⁴⁸) per table - the 50-billion proof point is roughly 0.02% of it.

Inside the walk: beam width & the ratio rule

The walk isn't a fixed top-k at each level. VAST carries a beam of the closest centroids and widens it toward the leaves, bounded by a per-level min/max and driven by a ratio rule: it keeps expanding a level while the distance to the farthest selected centroid stays within a set ratio (default 1.8) of the closest. The whole walk runs on 8-bit (int8) scalar-quantized codes; only the final leaf candidates are fetched at full precision and reranked. Tune the ratio and watch the work-versus-recall trade.

The beam, drawn by distance from the query

Every dot is a candidate centroid, placed by its distance from the query. The beam is whatever falls inside the dashed ring at ratio × d(closest), clamped to the level's min/max. Drag the ring or the slider.

1.80
1.0 · narrow, less work3.0 · wide, more work
1.80 × dquery

L2: 81 selected - 81 inside the ring, within 20-150.

in the beaminside ring, over maxoutside ring, added by minprunedd(closest)

Beam per level, on a 1,000-cell fan-out

↓ 4 × 1,000 children = 4.0K scored at L3

↓ 4 × 1,000 children = 4.0K scored at L2

↓ 81 × 1,000 children = 81K scored at L1

Quantized scans (int8)

90K

centroids + leaf vectors compared

Full-precision reranks

189

originals fetched at the leaf only

Corpus scanned

9.0e-6%

of 1.0T vectors

Work90K (range 29K-159K)
Recalldirection only - no number implied

A wider beam explores more branches, so a true neighbor just across a cluster boundary is less likely to be pruned - at the cost of more scans. The walk still touches a vanishing fraction of the corpus, and the leaf rerank restores full precision. Admins set the ratio with vsetting; a session can override min_per_level / max_per_level at query time.

Illustrative model of a 1T-vector table with 1,000-way branching. Beam bounds and the 1.8 default ratio are VAST-documented; the dot fields are toy data, so beam counts show the rule, not a measured index.

Benchmarked: 11× faster, ~91% lower cost

The architecture differences show up in numbers. On a 1-billion-vector, 128-dimension benchmark at 99% recall, VAST delivers roughly 11× the throughput of a leading disk-based shared-nothing system, at about 91% lower cost per thousand searches, because the vectors no longer have to sit in expensive RAM. At 50 billion vectors - a scale the comparison system was not tested at - VAST sustains over a thousand queries per second on just 8 CNodes. To see why RAM is the bill at this scale, size an index in the hyperscale memory calculator - a trillion-vector index can live on NVMe but not in DRAM.

VAST vector search, in benchmarks

More than 11× the throughput at 1 billion vectors (128-dim, 99% recall), about 91% lower cost per search, and a result at 50 billion vectors where the comparison system was not tested.

Throughput at 1B vectors
higher is better
Milvus 2.6disk index (AISAQ)~89 QPS
VASTsame test~1,000 QPS

1,000 ÷ 89 ≈ 11.2 Milvus-sized blocks of queries per second.

Cost per 1,000 searches
lower is better
Sharded baselineper 1,000~$0.033
VASTper 1,000~$0.003

The baseline bill is 11 VAST-sized bills: $0.033 ÷ $0.003 = 11, so about 91% lower.

Tested envelope - throughput vs dataset size
benchmark points only · tap a point
04008001,200QPS1B10B50Bvectors in the index (log scale)Milvus: not testedat 50B (scaling limits)~1,000~891,026 · 375 thr606 · 75 thr
VASTMilvus 2.6 (disk / AISAQ)no result

VAST at 50B vectors on 8 CNodes, 375 threads: 1,026 QPS at ~317ms.

Points are separate test runs, not a measured curve - nothing is drawn between them. The two 50B results show the concurrency trade: more threads, more total QPS, higher latency per query.

99%
recall at 1B vectors
50B
vectors on 8 CNodes
~1M/s
ingest via Parquet
~290K/s
background clustering

All figures are from benchmarks: 1B vectors at 128-dim and 99% recall, ~1,000 QPS vs Milvus 2.6 disk (AISAQ) ~89 QPS; ~$0.003 vs ~$0.033 per 1,000 searches; 50B vectors on 8 CNodes, 1,026 QPS at ~317ms (375 threads) and 606 QPS at ~114ms (75 threads); ingest ~1M vectors/s via Parquet; background clustering ~290K vectors/s.

Indexing at write, GPU search, and hybrid vector + SQL queries

Indexes are built at write time - zone maps, the vector index, and secondary indexes are all maintained on ingest, so there is no separate “build the index” batch job. Search runs on CPU or GPU algorithms, including NVIDIA CAGRA, without caching the whole index in GPU memory. Because data is laid out in small columnar chunks, the engine can prune irrelevant blocks early instead of scanning everything.

The payoff for co-location is the hybrid query: vector similarity and SQL metadata filters in a single statement - for example, the nearest neighbors to a query embedding, filtered to a brand and a price range, without round-tripping between two systems.

It is genuinely SQL, not a bolt-on API. Vectors are a first-class column (fixed-dimension PyArrow list types), and built-in distance functions - array_distance() for Euclidean and array_cosine_distance() for cosine similarity - drop straight into a query, so you WHERE-filter on metadata, ORDER BY similarity, and LIMIT to the top matches in one statement. Apps reach it over the Arrow stack (ADBC driver, the vastdb Python SDK), which keeps vectors and results in columnar form end to end.

GPUs accelerate more than model inference here - they speed up the whole vector pipeline. Index builds run about 4× faster on GPU (10 hours down to 2.5 hours) using NVIDIA's cuVS library, while ingesting on the order of a million vectors per second. For SQL, Sirius - an open-source GPU SQL engine from the University of Wisconsin-Madison, in technical preview on VAST - ran up to 20× faster than DuckDB on NVIDIA RTX PRO 6000 in benchmarks (the VAST DataBase page puts that next to its other Sirius figures).

GPUs across the whole vector pipeline

Not just inference: GPUs speed up ingest, index build and SQL query. Pick a stage to race GPU against CPU on the same clock.

1. Ingest
~1M vectors/s
1024-dim vectors
Index build on the clock
1/21
CPU build
0% built
GPU build
0% built
0h2.5h5h7.5h10h
0.0h
clock
0%
CPU
0%
GPU

Both builds start together on the same data.

In benchmarks: GPU index build ~4x faster using NVIDIA cuVS (10h to 2.5h); Sirius GPU SQL up to 20x faster than DuckDB on NVIDIA RTX PRO 6000; ~1M vectors/s ingest (1024-dim). The SQL race is drawn at the best case (20x); real speedups vary by query.

One engine, every retrieval technique

Production retrieval usually means bolting extra techniques onto a vector DB - hybrid fusion, filtered (ACORN), quantized rescoring, MMR diversity, GPU index builds. Because VAST stores vectors as a SQL column, most of these collapse into a single query path. Compare them technique by technique.

VAST vs conventional vector-DB techniques

The retrieval tricks teams bolt onto vector DBs, drawn as pipelines. Pick a technique and watch the conventional stack collapse into one query path.

Native - built into the engineNot needed - the problem doesn't ariseApp layer - still lives in your app
NativeHybrid search & score fusion

Combine semantic (vector) and lexical (BM25) signals into one ranking.

Conventional vector DB3 steps · 3 systems
Vector DB
ANN search
Search engine
BM25 search
App code
RRF / DBSF merge
results

Run a vector search and a lexical search separately, then fuse the lists in the app with RRF or DBSF (Qdrant, Weaviate).

On VAST3 steps · one query

one SQL query

VAST engine
vector similarity
VAST engine
lexical match
VAST engine
weighted fusion in SQL
results

Both axes run inside one SQL query; weighted fusion is expressible in SQL - no app-level merge. Default top-1000 candidate pool (vs ~100 on AWS S3 Vectors) gives 10× to rerank over.

Index families, head to head
practical scale, log axis · tap a family
1B100B10T

Solid = practical or proven scale; dashed = VAST's documented design ceiling (2⁴⁸ rows per table).

Index location
HNSW: RAM (graph + vectors)
VAST: NVMe flash (vectors + index)
Memory cost
HNSW: ≈4d + 8M B/vec (~91TB at 20B×1024d)
VAST: Near-zero per vector (only centroids in RAM)
Practical scale
HNSW: ≈1B / node
VAST: 50B proven · 256T ceiling
Ingest cost
HNSW: Hours-days graph rebuild
VAST: Fast unclustered landing table + background reclustering
Filter / metadata
HNSW: Post-filter or ACORN
VAST: Native SQL push-down (any column)
Best for
HNSW: <1B, high recall
VAST: Trillion-scale RAG, real-time updates

In this comparison, only hierarchical K-means scales beyond one node without sharding, and only its memory cost stays flat as the corpus grows.

What you can configure

The vector column and its index expose a small, typed set of knobs at schema, index, and query level.

Every knob, pinned to what it controls

One vector table and one hybrid query. Tap a knob to see where it lives - at schema, index, or query level. Tap the distance metric again to cycle its SQL function.

Schema level

Index level

Query level

Table docs

idint64
brandstring
price, title…typed metadata
embeddingfloat32 × dim

room for up to 31,936 typed metadata columns beside the vector

one embedding per row, up to 4,096 components, fixed when the column is created

secondary index sort keys:idbrandkey 3key 4max 4

Hybrid query (illustrative)

SELECT id, title
FROM docs
WHERE brand = 'acme'
  AND price < 50
ORDER BY
  array_cosine_distance(
    embedding, :query_vec)
LIMIT 10;
cluster scan for this session:adaptive (admin ratio default)or set per session
Schema · Vector dimensionup to 4,096set at column creation

Manhattan distance is not currently supported (under consideration). Configuration details are VAST-documented. The table, column names and query are illustrative.

New in VAST AI OS 5.5

VAST AI OS 5.5, generally available since August 6, 2026, ships the index this page describes as the Hyperscale Vector Index: the hierarchical, quantized, DASE-persisted index above, with memory-bounded queries at trillion-vector scale - so latency plateaus as the dataset grows instead of degrading. The release also adds 50+ aggregation functions to the native query engine, and SQL-native partitioning: a partition layout you can declare on a table when it helps, on top of the per-chunk pruning that lets most VAST DataBase tables skip partition design entirely. In benchmarks against Iceberg-based tables, partitioned tables ran highly selective queries more than 60% faster and updates 200× faster, with about 20% lower aggregate runtime in one manufacturing workload.

VAST AI OS 5.5 · generally available Aug 2026

Vector search and SQL analytics, upgraded

Three additions that shipped in 5.5: a hierarchical vector index, a larger query engine, and SQL-native partitioning. Walk a query through the index to see why it reads so little.

Hyperscale Vector Indexhierarchical, quantized, DASE-persisted
Native Query Engine50+ new aggregation functions
SQL-native partitioningbenchmarked against Iceberg tables
Inside the Hyperscale Vector Index
fragments read: 0 / 21
1/5

At rest: the whole index is persisted on DASE flash - nothing is pinned in RAM.

L2
L1
L0
DASE · NVMe flash - the index is persisted here, not pinned in RAM
quantized centroid fragment L0 fragment, full-precision vectors read by this query
Native Query Engine
50+

new aggregation functions for in-place SQL analytics

one cell per function - 50 shown, the release adds more

SQL-native partitioning

vs Iceberg-based tables, in benchmarks

  • >60%faster on highly selective queries
  • 200×faster updates
  • ~20%lower aggregate runtime in one manufacturing workload

The index tree is a schematic - real levels hold far more fragments, and the 7-of-21 count is illustrative.

End-to-end RAG on one platform

Because vectors are native, the whole retrieval loop runs on a single platform: ingest the source data, embed it with serverless functions or NVIDIA NIM, index it at write, and retrieve with hybrid queries - no ETL drift between the vectors and the documents they came from. That is the difference between a vector store you have to babysit and one that stays consistent by construction.

Key takeaways

In one line

Vectors live as a first-class SQL column next to source data on VAST, replacing bolt-on vector databases with atomic writes and a single hybrid query.

Key points

  • Inserting a row and its vector is one atomic transaction - strong consistency, unlike external vector DBs' ETL sync and two-step retrieval.
  • A hierarchical, distance-clustered index (not an in-RAM graph) prunes about 1,000x the candidates per level, so search stays logarithmic without sharding.
  • On a 1-billion-vector, 128-dimension, 99%-recall benchmark, VAST delivered roughly 11x the throughput of a leading disk-based system and about 91% lower cost per 1,000 searches.
  • GPU index builds run about 4x faster via NVIDIA cuVS (10 hours down to 2.5 hours), and the open-source Sirius GPU SQL engine ran up to 20x faster than DuckDB in benchmarks.

Questions to explore

  1. 01How many separate systems (vector DB, warehouse, ETL job) keep your embeddings in sync with source data?
  2. 02How large could your vector corpus grow, and would an in-memory graph index like HNSW still fit in RAM at that scale?
  3. 03Do your retrieval queries need to combine vector similarity with metadata filters (tenant, date, price) in a single request?

Common questions

How does search stay fast at trillion-vector scale?
Each index level prunes ~1,000x the candidates and the DASE design needs no sharding; the design target is sub-200 ms search at trillion-vector scale with no index freezes, and in benchmarks 50 billion vectors on 8 CNodes answered in about 114-317 ms.
What happens to recall when VAST prunes 1,000x the candidates away at each level?
The index keeps a "beam" of the closest centroids per level, widened by a tunable ratio rule (default 1.8); only final candidates are fetched at full precision and reranked, reaching 96-99% recall at 50 billion vectors.
Which distance metrics and configuration limits apply?
Euclidean, cosine distance, and negative inner product are built into SQL; vectors support up to 4,096 dimensions and up to 31,936 metadata columns; Manhattan distance is not currently supported.