VAST Data

VAST DataBase

Intermediate

A columnar table format - like Iceberg or Delta, but native to the platform - that runs analytics directly on exabyte-scale storage, with no separate Parquet files or metastore.

A columnar table format built into storage

The lakehouse pattern stores tables as files on object storage, then layers a table format (, Delta, Hudi) plus a separate metastore on top to add schema, , and time-travel.

VAST DataBase is a columnar table format too - analytical, schema-aware, ACID - but it is native to the platform: the table format, its metadata, and the data all live in DASE storage as one thing. The three properties below are what that buys; the “Beyond the lakehouse” section further down compares it with Iceberg and Delta point by point.

Three properties, each one flip away from the alternative

Flip each panel to compare with the usual approach. Table sizes are illustrative.

Columnar

SELECT SUM(c2)

Cells read 6 of 24

Tables stored as columns for fast analytical scans with predicate & projection pushdown.

Mutable & ACID

table (id · c2)

101 · 5
102 · 1
103 · 9

Run the UPDATE.

Real inserts / updates / deletes with ACID guarantees - not append-only like classic lake formats.

One copy

SQLPythonAgents
live tablesource of truth

Copies 1 · ETL hops 0

No metastore, no ETL hops - query the live source of truth in place.

The write path: Persistent Write Buffer → Low-Cost Flash

VAST DataBase unifies transactional and analytical workloads in one table format. It writes in rows, perfect for transactions, and stores in columns, optimized for analytics - and it is fully ACID compliant. Writes land row-by-row into a Persistent Write Buffer (Storage-Class Memory), a low-latency tier that removes write hotspots, then background processes reshape those records into small ~32 KB columnar chunks on low-cost QLC flash. The result is low-latency ingest and fast columnar scans from the same table, without the read-versus-write tradeoff that forces most shops to run a separate OLTP database and analytical warehouse.

Compute is stateless over NVMe-oF: any node serves any query, with no sharding or partition owners, and ACID is enforced through decentralized object- and file-level locks at exabyte / trillion-row scale. Because the columnarization runs off the critical write path, ingest never pays the column-store tax.

The transpose - write in rows, store in columns

Step through one batch of writes. Color marks the column, so watch each row's cells leave the buffer and regroup into one chunk per column on flash.

1/3Row ingest
CNode 1CNode 2CNode 3any nodetakes writesWRITE BUFFER · SCMuser_idc2c3countryLOW-COST FLASH · QLCuser_idmin/max 101–106c2min/max 1–9c3min/max a–fcountrymin/max BR–UScolumnarizein background101102103104105106519372abcdefUSDEJPFRBRIN✓✓✓✓✓✓
1 · Row ingest

Writes land in the Persistent Write Buffer

Inserts, updates and deletes hit a low-latency, persistent write buffer (Storage-Class Memory) row-by-row - ideal for transactions. Every CNode can absorb a write, so there are no partition owners and no write hotspots: writes commit immediately, durably, and become queryable.

color = column chunk footer: min/max + count, held in SCM✓ committed and queryable

Rows = transactions, columns = analytics - one fully ACID table format serving both, with atomic, unified permissions across tables, files and objects.

Six rows and four columns stand in for real chunks; values are illustrative.

Crucially, tables, files and objects all live in one namespace with atomic, unified permissions - so a table row, a Parquet object and a raw file are governed and transacted together rather than scattered across systems.

Inside the column: a 32 KB chunk that prunes itself

Each column lands on flash as a ~32 KB chunk with a footer of metadata - sorted projections, customer-defined sort keys, and per-chunk statistics (min/max and count) - held in SCM, with no separate metadata manager. A chunk is roughly 1/4000th the size of a Parquet row group, so min/max pruning skips almost everything for a selective query: the engine reads one chunk instead of scanning a whole row group. That fine granularity is also why table updates stay simple - there is no partitioning to design, no pruning or vacuuming to run, and cross-table works at scale.

Anatomy of a 32 KB column

Column chunk~32 KB on QLC flash

Footer metadata - in SCM

ProjectionsSort keysMin/max()Count

Self-describing, no metadata manager

Each chunk operates somewhat like Parquet, but the statistics, sort keys and projections travel with the data in SCM - there is no separate metadata service to scale or keep consistent. Customers can add their own sort keys for distributed sorting, projection and filtering, with index support built in.

Fine-grained: ~1/4000th of a Parquet row group

One Parquet row groupto scale
One 32 KB chunk - the hairline at the left edge. At ~1/4000th of the bar it is narrower than a pixel here.

Fire a query. Min/max statistics in each chunk's footer prune everything whose range cannot contain a match, so an exact lookup reads a single 32 KB chunk and a range query reads only the chunks its range overlaps - never a whole row group.

chunks scanned: 12 / 12

32 KB

0-999

32 KB

1000-1999

32 KB

2000-2999

32 KB

3000-3999

32 KB

4000-4999

32 KB

5000-5999

32 KB

6000-6999

32 KB

7000-7999

32 KB

8000-8999

32 KB

9000-9999

32 KB

10000-10999

32 KB

11000-11999

A Parquet row group is coarse - a selective query reads the whole group. VAST's 32 KB chunk is ~1/4000th that size, so selective queries touch a fraction of the data. Pick a query above to see min/max pruning skip the rest.

No partitioning toil

Non-partitioned datasets scan as fast as partitioned Parquet or Iceberg. Per-chunk min/max does the pruning, so there is no partition layout to design, maintain, or get wrong.

No pruning, no vacuuming

Updates rewrite only the affected 32 KB chunks - table updates stay simple and fast. There are no snapshot rewrites and no compaction or vacuum jobs to chase.

Cross-table CDC at scale

Fine-grained, mutable chunks make change data capture across tables simple - without the ETL limitations of legacy lake formats.

Simplified from VAST's published DataBase design (~32 KB columnar chunks, roughly 1/4000th the size of a Parquet row group, with per-chunk min/max statistics for pruning). Chunk ranges shown are illustrative.

Beyond the lakehouse

Iceberg, Delta Lake and Snowflake-style lakehouses pair Parquet files with a table format and a separate metastore - three loosely-coupled pieces to keep consistent. VAST DB is itself a native table format: metadata lives with the data, so there is no metastore bottleneck. Streaming and CDC into Iceberg / Delta spawn countless tiny Parquet files that constantly need compaction; VAST columnarizes off the critical write path into fixed 32 KB chunks on flash, which sidesteps the small-file problem entirely. And instead of append-only writes with snapshot rewrites, VAST does real-time mutable inserts, updates and deletes with query-in-place and atomic multi-table transactions - models and agents scan columnar tables directly on the source of truth, with no ETL copy into a separate warehouse.

Two table architectures, the same writes

Stream data in, update a row or compact, and watch both sides. Hover or tap a capability in the table below to see which part of each diagram it describes.

Iceberg / Delta

metadata tree + Parquet files

Catalog / metastoreseparate servicemetadata v1s1manifest listmanifestmanifestParquet data filesfilefilefilefileno small files yet

4

data files

1

snapshots

0

compactions

VAST DB

flat table, metadata with the data

Table (native format)no metastoreWrite buffer · SCMrows land here, then move to flashFlash · ~32 KB chunks, each with its own footer

0

chunks rewritten

0

rows in buffer

0

compactions

Press a button - both tables get the same writes.

VAST DB
Iceberg / Delta
Separate metastore
None - native table format, metadata lives with the data
External metastore (Hive/Glue/catalog) is a coordination bottleneck
Small-file problem
~32 KB columnar chunking on flash - no compaction debt
Streaming / CDC spawns countless tiny Parquet files needing compaction
In-place updates & deletes
Mutable rows - real UPDATE / DELETE at the storage layer
Append-only + snapshot rewrites; merge-on-read amplification
Real-time ingest
Row-by-row into SCM buffer, immediately queryable
Micro-batch commits; freshness gated by file/snapshot cadence
Atomic multi-table transactions
ACID across tables, files and objects in one namespace
Single-table snapshot isolation; cross-table atomicity not native
Query in place (no warehouse copy)
Columnar tables queried where they live - no warehouse copy
Often copy/ELT into a separate warehouse for fast analytics

Benchmark: point lookup, 10-billion-row table

VAST DB
120 ms
Iceberg
~1.6 s

Linear scale, lower is better.

Storage-layer predicate pushdown and hierarchical sorted projections give roughly O(log n) lookups, versus Iceberg scanning row groups.

Lookup figures come from a benchmark on a specific configuration. The architectural differences above (no metastore, mutability, no small files) are design properties. File and snapshot counts in the diagrams are illustrative.

Query engines: native, federated & pushdown

VAST has its own native query engine that runs in-place on the CNodes, executing SQL and vector search directly against the columnar tables, with heavy aggregations GPU-accelerated by Sirius - an open-source GPU SQL engine built on NVIDIA - on NVIDIA GPUs, in technical preview. Open engines such as Trino and Spark can also run natively on VAST serverless compute, or attach externally via push-down plugins that ship predicates and projections down to storage. BI tools reach the data through those SQL engines or via Arrow Flight SQL, and the Python SDK gives programmatic access - all of them meet at the same tables, shown in the diagram below.

Every engine, one table

Whatever tool a team already uses, it points at the same tables. SQL engines, the Python SDK, streaming events, and bulk imports all read and write one copy of the data over NVMe-over-Fabrics, under one governance model. Pick an engine or an ingest path in the diagram to see how it connects and where its data goes.

Every engine, one table - and the filter travels to the data

Query engines on one side, ingest paths on the other, one copy of the table in the middle. Pick a query engine to send c2 > 2 and [user_id, c2] down to storage; pick an ingest path to see where writes land.

CNODES stateless · any node serves any queryVAST query engine + Siriuswrite buffer · SCMuser_idc2c3country0–23–60–14–9one copy · over NVMe-oF on SCM + QLC flashTrinoApache SparkArrow Flight SQLPython SDKKafka eventsParquet importerSDK insert

Pushed down to storage

predicate c2 > 2
projection [user_id, c2]

4 of 16 chunks readOnly 2 of 4 columns are touched, and min/max skips the c2 chunks that cannot match.

Thin green line: the result stream. Wide grey band: what a full-table scan would ship.

VAST query engine

Native · runs on the CNodes

VAST's own query engine runs in place on the CNodes, executing SQL and vector search directly against the columnar tables with predicate and projection pushdown - no external engine to deploy. Heavy aggregations are GPU-accelerated by Sirius, an open-source GPU SQL engine built on NVIDIA cuDF, on NVIDIA GPUs (in technical preview).

Typical use: In-platform SQL analytics and vector retrieval with no separate query cluster; GPU acceleration for large aggregations, joins and statistical functions.

runs on the CNodes / serverlessexternal cluster, push-down pluginclient protocol / SDKingest chunk read pruned by min/max

Every path reads and writes the same columnar tables in place - no copies, no separate warehouse, one governance model - and every query path pushes predicates and projections down to storage.

Beyond these, VAST DB connects through Apache NiFi, Flink, Beam, Ray / Daft, Dremio, and LangGraph checkpoint storage. Chunk counts and ranges are illustrative.

GPU-accelerated SQL with Sirius

Sirius is an open-source GPU SQL engine led by the University of Wisconsin-Madison with NVIDIA support - it accelerates DuckDB by plugging in through the Substrait query-plan format and running relational operators on NVIDIA cuDF, with no query rewrites. VAST integrates it with the DataBase so heavy aggregations execute on GPUs at the compute layer; the integration has been in technical preview since August 2026. VAST's columnar layout and predicate / projection pushdown cut how much data the GPU has to touch in the first place; Sirius speeds up the work that remains.

Two Sirius figures come from two different benchmarks. In early benchmarks, VAST DataBase with Sirius took up to 44% less query time and up to 80% less query cost - the race below. Separately, Sirius ran up to 20× faster than DuckDB on an NVIDIA RTX PRO 6000 GPU. They measure different setups, so read them side by side rather than as one number.

CPU vs GPU: the same analytical query, relative time and cost

VAST pushes predicates and projections down at the storage layer - the same scan in both lanes - then hands the heavy aggregation to Sirius on the GPU. Bars are relative to the CPU run (= 100%).

Query time · relative

CPU SQL engine100%
scan
aggregate on CPU
Sirius (GPU, cuDF)56%
scan
GPU
up to 44% less time
0%25%50%75%100%

Query cost · relative

CPU
100%
Sirius
20% - up to 80% less cost
storage scan with pushdown (same in both) hand-off to the GPU

Pushdown shrinks what reaches the engine; Sirius makes the work that remains finish sooner and cost less.

The 44% time / 80% cost figures come from early benchmarks of VAST DataBase + Sirius on NVIDIA GPUs. Only the lane totals come from the benchmark; the split between stages is illustrative.

Don't confuse Sirius with KV cache

Sirius accelerates analytics (SQL on GPUs). It is not the KV-cache / inference story - that is NVIDIA Context Memory Storage (CMX) on the STX architecture, a separate piece of the VAST + NVIDIA stack. Two different GPUs-meet-data problems: Sirius is for queries, CMX is for inference context. See NVIDIA STX & CMX →

Sirius is one piece of VAST's end-to-end accelerated stack with NVIDIA:

Which GPU piece answers which problem

Two pieces work on the platform's data; CMX serves inference context and belongs to a separate architecture. Tap a row.

GPUs meet the data - queries and search

A different problem - inference context (STX)

Sirius

GPU SQL execution (cuDF) for analytics on the DataBase.

Apache Arrow & Arrow Flight: zero-copy at wire speed

is a standardized columnar in-memory format. Arrow Flight transports Arrow record batches over gRPC with parallel streaming and zero-copy, avoiding the ODBC / JDBC serialization overhead commonly cited at 60–90%. VAST's SDK is Arrow-native: queries return a streaming pyarrow.RecordBatchReader, so data flows from flash to your dataframe without a row-by-row reserialization tax.

One result set, two ways to your dataframe

Both lanes move on the same clock. Dashed orange stages are format conversions. Schematic - it counts stages, not milliseconds.

1/7Play to race the two paths

ODBC / JDBC driver

stage 1 of 6

  1. Columnar on flash
  2. Pivot to rows
  3. Serialize row by row
  4. One stream over the wire
  5. Deserialize rows
  6. Rebuild columns → DataFrame

Arrow Flight + the vastdb SDK

stage 1 of 3

Conversions so far, driver
0
Conversions, Arrow
0
Streams, Arrow
parallel

Tap an Arrow stage

Zero-copy over gRPC

Arrow Flight streams record batches in parallel without ODBC / JDBC serialization.

The 60-90% figure above is the commonly cited ODBC / JDBC serialization overhead; this diagram does not measure it.

The Python SDK in practice

The vastdb package (“vast-py”) installs with pip install vastdb. You connect with an endpoint plus access / secret keys, then run operations inside a session.transaction() block. The hierarchy is bucket → schema → table (PyArrow schemas): table.insert(pyarrow_table) writes rows, and table.select(...) returns a streaming reader with predicate and projection pushdown expressed via Ibis (e.g. (_.c2 > 2) & _.c3.isnull()).

pip install vastdb  ·  quickstart.py
import pyarrow as paimport vastdbfrom ibis import _ # 1. Connect: endpoint + access / secret keyssession = vastdb.connect(    endpoint="http://vip-pool.vast.example.com",    access="ACCESS_KEY",    secret="SECRET_KEY",) # 2. Everything runs inside an ACID transactionwith session.transaction() as tx:    # bucket -> schema -> table    schema = tx.bucket("ml").schema("features")    table = schema.table("user_events")     # 3. Insert a PyArrow table (row ingest -> SCM buffer)    batch = pa.table({        "user_id": [101, 102, 103],        "c2": [5, 1, 9],          # int column        "c3": ["a", None, "c"],   # nullable string    })    table.insert(batch)     # 4. Streaming read with predicate + projection pushdown (Ibis)    reader = table.select(        columns=["user_id", "c2"],        predicate=(_.c2 > 2) & _.c3.isnull(),    )     # 5. reader is a pyarrow.RecordBatchReader -> pandas    df = reader.read_all().to_pandas()    print(df)
quickstart.pyendpoint + keys1ACID transaction · CNodes2SCM write buffer3user_idc2c3c2 > 2 · [user_id, c2]4RecordBatchReader → DataFrame5

4 · Pushdown read

select() ships the predicate and the column list to storage, so only matching chunks of user_id and c2 are read.

See pushdown in the hub ↑

Arrow-native end to end: reads stream as a pyarrow.RecordBatchReader, so predicate and projection filtering happen in storage before any bytes cross the wire. Hover or tap a numbered block to see what it does.

Beyond the basics: semi-sorted projections for fast secondary lookups, S3 Parquet import without a client-side copy, the VAST Catalog (query the filesystem itself as a table), and snapshots for point-in-time reads.

Serving features & metadata to AI pipelines

Because the table format holds fresh, mutable data and large-scale analytics together, AI pipelines read current, consistent data with no ETL hops. Features and metadata are served in real time from the source of truth: an agent can update a record and immediately query it, training jobs read the same tables that production writes, and the Event Broker exposes streams as queryable tables - closing the loop between operational and analytical data.

One live table, four AI consumers, no ETL between them

Follow one record through the pipeline. Every consumer touches the same table, so there is no copy to fall behind. The record is illustrative.

1/4Play to follow one record

orders · live table

one copy

idcustomerflag
1041acme-
1042globex-
1043initech-

tables, files and objects in one namespace, under one set of permissions

An order event arrives on the Event Broker and lands as a row - queryable right away.

Streams as tables: The Event Broker surfaces event streams as queryable tables for online + offline use.

Key takeaways

In one line

VAST DataBase is a native, ACID columnar table format built into storage itself - no separate Parquet files, metastore, or ETL copy.

Key points

  • Writes land row-by-row in a Persistent Write Buffer (SCM) for low-latency ingest, then are reshaped into ~32 KB columnar chunks on QLC flash.
  • Each 32 KB chunk carries its own footer (sort keys, min/max, count) in SCM - about 1/4000th the size of a Parquet row group.
  • Heavy aggregations run GPU-accelerated via Sirius, an open-source GPU SQL engine from the University of Wisconsin-Madison built on NVIDIA cuDF, on NVIDIA GPUs (technical preview) - up to 44% less time and 80% less cost in early benchmarks.
  • Trino, Spark, Arrow Flight SQL and the Python SDK all query the same tables, with predicate and projection pushdown.

Questions to explore

  1. 01How many tiny Parquet files and compaction jobs does your streaming or CDC pipeline generate?
  2. 02How much of your analytics stack copies data from an operational database into a separate warehouse today?
  3. 03Would your BI tools or agents benefit from querying live, mutable data instead of a batch-refreshed copy?

Common questions

How does this avoid the small-file problem that plagues Iceberg or Delta under streaming ingest?
Columnarization happens asynchronously, off the critical write path, into fixed ~32 KB chunks on flash, so there's no pile-up of tiny Parquet files and no compaction jobs to run.
How fast are point lookups on large tables?
In a benchmark on a 10-billion-row table, a VAST DB point lookup took 120 ms versus ~1.6 s for Iceberg at equal concurrency.
Can it really do real-time updates and deletes, unlike typical lakehouse formats?
Yes - it supports mutable inserts, updates and deletes with ACID guarantees, unlike append-only lakehouse formats that rely on snapshot rewrites and merge-on-read.