VAST Data

DataEngine & Event Broker

Intermediate

Bring compute to the data - serverless functions triggered by events, on a Kafka-compatible broker built into the platform.

Bring compute to the data

The default pattern for processing data is backwards: data lands in storage, then a separate pipeline copies it out to a compute cluster, transforms it, and writes results back somewhere else. Every hop costs money, adds latency, and creates a stale copy. For AI this means embeddings and enrichment lag behind ingest, so the context your models retrieve is a batch window out of date.

inverts the model: run the function where the data already is. No ETL job to schedule, no second cluster to feed, no extra copy to keep consistent. The compute is event-driven and stateless, so it only runs when new data actually arrives.

How stale is the data your model sees?

The same objects land over time. A pipeline or batch job only touches them at its next run, so each one waits - hatched - before it is usable. Running compute per event closes the gap.

New data lands
objects written to storage
Move data → computeold way
copy out, transform, copy back
Scheduled batchold way
a cron / Spark job sweeps
Compute → dataDataEngine
a function per event, in place
time (schematic) →
1/25

Move data → compute

0of 0 landed objects not usable yet · 0 copies made

Scheduled batch

0of 0 landed objects not usable yet · waits for the schedule

Compute → data

0of 0 landed objects not usable yet · no copies, no schedule
Object landsStale - landed, not usable yetCopy out / back (a hop and a copy)Processed in place

Takeaway: moving data to compute costs N hops, N copies and N failure points; a schedule leaves context one window stale. Functions that run in place on each event are fresh by construction.

Schematic: arrival times, the sweep period and job lengths are illustrative.

DataEngine: serverless functions & triggers

GA

DataEngine is an event-driven serverless compute fabric that runs functions on data where it lands. Serverless functions & triggers went GA on Nov 11, 2025. It has three execution modes that compose into one in-situ pipeline.

Three execution modes, three lifetimes

The same stream of data events, seen by each mode. Pick one to see what it does - together they compose into one in-situ pipeline.

Event Triggers
an instant per event
Serverless Python Functions
short-lived, 0 → N → 0
00
Containerized Engines
long-lived
one long-lived engine
time (illustrative) →
Serverless Python FunctionsStandard Docker containers registered in VAST. They run in stateless containers that scale elastically and execute only when triggered.

Takeaway: triggers turn each data event into an invocation, functions exist only while there is an event to handle, and engines stay up for heavier work - all of it next to the data.

Illustrative: event times and run lengths are made up to show the shapes.

Drop a file, watch the platform react

Each write to a view becomes an event on the broker. Stateless function pods scale out to consume the events, write their result into a table next to the source path, and scale back to zero when the queue is empty.

Function does

Use case: Auto-embedding on ingest

1

View /docs

S3 · NFS · SMB

empty - drop a file

0 recent objects

2

Event Broker topic

trigger: object created

no events waiting

0 events waiting

3

Function pods

serverless · fn: embed

scaled to 0

0 running · peak 0

4

VAST DataBase table

docs_index

idpathvector
no rows yet

0 rows written

t=0Idle - no events, no pods running. Drop a file to start.

0

events published

0

pods running now

0

copies made

0

ETL jobs scheduled

Event tile (one per write)Function pod (stateless container)New table row

Takeaway: compute runs where the data lands and only while there is work - per object, as it arrives, with no copy out to a separate cluster and no scheduled batch job.

Illustrative: tick length, pod start and run times, the 6-pod cap, file names and output values are made up to show the shape of the flow.

A function's lifecycle: code → container → trigger → in-situ run

You ship a function the same way you ship any container - the only new steps are registering it in VAST and binding it to a data event. Step through the lifecycle below.

From handler.py to an in-situ run

The same function changes shape at every step - watch the artifact strip as you step through.

Step 1 of 5: Write Python locally

Artifact at each step

1

Write Python locally

Author a standard Python function - your embedder, classifier, transformer, or enrichment logic. No proprietary SDK lock-in; ordinary code and libraries.

illustrative code
def handle(event):
    obj = read(event.path)
    vec = embed(obj)   # NIM, OSS model…
    return {"vector": vec}
1/5

Code, names and the registry path are illustrative. Steps 3 and 4 show what you do, not exact command syntax.

Event Broker: Kafka-compatible, stateless, storage-backed

GA

The is a real-time event streaming engine built into the platform - it removes the need for a separate Kafka cluster and unifies streaming, storage, and analytics. Announced Feb 19, 2025 and available Mar 2025.

It is Kafka API-compatible, so existing producers and consumers work unchanged - but its brokers are stateless. There is no per-broker persistent state; durability and resilience come from VAST's global storage. Streams flow directly into tables for instant correlation across structured and unstructured data.

Add a broker, lose a broker

Run the same event on both designs. Stateful Kafka brokers have to copy partition data around; stateless VAST brokers just change who serves what, because the data already sits in shared storage.

Classic Kafka

stateful brokers
producerproducer

B1

P0P2P3P4

local disk

B2

P0P1P4P5

local disk

B3

P1P2P3P5

local disk

connector / ETL job

runs on a schedule

warehouse

0

replicas moved

Six partitions, each stored on two brokers' local disks (replication factor 2).

VAST Event Broker

stateless brokers
producerproducer

B1

P0P3

no local state

B2

P1P4

no local state

B3

P2P5

no local state

Shared VAST storage

every partition once, durable

P0P1P2P3P4P5
= DB table

0

partition data moved

Brokers only hold a serving assignment. All six partitions live once in shared VAST storage.

Leader replicaFollower replicaData being copiedLost with the brokerDown to one copyMoved / rebuilt copyServing assignment (no data)
Broker stateKafkaStateful brokers own partitions; data is pinned to disks per brokerVASTStateless brokers - no per-broker persistent state to manage
DurabilityKafkaReplication factor copies data across brokers (2–3×)VASTDurability comes from VAST global storage - no replica fan-out
ClusterKafkaSeparate Kafka cluster to size, patch, rebalanceVASTBuilt into the platform - no separate cluster to operate
Path to analyticsKafkaETL/connectors land streams into a warehouse before queryVASTStreams land directly as VAST DataBase tables - query immediately
RebalancingKafkaPartition reassignment moves data; slow and risky at scaleVASTNo data to move on broker change - state lives in storage
GA

Throughput

10x+

vs Kafka on like-for-like hardware

500M+

messages/sec on the largest clusters

136M

msgs/sec in a separate benchmark blog

Announced Feb 19, 2025 and available Mar 2025; 136M msgs/sec comes from a separate benchmark. Results vary with hardware, message size, and configuration.

Takeaway: when state lives in shared storage, adding or losing a broker changes only who serves a partition - there is no partition data to move, and the stream is already a queryable table.

Illustrative: six partitions, replication factor 2, placements and transfer times are chosen to show the mechanism, not measured.

Streams as tables: unifying streaming + analytics

Because brokers are stateless and storage is the source of truth, a stream is not a separate artifact you later into a warehouse - it is a queryable table the moment it lands. The same record can be consumed as a Kafka message and joined against historical structured data in one query, with no copy in between.

One landing zone, two access patterns

Events written through the Kafka API land as rows of a table. A consumer streams them in order, and a query joins the newest event against months of history - the same rows, with no connector and no copy in between.

Producers

write via the Kafka API

no new events yet

VAST DataBase table

customer_events

tskeyevent
Jun 02cust-17order
Jul 14cust-42return
Aug 09cust-08order
Sep 21cust-42order
history above · live stream below
(no live rows yet)

Kafka consumer

reads the stream in order

waiting for events

SQL / agent query

joins live with history

SELECT … JOIN history

1/5

Months of prior data already sit in the table. Press play to stream new events in.

Takeaway: no connector to maintain and no schema drift between the stream and the warehouse - correlating live and historical data is bounded by query time, not an ingest pipeline window.

Illustrative: keys, dates and events are made up.

Why this matters for AI

Together, DataEngine and the Event Broker make data active. Embeddings and enrichment happen automatically as data arrives, so the context your and agent systems retrieve stays fresh in real time - not as of the last batch run. The embeddings land as vectors next to their source rows in VectorStore; see RAG & Vector Search for how that retrieval works. Data that starts outside VAST gets here through SyncEngine, which copies it onto the platform, where DataEngine functions can then process it like any other write.

From “data written” to “data usable”

Follow one new document through the platform. The Event Broker carries the trigger; DataEngine turns it into fresh context. Tap any stage.

Data written → data usablestage 1 of 5
New data landsA new file, object or record is written to the platform. In a batch world it would now wait for the next pipeline run.
1/5

Takeaway: the broker turns new data into an immediate trigger, and DataEngine turns that trigger into fresh embeddings - so retrieval never waits for a batch window.

Key takeaways

In one line

DataEngine runs serverless functions where data lands, triggered by a built-in Kafka-compatible Event Broker, so embeddings and enrichment happen as data arrives.

Key points

  • Serverless functions & triggers went GA on Nov 11, 2025, with three modes - event triggers, Python functions, and containerized engines.
  • Functions ship as standard Docker containers, run stateless and elastic, and execute only when triggered - no separate compute cluster to feed.
  • The Event Broker (announced Feb 19, 2025, available Mar 2025) is Kafka API-compatible, but its brokers are stateless, with durability from VAST's global storage.
  • Because brokers are stateless, a stream is a queryable VAST DataBase table the moment it lands, so live events can be joined against history in one query.

Questions to explore

  1. 01How much of your data processing still runs as scheduled batch jobs that leave embeddings or enrichment stale between runs?
  2. 02Do you run a separate Kafka cluster purely to move data into a warehouse before it can be queried?
  3. 03Could your RAG or agent systems benefit from correlating a live event stream against months of history in a single query?

Common questions

How much throughput can the Event Broker handle?
It delivers 10x+ the throughput of Kafka on like-for-like hardware and 500M+ messages/sec on the largest clusters; a separate benchmark measured 136M messages/sec. Results vary with hardware, message size, and configuration.
Do existing Kafka producers and consumers need to be rewritten?
No - it is Kafka API-compatible, so existing producers and consumers work unchanged; the difference is that brokers are stateless and durability comes from VAST's storage instead of broker-local replicas.
What happens to a function when there's no data to process?
Serverless functions run in stateless containers that execute only when triggered, scaling from 0 to N and back to 0 - so idle functions consume nothing.