VAST Data
DataEngine & Event Broker
IntermediateBring compute to the data - serverless functions triggered by events, on a Kafka-compatible broker built into the platform.
Bring compute to the data
The default pattern for processing data is backwards: data lands in storage, then a separate pipeline copies it out to a compute cluster, transforms it, and writes results back somewhere else. Every hop costs money, adds latency, and creates a stale copy. For AI this means embeddings and enrichment lag behind ingest, so the context your models retrieve is a batch window out of date.
inverts the model: run the function where the data already is. No ETL job to schedule, no second cluster to feed, no extra copy to keep consistent. The compute is event-driven and stateless, so it only runs when new data actually arrives.
How stale is the data your model sees?
The same objects land over time. A pipeline or batch job only touches them at its next run, so each one waits - hatched - before it is usable. Running compute per event closes the gap.
Move data → compute
Scheduled batch
Compute → data
Takeaway: moving data to compute costs N hops, N copies and N failure points; a schedule leaves context one window stale. Functions that run in place on each event are fresh by construction.
Schematic: arrival times, the sweep period and job lengths are illustrative.
DataEngine: serverless functions & triggers
GADataEngine is an event-driven serverless compute fabric that runs functions on data where it lands. Serverless functions & triggers went GA on Nov 11, 2025. It has three execution modes that compose into one in-situ pipeline.
Three execution modes, three lifetimes
The same stream of data events, seen by each mode. Pick one to see what it does - together they compose into one in-situ pipeline.
Takeaway: triggers turn each data event into an invocation, functions exist only while there is an event to handle, and engines stay up for heavier work - all of it next to the data.
Illustrative: event times and run lengths are made up to show the shapes.
Drop a file, watch the platform react
Each write to a view becomes an event on the broker. Stateless function pods scale out to consume the events, write their result into a table next to the source path, and scale back to zero when the queue is empty.
Use case: Auto-embedding on ingest
View /docs
S3 · NFS · SMB
0 recent objects
Event Broker topic
trigger: object created
0 events waiting
Function pods
serverless · fn: embed
0 running · peak 0
VAST DataBase table
docs_index
0 rows written
t=0Idle - no events, no pods running. Drop a file to start.
0
events published
0
pods running now
0
copies made
0
ETL jobs scheduled
Takeaway: compute runs where the data lands and only while there is work - per object, as it arrives, with no copy out to a separate cluster and no scheduled batch job.
Illustrative: tick length, pod start and run times, the 6-pod cap, file names and output values are made up to show the shape of the flow.
A function's lifecycle: code → container → trigger → in-situ run
You ship a function the same way you ship any container - the only new steps are registering it in VAST and binding it to a data event. Step through the lifecycle below.
From handler.py to an in-situ run
The same function changes shape at every step - watch the artifact strip as you step through.
Step 1 of 5: Write Python locally
Artifact at each step
Write Python locally
Author a standard Python function - your embedder, classifier, transformer, or enrichment logic. No proprietary SDK lock-in; ordinary code and libraries.
def handle(event):
obj = read(event.path)
vec = embed(obj) # NIM, OSS model…
return {"vector": vec}Code, names and the registry path are illustrative. Steps 3 and 4 show what you do, not exact command syntax.
Event Broker: Kafka-compatible, stateless, storage-backed
GAThe is a real-time event streaming engine built into the platform - it removes the need for a separate Kafka cluster and unifies streaming, storage, and analytics. Announced Feb 19, 2025 and available Mar 2025.
It is Kafka API-compatible, so existing producers and consumers work unchanged - but its brokers are stateless. There is no per-broker persistent state; durability and resilience come from VAST's global storage. Streams flow directly into tables for instant correlation across structured and unstructured data.
Add a broker, lose a broker
Run the same event on both designs. Stateful Kafka brokers have to copy partition data around; stateless VAST brokers just change who serves what, because the data already sits in shared storage.
Classic Kafka
stateful brokersB1
local disk
B2
local disk
B3
local disk
connector / ETL job
runs on a schedule
warehouse
0
replicas moved
Six partitions, each stored on two brokers' local disks (replication factor 2).
VAST Event Broker
stateless brokersB1
no local state
B2
no local state
B3
no local state
Shared VAST storage
every partition once, durable
0
partition data moved
Brokers only hold a serving assignment. All six partitions live once in shared VAST storage.
Throughput
10x+
vs Kafka on like-for-like hardware
500M+
messages/sec on the largest clusters
136M
msgs/sec in a separate benchmark blog
Announced Feb 19, 2025 and available Mar 2025; 136M msgs/sec comes from a separate benchmark. Results vary with hardware, message size, and configuration.
Takeaway: when state lives in shared storage, adding or losing a broker changes only who serves a partition - there is no partition data to move, and the stream is already a queryable table.
Illustrative: six partitions, replication factor 2, placements and transfer times are chosen to show the mechanism, not measured.
Streams as tables: unifying streaming + analytics
Because brokers are stateless and storage is the source of truth, a stream is not a separate artifact you later into a warehouse - it is a queryable table the moment it lands. The same record can be consumed as a Kafka message and joined against historical structured data in one query, with no copy in between.
One landing zone, two access patterns
Events written through the Kafka API land as rows of a table. A consumer streams them in order, and a query joins the newest event against months of history - the same rows, with no connector and no copy in between.
Producers
write via the Kafka API
VAST DataBase table
customer_events
Kafka consumer
reads the stream in order
waiting for events
SQL / agent query
joins live with history
SELECT … JOIN history
Months of prior data already sit in the table. Press play to stream new events in.
Takeaway: no connector to maintain and no schema drift between the stream and the warehouse - correlating live and historical data is bounded by query time, not an ingest pipeline window.
Illustrative: keys, dates and events are made up.
Why this matters for AI
Together, DataEngine and the Event Broker make data active. Embeddings and enrichment happen automatically as data arrives, so the context your and agent systems retrieve stays fresh in real time - not as of the last batch run. The embeddings land as vectors next to their source rows in VectorStore; see RAG & Vector Search for how that retrieval works. Data that starts outside VAST gets here through SyncEngine, which copies it onto the platform, where DataEngine functions can then process it like any other write.
From “data written” to “data usable”
Follow one new document through the platform. The Event Broker carries the trigger; DataEngine turns it into fresh context. Tap any stage.
Takeaway: the broker turns new data into an immediate trigger, and DataEngine turns that trigger into fresh embeddings - so retrieval never waits for a batch window.
Key takeaways
In one line
DataEngine runs serverless functions where data lands, triggered by a built-in Kafka-compatible Event Broker, so embeddings and enrichment happen as data arrives.
Key points
- Serverless functions & triggers went GA on Nov 11, 2025, with three modes - event triggers, Python functions, and containerized engines.
- Functions ship as standard Docker containers, run stateless and elastic, and execute only when triggered - no separate compute cluster to feed.
- The Event Broker (announced Feb 19, 2025, available Mar 2025) is Kafka API-compatible, but its brokers are stateless, with durability from VAST's global storage.
- Because brokers are stateless, a stream is a queryable VAST DataBase table the moment it lands, so live events can be joined against history in one query.
Questions to explore
- 01How much of your data processing still runs as scheduled batch jobs that leave embeddings or enrichment stale between runs?
- 02Do you run a separate Kafka cluster purely to move data into a warehouse before it can be queried?
- 03Could your RAG or agent systems benefit from correlating a live event stream against months of history in a single query?
Common questions
- How much throughput can the Event Broker handle?
- It delivers 10x+ the throughput of Kafka on like-for-like hardware and 500M+ messages/sec on the largest clusters; a separate benchmark measured 136M messages/sec. Results vary with hardware, message size, and configuration.
- Do existing Kafka producers and consumers need to be rewritten?
- No - it is Kafka API-compatible, so existing producers and consumers work unchanged; the difference is that brokers are stateless and durability comes from VAST's storage instead of broker-local replicas.
- What happens to a function when there's no data to process?
- Serverless functions run in stateless containers that execute only when triggered, scaling from 0 to N and back to 0 - so idle functions consume nothing.