VAST Data
VAST DataSpace
IntermediateA global namespace that makes the same data addressable across edge, on-prem and every cloud - one source of truth, strictly consistent, with no copies to keep in sync.
The data-gravity problem
AI workloads now run wherever GPU capacity is available - on-prem, in one cloud this week and another the next, out at the edge where data is generated. But the data has gravity: it is large, it is expensive to move, and the usual answer is to copy whole datasets to every site that needs them. Those copies multiply storage and egress cost, and they drift out of date the moment one site writes.
VAST DataSpace is a - a data fabric that unifies many VAST clusters worldwide into one consistently addressable space. The same files, objects and tables are reachable from every site at local all-flash speed, and data moves only when it is actually needed.
What copying the data to every GPU site costs
Add the sites that need to run AI jobs. The usual answer is a full copy of the dataset at each one - watch the three problems grow together. Units are whole datasets, illustrative.
where the data lives
GPUs free this week
GPUs free next week
where data is generated
Data gravity
Timeline at the newest site: Cloud A
time → the GPUs idle through the orange part
Datasets are too large and costly to shuttle between sites on demand, so compute ends up waiting on staging.
Copy sprawl
Stored: 2× the dataset
Egress to fill the copies: 1× the dataset
Replicating to every location multiplies storage and egress cost - and every copy is a chance to read stale data.
Multi-site by default
2
separate views of the data
1
view the jobs need
1 copy can drift the moment one site writes.
Training and inference span on-prem, multiple clouds and the edge - they need one view of the data, not many.
One namespace, no copies
Instead of pushing a full copy of the dataset to every site, DataSpace keeps one authoritative copy in the namespace and lets each site stream only the bytes it touches - or sync metadata only, moving no data until it is read. Toggle the two models to see the difference in copies, consistency and data movement. Data that lives outside VAST today - on an old file share, in a cloud bucket or a collaboration app - gets into the namespace with SyncEngine: DataSpace connects VAST sites, SyncEngine brings outside data in.
Copy everywhere vs stream what you touch
Same four sites, same jobs. Watch what crosses the links - and the meter of data moved - under each model.
Data moved between sites
illustrative
full run, for reference
- Copies of the data
- 4×
- Consistency
- Stale risk
- Egress on sync
- Per copy
Traditional: replicate the whole dataset everywhere
Each site gets a full copy of the dataset. Storage cost multiplies with every location, egress is paid on every sync, and the copies drift out of date the moment one site writes - so jobs can read stale data.
Illustrative: the dataset is drawn as 8 blocks and each job touches 1-2 of them; the meter counts whole datasets moved, not measured traffic.
How global access stays consistent
VAST is blunt that “eventually consistent isn't consistent enough.” DataSpace uses an model with decentralized . Per global folder, one cluster is the Origin that holds the authoritative copy and the write lease; the others are Satellites that cache locally under read leases. Step through a remote write to see why an update in one site is visible everywhere - with no lag.
A remote write, message by message
Keys are leases: the Origin holds the write lease, each Satellite caches under a read lease. Pick which site writes, then step through what travels over the links.
One namespace, one authoritative copy
Per global folder, one cluster is the Origin - it holds the authoritative copy and the write lease. Every other site is a Satellite that caches the data locally on flash under a read lease, so the same file is addressable everywhere at local speed.
Strictly consistent - decentralized read/write leases make an update in one site visible everywhere, with no eventual-consistency lag across clusters.
Schematic: message order is shown, not timings or distances.
Roles are per-folder, so a single cluster can be Origin for some data and Satellite for others. A remote write costs a round-trip to the Origin, so VAST recommends placing the Origin near write-heavy clients; write leases are also becoming portable so they can live where data is created and migrate afterward.
Caching, prefetch & intelligent streaming
Each global folder has its own cache capacity and prefetch policy, so admins decide how aggressively a Satellite warms its local flash. Data is moved only when required - the namespace is global, the bytes stay put until something reads them.
Three caching policies, one job
A job starts at the AWS Satellite. Switch the folder's policy to see when its bytes arrive - before the job, on first read, or only the names until something is touched.
On-demand (default): Satellites fetch and cache data the first time it is read, then serve it locally on flash for subsequent reads.
Data moved between sites
to AWS · illustrative
names sync, bytes on touch
First read is served from
the Origin, then cached
Illustrative: the dataset is drawn as 8 blocks and the job touches 2 of them.
Snapshots & replication
The same write-in-free-space design that powers the platform makes data protection cheap across the fabric. Snapshots take no data or metadata copy, and replication ships only the byte ranges that changed.
One dataset, protected across sites
Twelve blocks at a primary site and its replica. Write to change a byte range, then replicate, snapshot or clone. Blocks and ranges are illustrative.
Site A · primary
cyan = changed byte range
Site B · replica
dashed orange = change still on its way
Site C · no clone yet
points at the same blocks - no data duplicated
- Copied by snapshots
- none
- Shipped so far
- 5.3%
- A full re-copy ships
- 100%
Two earlier writes are already on the replica. Write again, then replicate, snapshot or clone.
Data follows the GPUs
Because one namespace spans every site, jobs can be scheduled wherever accelerators are free and stream the data they need from the Origin - rather than forcing a dataset migration first. Training runs across sites against a globally consistent view, checkpoints land without stalling GPUs, and inference deploys close to where data is generated with embeddings kept current for RAG.
In a demonstration, DataSpace connected clusters roughly 10,000 km apart (US and Japan) - with TPUs on one side and GPUs on the other - sharing one namespace. Cloud instances run natively in AWS, Azure and Google Cloud, and as of November 2025 DataSpace ships as a fully managed VAST AI OS service on Google Cloud with TPU support.
Send the job to the GPUs, not the dataset
One namespace across an on-prem Origin and four Satellites. Pick an idea, then use its control. Which sites have free GPUs, and which clouds play which role, is illustrative.
Schedule on free capacity: Send the job to whichever site has GPUs or TPUs available; the data streams to it on demand.
Pick a site with free accelerators to send the job there.
Key takeaways
In one line
A global namespace makes the same files, objects and tables addressable at local speed everywhere, with one authoritative copy instead of many.
Key points
- Each global folder has one Origin cluster holding the authoritative copy and write lease; other sites are Satellites caching under read leases.
- DataSpace is strictly, not eventually, consistent - a remote write invalidates every read lease before satellites refetch the updated bytes.
- A single cluster supports up to about 1,000,000 snapshots with no data or metadata copy; replication ships only changed byte ranges.
- In a demo, clusters about 10,000 km apart (US and Japan) shared one namespace across TPUs and GPUs.
Questions to explore
- 01How much do you spend on storage and egress replicating datasets across regions or clouds today?
- 02How often do multi-site jobs read stale data because of replication lag?
- 03Could you run training or inference wherever GPU or TPU capacity is free, instead of migrating data first?
Common questions
- Doesn't a remote write have to round-trip to the Origin, adding latency?
- Yes - a remote write costs a round-trip to the Origin, so VAST recommends placing the Origin near write-heavy clients; write leases are also becoming portable to migrate with the data.
- How is this different from simply replicating data to every site?
- Copying multiplies storage and egress cost and drifts stale the moment one site writes; DataSpace keeps one authoritative copy and streams only the bytes actually touched, or syncs metadata only.
- Where can DataSpace run?
- Cloud instances run natively in AWS, Azure and Google Cloud, and as of November 2025 DataSpace ships as a fully managed VAST AI OS service on Google Cloud.