VAST Data

VAST DataSpace

Intermediate

A global namespace that makes the same data addressable across edge, on-prem and every cloud - one source of truth, strictly consistent, with no copies to keep in sync.

The data-gravity problem

AI workloads now run wherever GPU capacity is available - on-prem, in one cloud this week and another the next, out at the edge where data is generated. But the data has gravity: it is large, it is expensive to move, and the usual answer is to copy whole datasets to every site that needs them. Those copies multiply storage and egress cost, and they drift out of date the moment one site writes.

VAST DataSpace is a - a data fabric that unifies many VAST clusters worldwide into one consistently addressable space. The same files, objects and tables are reachable from every site at local all-flash speed, and data moves only when it is actually needed.

What copying the data to every GPU site costs

Add the sites that need to run AI jobs. The usual answer is a full copy of the dataset at each one - watch the three problems grow together. Units are whole datasets, illustrative.

Sites running AI: 2
On-prem

where the data lives

sourceGPUs
Cloud A

GPUs free this week

full copyGPUs
Cloud B

GPUs free next week

Edge

where data is generated

Data gravity

Timeline at the newest site: Cloud A

staging the datasetjob

time → the GPUs idle through the orange part

Datasets are too large and costly to shuttle between sites on demand, so compute ends up waiting on staging.

Copy sprawl

Stored: 2× the dataset

Egress to fill the copies: 1× the dataset

Replicating to every location multiplies storage and egress cost - and every copy is a chance to read stale data.

Multi-site by default

2

separate views of the data

1

view the jobs need

1 copy can drift the moment one site writes.

Training and inference span on-prem, multiple clouds and the edge - they need one view of the data, not many.

One namespace, no copies

Instead of pushing a full copy of the dataset to every site, DataSpace keeps one authoritative copy in the namespace and lets each site stream only the bytes it touches - or sync metadata only, moving no data until it is read. Toggle the two models to see the difference in copies, consistency and data movement. Data that lives outside VAST today - on an old file share, in a cloud bucket or a collaboration app - gets into the namespace with SyncEngine: DataSpace connects VAST sites, SyncEngine brings outside data in.

Copy everywhere vs stream what you touch

Same four sites, same jobs. Watch what crosses the links - and the meter of data moved - under each model.

On-PremSource datasetv1full datasetEdgeSiteemptyAWSSiteemptyAzureSiteemptyGoogle CloudSiteempty

Data moved between sites

illustrative

Copy to every site0.00 datasets
DataSpace0.75 datasets

full run, for reference

01234
Copies of the data
4×
Consistency
Stale risk
Egress on sync
Per copy
1/4The dataset lives at one site; every other site is empty
Traditional: replicate the whole dataset everywhere

Each site gets a full copy of the dataset. Storage cost multiplies with every location, egress is paid on every sync, and the copies drift out of date the moment one site writes - so jobs can read stale data.

Full-copy blocksRead requestStreamed blockWriteLocal flash, filled share

Illustrative: the dataset is drawn as 8 blocks and each job touches 1-2 of them; the meter counts whole datasets moved, not measured traffic.

How global access stays consistent

VAST is blunt that “eventually consistent isn't consistent enough.” DataSpace uses an model with decentralized . Per global folder, one cluster is the Origin that holds the authoritative copy and the write lease; the others are Satellites that cache locally under read leases. Step through a remote write to see why an update in one site is visible everywhere - with no lag.

A remote write, message by message

Keys are leases: the Origin holds the write lease, each Satellite caches under a read lease. Pick which site writes, then step through what travels over the links.

Writer
On-PremOrigin · write leasev1authoritativeEdgeSatellitev1in syncAWSSatellitewriterv1in syncAzureSatellitev1in syncGoogle CloudSatellitev1in sync
1/4One namespace, one authoritative copy
One namespace, one authoritative copy

Per global folder, one cluster is the Origin - it holds the authoritative copy and the write lease. Every other site is a Satellite that caches the data locally on flash under a read lease, so the same file is addressable everywhere at local speed.

Write leaseRead leaseLease recalledWrite / ackInvalidationChanged bytes

Strictly consistent - decentralized read/write leases make an update in one site visible everywhere, with no eventual-consistency lag across clusters.

Schematic: message order is shown, not timings or distances.

Roles are per-folder, so a single cluster can be Origin for some data and Satellite for others. A remote write costs a round-trip to the Origin, so VAST recommends placing the Origin near write-heavy clients; write leases are also becoming portable so they can live where data is created and migrate afterward.

Caching, prefetch & intelligent streaming

Each global folder has its own cache capacity and prefetch policy, so admins decide how aggressively a Satellite warms its local flash. Data is moved only when required - the namespace is global, the bytes stay put until something reads them.

Three caching policies, one job

A job starts at the AWS Satellite. Switch the folder's policy to see when its bytes arrive - before the job, on first read, or only the names until something is touched.

On-demand (default): Satellites fetch and cache data the first time it is read, then serve it locally on flash for subsequent reads.

On-PremOriginall 8 blocksEdgeSatelliteidleAWSSatellitecoldAzureSatelliteidleGoogle CloudSatelliteidle

Data moved between sites

to AWS · illustrative

On-demand0.00 datasets
Full prefetch1.00 datasets
Metadata-only0.25 datasets

names sync, bytes on touch

0.25.5.751

First read is served from

the Origin, then cached

1/4On-demand: nothing moves ahead of time - the Satellite's flash is cold
Names (metadata)Read requestStreamed on readPrefetched

Illustrative: the dataset is drawn as 8 blocks and the job touches 2 of them.

Snapshots & replication

The same write-in-free-space design that powers the platform makes data protection cheap across the fabric. Snapshots take no data or metadata copy, and replication ships only the byte ranges that changed.

One dataset, protected across sites

Twelve blocks at a primary site and its replica. Write to change a byte range, then replicate, snapshot or clone. Blocks and ranges are illustrative.

Site A · primary

cyan = changed byte range

snapshotsnone yet
replica in step

Site B · replica

dashed orange = change still on its way

Site C · no clone yet

points at the same blocks - no data duplicated

Copied by snapshots
none
Shipped so far
5.3%
A full re-copy ships
100%

Two earlier writes are already on the replica. Write again, then replicate, snapshot or clone.

Data follows the GPUs

Because one namespace spans every site, jobs can be scheduled wherever accelerators are free and stream the data they need from the Origin - rather than forcing a dataset migration first. Training runs across sites against a globally consistent view, checkpoints land without stalling GPUs, and inference deploys close to where data is generated with embeddings kept current for RAG.

In a demonstration, DataSpace connected clusters roughly 10,000 km apart (US and Japan) - with TPUs on one side and GPUs on the other - sharing one namespace. Cloud instances run natively in AWS, Azure and Google Cloud, and as of November 2025 DataSpace ships as a fully managed VAST AI OS service on Google Cloud with TPU support.

Send the job to the GPUs, not the dataset

One namespace across an on-prem Origin and four Satellites. Pick an idea, then use its control. Which sites have free GPUs, and which clouds play which role, is illustrative.

Schedule on free capacity: Send the job to whichever site has GPUs or TPUs available; the data streams to it on demand.

On-PremOriginv1one copyEdgeGPUs busyqueue fullAWSGPUs freeavailableAzureGPUs busyqueue fullGoogle CloudGPUs freeavailable

Pick a site with free accelerators to send the job there.

Streamed on readMirrored writeNew version visible

Key takeaways

In one line

A global namespace makes the same files, objects and tables addressable at local speed everywhere, with one authoritative copy instead of many.

Key points

  • Each global folder has one Origin cluster holding the authoritative copy and write lease; other sites are Satellites caching under read leases.
  • DataSpace is strictly, not eventually, consistent - a remote write invalidates every read lease before satellites refetch the updated bytes.
  • A single cluster supports up to about 1,000,000 snapshots with no data or metadata copy; replication ships only changed byte ranges.
  • In a demo, clusters about 10,000 km apart (US and Japan) shared one namespace across TPUs and GPUs.

Questions to explore

  1. 01How much do you spend on storage and egress replicating datasets across regions or clouds today?
  2. 02How often do multi-site jobs read stale data because of replication lag?
  3. 03Could you run training or inference wherever GPU or TPU capacity is free, instead of migrating data first?

Common questions

Doesn't a remote write have to round-trip to the Origin, adding latency?
Yes - a remote write costs a round-trip to the Origin, so VAST recommends placing the Origin near write-heavy clients; write leases are also becoming portable to migrate with the data.
How is this different from simply replicating data to every site?
Copying multiplies storage and egress cost and drifts stale the moment one site writes; DataSpace keeps one authoritative copy and streams only the bytes actually touched, or syncs metadata only.
Where can DataSpace run?
Cloud instances run natively in AWS, Azure and Google Cloud, and as of November 2025 DataSpace ships as a fully managed VAST AI OS service on Google Cloud.