VAST Data
VAST SyncEngine
FundamentalFind the data scattered across file shares, object stores and SaaS apps, then move what matters onto the platform - verified, resumable, and kept in sync until cutover.
The last-mile data problem
finds data across your file shares, object stores and SaaS apps, and moves the data you choose onto the VAST platform - keeping it in sync until you cut over.
It has been available to all VAST customers since August 2025, at no additional cost, and ships on its own release cadence (version 3.4 as of September 2026).
Most of the data an AI project could use already exists. It sits in older file servers, object buckets and collaboration apps, invisible to AI pipelines until someone brings it to where those pipelines run. That last step hurts in four ways:
- You can't use what you can't see. Nobody can say how many PDFs there are, where, how old, or who owns them - so projects start with weeks of discovery.
- Moving it is a project, not a command. Hand-written scripts and separate tools for listing, copying and checking; a job that dies halfway starts over; nobody is sure the copy is complete.
- File count, not just size, sets the clock. Hundreds of millions of small files take far longer to move than bandwidth math suggests.
- Copies go stale. A one-time copy is out of date the next morning, and a RAG assistant answers from last quarter's documents.
See first, then move
SyncEngine pairs two jobs. The first is seeing: a metadata indexing job walks a POSIX file system and writes one row per file - path, size, owner, timestamps - into a table. That table is a of the file system's metadata: no file contents move, and you query it with SQL engines or the vastdb SDK to size the job and decide what is worth bringing. The second is moving: scale-out workers copy the chosen buckets, folders, spaces and paths onto VAST, retry what fails, and keep the copy current.
Think of it as moving house: survey the rooms first, then hire the movers - and keep forwarding the mail until moving day. Step through one illustrative estate below.
From scattered data to AI-ready, one step at a time
Follow one illustrative estate - a file share, two object buckets and three SaaS sources - through SyncEngine. Tap a step, or press play.
Index rows
0M
file share only
Planned move
-
not chosen yet
Bytes moved
0 B
nothing copied yet
Failed files
0
A metadata indexing job walks the file share and writes one row per file into a VAST DataBase table - path, size, owner, dates. No file contents are copied, so bytes moved stays at zero.
VAST DataBase table - one row per file (illustrative rows)
| path | size | mtime | ext |
|---|---|---|---|
| /proj/atlas/specs/design-v4.pdf | 8.1 MB | 2025-11-03 | |
| /proj/atlas/scratch/run_0192.tmp | 2.4 GB | 2023-02-17 | tmp |
| /legal/contracts/msa-2026.docx | 310 KB | 2026-06-21 | docx |
| /it/images/base-os-2021.iso | 5.7 GB | 2021-08-09 | iso |
| /support/kb/export-0907.md | 46 KB | 2026-09-07 | md |
SELECT extension, count(*), sum(size) FROM index_run_1 GROUP BY extension;
Query the index with SQL engines or the vastdb SDK to size the move, spot cold data and decide what to leave behind. Each indexing run creates one table, for one source path.
Sizes, counts and file names are illustrative. Metadata indexing covers POSIX file systems; the other sources are copied by their own migration jobs.
What it connects to
Each migration reads one source - a bucket, a space, a folder or a path - and writes to one destination. Object stores and collaboration apps land in VAST S3; NFS exports land in VAST NFS. For collaboration apps, SyncEngine also converts content into formats AI pipelines read easily and turns sharing settings into S3 ACLs.
| Source | Lands in | Access | What carries over |
|---|---|---|---|
| S3-compatible bucketsAWS S3, MinIO, IBM COS and others | VAST S3 | S3 API with access keys | Object tags copied by default; S3 ACLs and versions are not carried over. AWS S3 can also sync from bucket events via SQS. |
| Google Cloud Storageone bucket per migration | VAST S3 | Google Cloud service account | Same S3 destination options as the other object sources. |
| Confluence Cloudone space per migration | VAST S3 | Atlassian API token | Pages, blog posts, comments and attachments; converted to Markdown; permissions become S3 ACLs. |
| Google Driveone folder per migration | VAST S3 | Google Cloud service account | Converted to Markdown by default (PDF and Markdown files kept as they are); permissions become S3 ACLs. |
| OneDrivewhole tenant or selected users | VAST S3 | Microsoft Entra ID app (Microsoft Graph) | Original file format plus a JSON metadata object; permissions become S3 ACLs. |
| NFS export (NFSv3)native NFSv3 scan | VAST NFS | rsync copy | POSIX ACLs kept; NFSv4.1 ACLs are not. |
| POSIX pathmounted on every worker | POSIX path | rsync copy | File system to file system, for example onto a mounted VAST view. |
| POSIX path (index only)one path per indexing job | VAST DataBase table | VAST DataBase connection | Metadata only - path, size, owner, timestamps and more - with no file contents copied. |
- S3-compatible buckets→ VAST S3
AWS S3, MinIO, IBM COS and others · S3 API with access keys
Object tags copied by default; S3 ACLs and versions are not carried over. AWS S3 can also sync from bucket events via SQS.
- Google Cloud Storage→ VAST S3
one bucket per migration · Google Cloud service account
Same S3 destination options as the other object sources.
- Confluence Cloud→ VAST S3
one space per migration · Atlassian API token
Pages, blog posts, comments and attachments; converted to Markdown; permissions become S3 ACLs.
- Google Drive→ VAST S3
one folder per migration · Google Cloud service account
Converted to Markdown by default (PDF and Markdown files kept as they are); permissions become S3 ACLs.
- OneDrive→ VAST S3
whole tenant or selected users · Microsoft Entra ID app (Microsoft Graph)
Original file format plus a JSON metadata object; permissions become S3 ACLs.
- NFS export (NFSv3)→ VAST NFS
native NFSv3 scan · rsync copy
POSIX ACLs kept; NFSv4.1 ACLs are not.
- POSIX path→ POSIX path
mounted on every worker · rsync copy
File system to file system, for example onto a mounted VAST view.
- POSIX path (index only)→ VAST DataBase table
one path per indexing job · VAST DataBase connection
Metadata only - path, size, owner, timestamps and more - with no file contents copied.
Documented source and destination paths in SyncEngine 3.4 (September 2026). Object and SaaS sources land only in VAST S3 buckets.
Inside a migration: control plane and workers
SyncEngine runs on hosts you provide, next to the platform rather than inside it. A control plane handles scheduling, coordination and monitoring; the workers, spread across several hosts, do the copying. Each worker runs many parallel transfer threads (by default twice its CPU count), and a worker label ties a worker to specific migrations. The software is included with VAST; the hosts are yours.
One host
Control plane
- API and web UI
- Internal database (PostgreSQL) for jobs and state
- Job queue (Redis) that hands work to workers
- Monitoring with Prometheus, Grafana and Loki
Two or more hosts
Workers
Worker 1
threads
Worker 2
threads
Worker n
threads
Each worker copies from the source to VAST with many parallel threads. Add hosts to add copy streams.
Runs as containers on Docker or Podman, or from a Helm chart on Kubernetes. The minimum layout is three hosts: one control plane and two workers.
Jobs are built for long runs: they can be paused and resumed, failed files can be replayed without re-copying the rest, rate limits protect busy sources and SaaS APIs, and progress, logs and metrics flow to Grafana dashboards.
Why file count matters as much as bytes
Every file costs a fixed amount of work before and after its bytes move - listing, opening, creating, checking, closing. With large files that overhead disappears into the transfer; with millions of small files it dominates, and the job runs far below what the network could carry. Parallel threads hide the per-file wait, which is why the tuning guidance depends on file size: 80 to 128 threads per worker for many small files, 16 to 32 for large files. Past a point, more threads stop helping because the source, the network or the destination is the limit.
Same bytes, different file counts
Keep the total fixed and change how it is split into files. Every file pays a fixed overhead - list, open, create, checksum, close - on top of moving its bytes. Illustrative values: 100 Gb/s network, 0.5 GB/s per copy process.
Total data
Per-file overhead (illustrative)
Throughput
13 GB/s
100% of the link
Copy time
2.2 h
network-bound
Bytes / bandwidth
2.2 h
what size alone predicts
Copy streams
64
2 workers x 32
Docs guidance: 16-32 for large files, 80-128 for many small files
Workers
Throughput vs processes per worker
File count sets the clock as much as bytes. With 100 MB files the copy streams keep the link full, so more processes add nothing - the network sets the time.
Planning a move: full copy, resync, cutover
A migration is not one copy. The first pass moves everything selected, which can take days while users keep working on the old system. Each after that carries only what changed since the last pass, so the passes get short. The is then a short pause: stop writes on the old system, run one final incremental pass, and point users and applications at VAST. SyncEngine shortens the cutover window; it does not remove it.
A skip-existing option avoids re-copying data that is already at the destination, and on S3 resyncs, objects with the same Last-Modified time on both sides are skipped. Values in the planner are illustrative.
Migration and cutover planner
A full copy, then scheduled resync passes that carry only changes, then a short freeze for the final delta. Every rate here is an illustrative assumption you can change to match your own environment.
Source type
NFSv3 export or POSIX path, copied with rsync on each worker. Picking a type loads its illustrative defaults below.
Source request limit (req/s)
The docs quote 1000-3000/s for AWS S3 and 5-10/s for Google Drive and Confluence
Network, source to VAST
Docs guidance per worker: 16-32 for large files, 80-128 for many small files
Resync every (scheduled)
Cutover window target
Bytes over the link set the pace - more workers or processes won't shorten the copy.
Full copy
22 h
1 PB in 5 million files
Bytes / bandwidth alone
22 h
the naive estimate
Resync passes
1
each scans every file: 26 s
Final cutover
56 s
fits the 15 min target
Timeline in hours - 24 h end to end
After 1 resync pass, the changes left to copy during the freeze fit the 15 min window.
Event-driven sync: only AWS S3 sources can sync from bucket events (via SQS). Every connector supports the scheduled resync modeled here (default every 24 h).
Illustrative model: copy time is the largest of bytes over the network, the source's limits (one request per file), and file count x (overhead + transfer) spread over all copy streams. Each resync pass scans every file (0.5 ms per file for this source type) and copies what changed since the previous pass began. One process streams 0.5 GB/s. None of these are measured SyncEngine figures.
Staying current after the move
Every connector can run in two modes. Manual runs a migration on demand. Automatic starts the first sync right away, then resyncs on a schedule set in hours - every 24 hours by default. For AWS S3 sources there is a third option: event-driven sync that reads the bucket's change events from an Amazon SQS queue, so new objects follow soon after they are written rather than at the next scheduled pass; any event that is missed is caught by the next full scan.
Deletions are handled deliberately. By default, a file deleted at the source stays on VAST. Mirroring deletions is an opt-in setting that makes the source the single source of truth - and files it removes at the destination do not come back if the setting is turned off later. For a assistant, the resync cadence sets how stale its sources can get before the pipeline even sees a change.
Staying current: how fresh is the copy on VAST?
The source keeps changing after the first copy. Pick a source and a run mode, then play a week of edits: each one waits - hatched - until a sync brings it over, and the search index can only catch up after that.
Source
Lands in a VAST S3 bucket.
Run mode
Worst-case staleness
22 h
An edit just after a sync waits up to the full 24 h.
Average staleness
15.3 h
How long this week's edits waited to reach VAST.
Changes still pending
0
of 0 edits so far - invisible to search until synced.
After each sync
- Change at the source
- SyncEngine copies it to VAST
- A DataEngine trigger fires
- InsightEngine re-embeds it
- Searchable
SyncEngine doesn't embed anything itself - it lands the change, and the trigger starts the embedding. So staleness adds up: the sync lag set by the run mode, plus the processing time after it.
Takeaway: a one-time copy starts going stale with the first edit. A scheduled resync caps staleness at the interval - the first run starts right away and later ones follow every N hours. For AWS S3, event-driven sync keeps the copy close to current.
Illustrative: edit times, sync-run lengths, the event lag and the embedding delay are schematic; real lag depends on the source, the queue and the amount of changed data.
How SyncEngine fits the AI OS
SyncEngine is the on-ramp. Its job ends when data lands in a VAST S3 bucket or NFS view; what happens next is the rest of the platform. A DataEngine trigger on the bucket starts a pipeline, chunks and embeds the documents, and the vectors and metadata are stored in VAST DataBase for retrieval by RAG applications and agents. SyncEngine itself does not chunk or embed. The vendor-neutral version of this flow is on the data pipeline page, and the InsightEngine page covers the embedding side.
DataSpace means you never copy data between VAST sites; SyncEngine makes the one copy that brings outside data in. They solve different problems: SyncEngine gets data onto VAST, and DataSpace makes it available at every site once it is there.
Among other ways to move data, SyncEngine is the tool VAST includes for bringing data onto its own platform. Vendor-neutral migration tools focus on analyzing and moving data between any platforms, and cloud transfer services do the same job for their own clouds.
Key takeaways
In one line
SyncEngine brings data from file shares, object stores and SaaS apps onto VAST - see it first, move what you choose, and keep it in sync until cutover.
Key points
- It pairs a metadata index - POSIX file-system metadata written to a VAST DataBase table you query with SQL - with scale-out workers that copy the chosen data onto VAST S3 or NFS.
- Documented sources are S3-compatible buckets, Google Cloud Storage, Confluence Cloud, Google Drive, OneDrive, NFSv3 and POSIX; collaboration content can be converted to Markdown, with permissions carried over as S3 ACLs.
- File count matters as much as bytes: per-file overhead dominates small-file moves, so worker thread counts are tuned to the file-size mix.
- After the first full copy, scheduled resyncs (or event-driven sync for AWS S3) carry only changes, so cutover is a short final pass; DataEngine and InsightEngine then embed the landed data.
Questions to explore
- 01How many separate file, object and SaaS systems hold data your AI projects could use, and does anyone have one view of them?
- 02What does your largest migration look like in file count, not just terabytes?
- 03How stale can source documents get before your RAG answers are wrong?
Common questions
- A share holds hundreds of millions of small files. Why can it take much longer to migrate than an archive of large videos of the same total size, and what helps?
- Every file carries fixed overhead - listing, opening, creating, checking - so small files make the job overhead-bound rather than bandwidth-bound; more parallel threads per worker, and more workers, hide that overhead until the source, network or destination becomes the limit.
- DataSpace promises no copies to keep in sync. Why does VAST also include a sync engine?
- DataSpace removes copies between VAST sites; SyncEngine brings data in from systems outside VAST - it makes the one copy onto the platform and keeps it current until the old system is retired.