VAST Data

VAST Architecture: DASE

Intermediate

Disaggregated, Shared-Everything - the storage design that lets compute and capacity scale independently and keeps GPUs fed.

The scaling wall: shared-nothing & shared-disk

For two decades, scale-out storage came in two flavors, and both hit a wall. Shared-nothing systems give each server its own direct-attached drives, so capacity and compute are welded together and nodes must forward east-west traffic for data they do not own. Shared-disk systems let many nodes mount the same storage, but distributed locking and cache coherence become the bottleneck. Flip between them below to see the tradeoffs.

Send one read - and grow the cluster - in each design

The same client read runs through a legacy design and DASE side by side. Watch where the request has to travel, then add compute or capacity to see what each design makes you add.

1/5

Scale-out servers, each with direct-attached drives. Data is partitioned across nodes, so any node must forward requests for data it does not own.

clienteast-west mesh · 6 linksABCD

Ready. The data is striped across all 4 nodes, and every node is linked to every other - 6 east-west links.

East-west messages0Compute4Capacity4
DASE

Disaggregated, Shared-Everything: stateless compute (CNodes) reaches all data in shared NVMe enclosures directly over NVMe-oF. Any node serves any request - no node owns data, so no inter-node chatter.

clientNVMe-oF fabricCNode 1CNode 2CNode 3CNode 4DBoxDBoxDBox

Ready. Every CNode sees every NVMe SSD directly over NVMe-oF.

East-west messages0CNodes4DBoxes3
Scaling✕Compute & capacity coupled - add both together✓Compute & capacity scale independently
Network✕All-to-all east-west traffic over a full node mesh✓No east-west forwarding - remote NVMe ≈ local
Software✕Cache-coherence / data-locality logic required✓No cache coherence - state lives only in enclosures
Failure✕Lost node = lost local data = rebuild✓Stateless node holds no data - no rebuild on failure
client request / replyeast-west mesh trafficlock / coherence messagesdirect NVMe-oF readthe requested bytes

Schematic: node, drive and message counts are what this diagram draws, not measurements.

DASE: disaggregated, shared-everything

VAST introduced in 2016. It splits compute logic from persistent capacity: compute runs as stateless containers, while all data, metadata and state live in shared NVMe enclosures that every compute node reaches directly over NVMe-oF. Any node can serve any request, so there is no inter-node chatter and no cache-coherence software. What makes this practical is networking roughly 100× faster than the 2000s - fast enough that remote NVMe behaves like local storage.

DASE topology - follow one read

Clients, stateless CNodes, the NVMe-oF fabric, and the DBoxes that hold all state. Step through a read: it never crosses from one CNode to another.

1/5Play to follow one read
clientclientclientCNode 1statelessNFS · SMB · S3 · DBCNode 2statelessNFS · SMB · S3 · DBCNode 3statelessNFS · SMB · S3 · DBCNode 4statelessNFS · SMB · S3 · DBNVMe-oF fabric · Ethernet or InfiniBandDBox 1DNodeDNodeSCM · metadata + write bufferQLC flashDBox 2DNodeDNodeSCM · metadata + write bufferQLC flashDBox 3DNodeDNodeSCM · metadata + write bufferQLC flash
1 · Topology

Every CNode reaches every DBox through the NVMe-oF fabric. No CNode owns a drive, and all state lives in the enclosures.

CNode-to-CNode messages 0Data held by CNodes noneclient requestmetadata lookup (SCM)data read (QLC)

Shared-everything, not just shared disk

The “shared-everything” in DASE is the part people underestimate. Older “shared-disk” clusters share the blocks, but each node still keeps its own private cache and coordinates ownership through a distributed lock manager - so they spend their time on cache-coherence chatter, and the lock manager becomes the ceiling. VAST shares four things, not one:

Four layers - and how many of them every node shares

The same three nodes over the four layers of a storage cluster. Switch architectures to see which layers stay one shared piece and which split per node. Tap a layer for what it means.

Shared layers 4 of 4

CNode 1
CNode 2
CNode 3

Shared metadata

The V-Tree metadata lives in shared SCM and is updated with atomic operations on that persistent media - not via a chatty cache-coherence or lock-manager protocol. Every node reads and writes the one global namespace.

Sharing data and metadata is what makes every a perfectly parallel, interchangeable worker. Any CNode can serve any client request for any object at any time, because none of them depend on local state. Throughput scales almost linearly as you add nodes - there is no shared coordinator to bottleneck on. The same property makes a component failure a non-event - see Fail-in-place below.

CNodes, DNodes & the NVMe-oF fabric

CNodes and are logical constructs: both are containers, not fixed hardware roles, and on many deployments they run on the same physical boxes. CNodes (VAST Servers) are the stateless compute containers: they terminate the protocols, run the backend data services, and carry all of the system's compute. DNodes are the lightweight containers that own the NVMe devices, the SSDs and QLC flash inside the enclosures (). State lives only on those devices. The fabric between them is NVMe-over-Fabrics running on Ethernet or InfiniBand.

Anatomy of a DASE cluster

Tap any part of the cluster, or any line below, to see what it does. CNodes and DNodes are containers, not fixed hardware roles.

clientclientclientCNode 1statelessNFS · SMB · S3 · DBCNode 2statelessNFS · SMB · S3 · DBCNode 3statelessNFS · SMB · S3 · DBCNode 4statelessNFS · SMB · S3 · DBNVMe-oF fabric · Ethernet or InfiniBandDBox 1DNodeDNodeSCM · metadata + write bufferQLC flashDBox 2DNodeDNodeSCM · metadata + write bufferQLC flashDBox 3DNodeDNodeSCM · metadata + write bufferQLC flash

CNodes (VAST Servers)

NVMe-oF fabric

DNodes (inside DBoxes)

Stateless compute = elastic, no-rebuild scaling

Statelessness is the property that pays for itself. Because CNodes hold no data, the cluster scales linearly, upgrades fully online, and grows or shrinks on demand. Failure handling is under Fail-in-place below.

What statelessness buys - three scenarios on one cluster

Stateless CNodes on top, enclosures holding all state below. Pick a scenario and use its control. Node counts and units are illustrative.

CNodes · stateless compute

CNode 1serving
CNode 2serving
CNode 3serving
NVMe-oF - every CNode sees every NVMe SSD

Enclosures · all data and state

DBox 1DNodesDNode 1ADNode 1B
DBox 2DNodesDNode 2ADNode 2B
Linear scaling

Add CNodes for throughput, add enclosures for capacity - independently. The two axes never have to grow together.

Throughput
3 units
Capacity
2 units
Must grow together
no

Throughput - grows with CNodes

Capacity - grows with enclosures

The write path: SCM → reduce → QLC

Every write lands in Storage Class Memory (SCM) first and is acked from there. Only later, off the critical path, is the data reduced and flushed to dense QLC flash in very large stripes. That deferral is what lets VAST absorb writes on low-endurance QLC without write amplification or garbage collection. Step through the three stages below.

1Client
2CNode DRAM
3SCM SSDs
4QLC flash

↑ client acked from SCM

Stage 1

DRAM buffer → parity → SCM, ack client

A CNode buffers the incoming write in DRAM. On a full stripe - or after a ~300µs timeout, zero-padded - it computes parity and writes data + parity across the SCM SSDs, then acks the client. All writes land in low-latency SCM first and are acknowledged from there.

client sees the ack here

Buffering in SCM, deferring data reduction, and flushing to QLC in very large stripes yields a ~10-year QLC lifespan and HDD-like cost on all-flash.

Locally-decodable erasure coding

Classic Reed-Solomon forces a tradeoff: wide stripes are capacity-efficient but expensive to rebuild, because reconstructing one lost strip means reading the whole stripe. VAST's computes parity strips from different data-strip sets spanning more than one stripe, so with P parity strips you rebuild a lost strip by reading only ~1/P of the surviving data strips plus the P parity strips - about 41 of 149 for 146+4. Compare schemes below.

146D + 4P

filled = data strips read on rebuild (~37 of 146, plus all 4 parity) · outlined = parity strips

Capacity overhead = P / (D+P)2.7%

lower is better

Data strips read on rebuild ≈ 1 / P25.3% (~37/146 + 4 parity)

lower is better

VAST locally-decodable wide stripe: parity strips span multiple stripes, so a lost strip is rebuilt by reading only ~1/P of the surviving data strips plus the 4 parity strips.

Wide stripes normally force a rebuild to read the whole stripe. Locally-decodable coding breaks that Reed-Solomon tradeoff: you get wide-stripe efficiency and cheap rebuilds at once.

A single-strip rebuild in 146+4 reads ~37 of 146 data strips plus the 4 parity strips - about a quarter of what a Reed-Solomon stripe of the same width reads - at 2.7% overhead. VAST's per-stripe Markov model puts MTTDL at 42M+ years.

Fail-in-place: failures become a non-event

A failed CNode held no data and no exclusive locks, and any CNode reaches any SSD, so the survivors keep serving the same requests with no rebuild, no ownership failover, and no job restart. Each DBox is fronted by two DNodes, so a DNode failure is a non-event too: its partner keeps the drives reachable - which is also what lets the rolling upgrade above stay online. Fail a CNode or a DNode below.

Shared-everything: any node, any request, any time

Six requests flow from clients through the CNodes, across the shared fabric, and through any DNode into the DBox holding their data. Every CNode reaches every DNode, and several CNodes can use the same DNode at once. Tap a CNode to fail it and watch only its requests move to the survivors; tap a DNode to fail it and watch its partner in the same DBox pick up its requests - the job never stops.

shared metadatashared datashared disk (NVMe-oF)R1R2R3R4R5R6CNode 12 reqCNode 21 reqCNode 32 reqCNode 41 reqDBox 12 pathsDNode 1ADNode 1BDBox 22 pathsDNode 2ADNode 2B
All 4 CNodes serving in parallel; every CNode reaches every DNode. Because every node sees the same shared data and metadata, any CNode can answer any request for any byte, through either DNode of the DBox that holds it - there is no “owner” to route to.
request pathrerouted pathfailed - tap again to restorerequests servedfabric lanes: what every CNode shares, directly over NVMe-oF

Illustrative model of VAST's stateless-CNode failover and dual-DNode path resiliency. Request placement and DBox layout are simplified for clarity.

Drives fail in place, because rebuilds are cheap. A lost strip is reconstructed by reading only ~1/P of the data strips plus the parity, and the 146+4 scheme tolerates several concurrent losses with an MTTDL measured in tens of millions of years, so VAST does not treat a dead SSD as an emergency. Failed media is simply left in place: the system routes around it, keeps serving at full speed, and you swap hardware in planned batches during a maintenance window, or run an enclosure to end of life without ever replacing individual drives. There is no urgent single-drive rebuild race and no degraded-performance window to sweat. Fail drives below and watch the system stay healthy.

Stripes across drives - and the many-to-many rebuild

Each drive holds strips from many different stripes, plus a slice of spare space. Fail a drive and follow the CNodes as they rebuild its strips into spare slices across the enclosure - while the dead drive stays put.

1/41. Healthy
CNodes
System status: HEALTHY
8 CNodes - stateless, no dataC1C2C3C4C5C6C7C8
Rebuild work done (relative - no time scale)0%
strips lost 0stripes readable allstrips rebuilt 0/4

All drives nominal - data is protected

Every stripe is spread across many drives, one strip per drive, so any drive can be lost without losing data. Tap a drive to choose which one fails, or step forward.

strip of some stripethe 4 stripes that lost a stripfailed drive spare sliceread: surviving strip → CNodewrite: CNode → spare slice

Absorbed in software

Wide erasure coding tolerates drive loss with no downtime and no data loss. A failure is a non-event for availability.

Many-to-many CNode rebuild

The CNodes do the rebuild, reconstructing lost data into reserved free space across all surviving drives - not a single hot-spare. Shared-everything means every CNode can contribute.

Replacement is deferred

The dead drive stays in the chassis. No emergency truck roll: swap failed media in planned batches, or run to end of life.

Illustrative model of fail-in-place behavior: 8-strip stripes, each rebuild reading ~1/P of a stripe (P = 2 here); real VAST stripes are far wider. A 146+4 wide stripe has an MTTDL of more than 42M years.

Data reduction: deduplication, compression & similarity

Three techniques stack here, and they are not interchangeable. Deduplication removes blocks that are byte-for-byte identical to one already stored - powerful when whole copies repeat, but coarse: a single changed byte makes two blocks different, so near-duplicates slip straight through. Compression shrinks the bytes within a block using an encoder; it only ever sees one block at a time, a purely local view, so it cannot touch the redundancy that lives between similar blocks.

VAST's differentiator is the third technique: global similarity reduction. A similarity hash returns the same value for blocks that are merely similar, not identical. The first block of a hash becomes a compressed reference holding the full dictionary; every later similar block is stored as a small delta against it (ZSTD, without the dictionary overhead), capturing the mid-size repetitions that dedup and compression alone leave on the table. It runs on fine, dynamic chunks (on the order of 32 KB) rather than one fixed block, giving the hash far more chances to match. Crucially the hash table is global across the whole cluster and lives in SCM - so the all-flash DASE design lets the similarity search span the entire namespace as one reduction realm, a vastly larger scope than legacy systems whose dedup tables must fit in controller DRAM. A bigger search space finds more redundancy, which is exactly why “all data reduction is not created equal”. Toggle the techniques to see what each catches.

VAST applies all three

Every write is globally compressed and deduplicated; similarity reduction is an optional layer on top. Toggle the techniques to compare what each one catches - the Similarity reduction view is VAST's full combined result.

Deduplication: Deduplication hashes every block and collapses any block whose hash COLLIDES with one already stored - so byte-for-byte identical copies become a single pointer. It excels on data with many exact copies, but it is coarse: a single changed byte yields a different hash, so near-duplicates slip through and are kept in full.

Storage footprint

1.9:149% smallerin illustrative units - not a measured ratio

What was written6.0 units
Stored on VAST3.09 units
Deduplication relies on an exact hash collision - one changed byte and the hashes differ. A similarity hash instead returns the same value for blocks that are merely similar. The first block of a similarity group becomes the reference; later similar blocks are stored as ZSTD deltas. Both hash tables are global across the cluster and live in SCM, so reduction works against the entire namespace.

Why VAST's scope is global, not local

Legacy systems confine dedup to a small local domain because the lookup tables have to fit in controller DRAM. VAST keeps all reduction metadata in Storage Class Memory across the enclosures, so any compute node can compare a new chunk against the entire cluster as one reduction realm. That far larger similarity scope, on fine ~32 KB chunks, is what finds the near-duplicates legacy systems never see - the core thesis of “all data reduction is not created equal”. Metadata points at bytes, not blocks, so a chunk that reduces to 29,319 bytes occupies 29,319 bytes with no packing waste.

Conceptual illustration; unit sizes are illustrative. Reduction ratios vary by workload, from ~2:1 on CGI/animation and ~4:1 on Splunk up to ~22:1 on backup data.

The Element Store & the V-Tree

The maps physical flash into user-facing “Elements” - files, S3 objects, tables and block volumes - in one namespace. The is the metadata structure behind it: a wide, shallow B-tree variant whose nodes live on SCM and point to bytes, not blocks. Each level fans out roughly 512×, so a handful of hops can address on the order of 100 trillion elements, and because pointers reference exact byte ranges, copy-on-write snapshots and clones cost almost nothing. Take a snapshot below to see the new root reuse existing subtrees.

7 hops to any of ~100 trillion elements

Pick a protocol to start a lookup - every protocol resolves to the same Elements, so the path through the V-Tree is the same. Then take a snapshot and modify a file to see copy-on-write share everything except one path.

Protocol & data access layer - tap to look up

1/9Lookup via NFS - press play
V-Tree metadata · SCMlookup via NFSL1rootrootL2… ×512level 2L3… ×512level 3L4… ×512level 4L5… ×512level 5L6… ×512level 6L7… ×512leavesData · QLC flash
Hops 0/7Fan-out ~512× per levelAddressable ~100T elements
lookup path (same for every protocol)snapshot root's shared pointerscopy-on-write path… ×512 = the rest of each level, not drawntinted band = V-Tree metadata in SCM

Wide and shallow

The V-Tree is 7 levels deep and each level fans out ~512x, so any of ~100 trillion elements is reached in 7 or fewer hops - versus the tens or hundreds of redirections a deep B-tree needs. The whole structure lives in SCM as one global, shared-everything namespace, so no cache-coherence chatter between nodes is required.

Copy-on-write snapshots

Take a snapshot to add a second root that reuses every level below it - pointer-based, so no bytes are duplicated.

Simplified from VAST's published Element Store / V-Tree design (7 levels, ~512x fan-out per level, ~100 trillion elements); only a few nodes per level are drawn. Every protocol resolves to the same Elements with equal performance and equal access.

One namespace, every protocol

Because every protocol resolves to the same Element, NFS, SMB, S3 and GPUDirect Storage clients all reach the same bytes simultaneously, with no copies and no gateway translation layer. That matters for AI because a pipeline rarely speaks one protocol: data is ingested over S3, curated over file, streamed to GPUs over GPUDirect, and served back to inference and RAG, all against a single copy. Step through the stages to see which protocol each uses and why a unified namespace removes an entire class of copy-and-sync work.

One copy of the data, four doors into it

Each pipeline stage speaks its own protocol, but every protocol resolves to the same Elements. Step through the stages (or tap one), then switch to separate stores to see the copies that come back.

1/4
One dataset · one namespacefiles and objects are the same Elementstouched via✓ S3NFS / SMBGPUDirectS3 / NFS1 IngestS32 Curate / preprocessNFS / SMB3 TrainGPUDirect (NFS / S3)4 Serve / RAGS3 / NFS
Copies of the dataset 1Copy / sync jobs noneactive stage's protocol path
1. IngestS3

Apps, sensors, and data pipelines land raw data over S3, the lingua franca of cloud-native ingest. It arrives directly in the namespace, with no landing-zone bucket to later copy out of.

Without multi-protocol access, each stage needs its own silo and you copy or sync data between an object store, a file store, and a training cache, paying for the capacity several times and fighting drift. VAST serves NFS, SMB, S3, and GPUDirect Storage against one set of bytes, so the whole loop, ingest to inference, runs on a single copy.

Why this matters for AI / GPU workloads

AI infrastructure has an asymmetric appetite: training (dataset reads and bursty checkpoint writes) and inference want enormous, steady throughput to keep GPUs fed, while datasets grow into the petabytes on a different curve. DASE matches that shape.

Three properties, three pictures

How DASE lines up with what AI pipelines need. Try each one; grid sizes are illustrative.

Asymmetric scaling

throughput →

capacity →

Reachable mixes of capacity and throughput: 25 of 25

Scale GPU-facing throughput by adding CNodes without buying more petabytes - and grow capacity without buying more compute. That is exactly the asymmetry AI pipelines need.

Cheap snapshots & clones

live root snapshot root data

Data copied by the snapshot: none

Pointer-based, copy-on-write snapshots cost almost nothing, so reproducible training datasets and experiment branches are practically free.

Near-zero capacity tax at PB scale

One 146+4 stripe, strip by strip

146 data strips4 parity

Parity overhead 2.7%

Wide locally-decodable erasure coding keeps overhead near 2-3% even on huge drives, with rebuilds cheap enough to run on multi-PB clusters.

Key takeaways

In one line

DASE splits stateless compute from shared NVMe storage, so capacity and throughput scale independently and a failed compute node needs no rebuild.

Key points

  • CNodes are stateless compute containers holding no data, so a failed CNode needs no rebuild - requests just reroute to survivors.
  • Locally decodable erasure coding (146+4) runs at ~2.7% overhead, and a rebuild reads only about a quarter of the stripe.
  • Global similarity reduction finds near-duplicate, not just identical, blocks across the whole cluster using fine ~32 KB chunks and a hash table in SCM.
  • Writes land first in Storage-Class Memory and are acked from there; reduction and erasure-coded flushing to QLC happen later, off the critical path.

Questions to explore

  1. 01How often do drive or node failures cause a degraded-performance rebuild window in your environment today?
  2. 02What capacity overhead does your current erasure coding or RAID scheme carry compared to roughly 2-3%?
  3. 03Do your compute and capacity needs grow at different rates, and can today's storage scale them independently?

Common questions

Doesn't sharing everything across nodes create the same coordination bottleneck as shared-disk systems?
No - metadata lives in SCM and is updated atomically rather than via locks, and data has no owner or shards, so nodes never need cache-coherence chatter to serve a request.
How does fail-in-place avoid data loss if drives sit dead for a while?
The 146+4 scheme tolerates multiple concurrent losses, with an MTTDL of 42M+ years, so a dead drive isn't urgent to replace.
How does low-cost QLC flash last in this design?
A ~10-year QLC lifespan comes from buffering writes in SCM, deferring data reduction, and flushing to QLC in very large stripes, which limits write amplification.