Applications

Data Pipeline for AI

Fundamental

Models get the headlines, but a model is only as good as the data fed into it - and only as fast as the storage feeding the GPUs. This page starts with the plumbing in plain English, then goes deep on how AI data pipelines, lakes, and lakehouses actually work, and why high-performance storage is the hidden lever on cost and speed.

Start with the plumbing

A is plumbing for data: pipes that move it from where it's created (apps, sensors, the web) to where it's used (dashboards, models), cleaning and reshaping it along the way. Where the water is stored matters too - and there are three common reservoirs.

Three reservoirs - what actually sits inside each

A warehouse keeps modeled tables, a lake keeps raw files of every type, and a lakehouse keeps the same raw files with a table-format layer on top. Drawings are schematic.

Data warehouse

A tidy library

modeled tables only

Structured, schema-on-write tables. Data is cleaned and modeled before it lands.

Strong: Fast, reliable SQL & BI dashboards.

Watch out: Rigid; struggles with raw text, images, and exploratory AI work.

Data lake

A giant storage unit

raw files, any type

Raw files of any type (text, images, video, logs) dumped cheaply into object storage.

Strong: Holds everything; cheap; perfect for ML training corpora.

Watch out: No guarantees - easily becomes a “data swamp” with no order.

Lakehouse

A renovated warehouse on the lake

table format over raw files

A lake plus an open table format (Iceberg/Delta/Hudi) that adds database guarantees.

Strong: Warehouse reliability over lake-scale raw data - one system for BI and AI.

Watch out: Younger ecosystem; needs governance discipline to stay clean.

warehouse: structured, schema-on-writelakehouse: guarantees on raw fileslake: raw, any type

ETL vs ELT - the same steps, a different order

Switch modes and watch Transform and Load trade places.

EExtract
TTransform
LLoad

The warehouse holds

polished resultonly

ETL - Extract, Transform, Load

Clean and reshape data before loading it. Classic warehouse approach - you only store the polished result.

Great when the schema is known and storage is precious.

A normal pipeline vs. an AI pipeline

A traditional analytics pipeline ends at a clean table a human reads. An AI pipeline keeps going: the data has to become something a model can learn from or retrieve. That adds a whole new set of stages.

Same start, different finish

Both pipelines ingest and clean. The classic one stops at a table a human reads; the AI one keeps going and forks.

Shared start

Ingest

pull from apps & databases

Clean & join

fix types, merge sources

then the paths split ↓

Classic analytics pipeline

Model & aggregate

build SQL tables

Dashboard

a human reads a chart

end: a human reads it

AI-centric pipeline (the extra stages)

Curation

select & weight sources by quality, not just volume

Deduplication

remove near-duplicate docs that waste compute & memorize

Tokenization

convert text to integer tokens, pack & shard for the GPUs

training corpussee the refinery ↓

Embedding generation

encode chunks into vectors for retrieval (RAG)

Vector indexing

build ANN indexes so similar vectors are findable fast

retrieval index

Feature store

compute & serve the same features offline and online

online + offline features

Labeling / RLHF

human & AI feedback to align and rank model outputs

alignment data

Same raw data, but it forks into a training corpus, a retrieval index, and a - each with its own quality bar.

Both forks start with the same unglamorous step: finding the data and bringing it in from wherever it lives today - old file shares, object buckets, collaboration apps - then keeping that copy in sync as the source changes. On VAST, that on-ramp is SyncEngine.

The data refinery: training data is manufactured, not found

Common Crawl is 9.5+ PB of raw web, ~250 TB and 2B+ pages in a single monthly snapshot - and over 80% of GPT-3's training tokens came from it. But you can't train on raw web; most of it is junk, boilerplate, and duplicates. Hugging Face's FineWeb distilled ~15 trillion high-quality tokens from 96+ snapshots - and beat bigger datasets by filtering and deduplicating harder, not crawling more. Watch the volume shrink at each stage.

The data refinery

Raw web in, training-ready tokens out. The stream's width is the data volume; each filter peels most of it away. Turn a filter off and see what gets through instead.

Filter stages
Raw crawl250 TBLang & quality250 TBDedup45 TBCurate16 TBTokenize8.7 TB−205 TBother langs, junk, boilerplate−29 TBduplicates−7.1 TBlow quality8.7 TB→ 15.00 T tokenstraining-ready
kept data (width = TB)discardedjunk that got through (filter off)

Survives the pipeline

3.46%

of the raw 250 TB snapshot

Training-ready output

15.00 T tokens

all filters on → FineWeb-scale ~15 T quality tokens

Turn filters off and the surviving volume balloons - but those extra “tokens” are duplicates and junk that hurt the model. FineWeb beat bigger datasets by filtering harder, not crawling more. Illustrative model; survival fractions approximate the real Common Crawl → FineWeb funnel.

Feeding the GPUs: compute is only as fast as the data

A training run is a race to keep thousands of GPUs busy. If storage can't stream training shards fast enough, GPUs sit idle and (MFU) - the fraction of theoretical compute you actually use - collapses. Idle GPUs at $3/hr each are pure burned cash. This is why high-performance storage is an AI-infrastructure decision, not a commodity one.

Keeping the GPUs fed

A GPU at $3/hr earns nothing while it waits on data. If storage can't supply bytes as fast as the fleet consumes them, utilization (MFU) drops and idle GPU-time burns cash. Pick a storage tier and watch the fleet.

Storage tier

fast but small, manual staging, no sharing

Provisioned bandwidthtier ceiling 120 GB/s60 GB/s
GPUs in the runeach wants ~4 GB/s64
Storage is the bottleneck: only 60 GB/s feeding 256 GB/s of demand → 23% MFU. The other GPUs idle.
60 GB/s delivereddemand 256 GB/s (dashed)Local NVMe≤120 GB/s64 GPUs23% fed

Pipe cross-section is proportional to bandwidth (thickness ∝ √GB/s). The colored bar at the outlet is the tier's ceiling.

GPU fleet (15/64 fed)

23% MFU

Burn rate while starved

$147/hr

49 idle GPUs at $3/hr each

Wasted since you arrived

$0.00

ticking in real time

Switch to the parallel all-flash filesystem and the ceiling jumps to ~1,000 GB/s - the fleet stays green even at full scale. Compute is only as fast as the data feeding it: high-performance storage is what turns expensive GPUs into useful FLOPs (the VAST value prop). Illustrative model; per-GPU demand and tier ceilings are approximate.

The other heavy I/O: checkpoints

Training also writes checkpoints - full snapshots so a crashed run can resume instead of starting over. With Adam mixed precision a checkpoint is roughly 16 bytes/param, so a 70B model is ~1 TB+. Frontier runs checkpoint frequently for fault tolerance, and the whole fleet stalls while writing - so fast write and fast read (for recovery) directly protect expensive GPU-hours.

Synchronous vs asynchronous checkpoints - same clock

Two checkpoints in the same window. Watch what the GPUs do while each one is written. Durations are illustrative units, not measurements.

1/25t = 0 of 24 (illustrative units)

At time 0: synchronous has stalled the GPUs for 0 units; asynchronous has paused them for 0.

Synchronous checkpoint

GPUs
training job
Storage
checkpoint writes
GPU time spent training0.0 / 0 · 0.0 idle

The classic approach: training pauses, copies the full model + optimizer state out of GPU memory, and waits for it to land on storage before resuming. Every GPU in the job sits idle for the whole write. On a large run that's thousands of GPUs burning money to do nothing - checkpoint a 70B every 30 minutes and the stall tax adds up fast.

Asynchronous checkpoint

GPUs
training job
Storage
checkpoint writes
GPU time spent training0.0 / 0 · 0.0 idle

The modern default ( DCP, NeMo, ). The state is snapshotted to fast CPU/GPU memory in a fraction of a second, then flushed to storage in the background while training keeps running. The GPU stall shrinks from minutes to near-zero, so teams can checkpoint more often (better fault tolerance) for less wasted compute. It leans even harder on high-bandwidth storage to absorb the burst.

GPUs trainingGPUs idle (stall)snapshot to CPU/GPU memoryblocking write to storagebackground flush

Inference has a pipeline too - and a flywheel

Serving a model isn't a single forward pass. A modern request often retrieves context with RAG - embed the query, search a vector index, stitch results into the prompt - and looks up online features from a feature store. And every interaction feeds the most valuable asset in AI: the data flywheel.

The data flywheel

Every interaction becomes training data, and every better model brings more interactions. Play a few turns.

1/15
more usagemore databetter modelturn 1Users hit the modelLog interactionsCurate & labelRetrain / fine-tuneShip a better model

… and the loop closes: more usage → more data → a better model → more usage. Whoever spins this fastest wins, and it all runs on the data pipeline.

The technology landscape

These are the names you'll meet in any AI-data conversation. Open table formats (Iceberg/Delta/Hudi) deserve a special call-out: they add , schema evolution, and time-travel on plain files in object storage - “database guarantees without a database” - turning a messy lake into a reliable .

The landscape, placed on the pipeline

Each tool sits where it does its job. Orchestrators run across every stage and managed platforms bundle several. Tap a tool to see its role.

Orchestration
schedules every stage below

1Ingest & stream

data arrives

2Storage & formats

table format over files over buckets

3Process & transform

clean, join, curate

4Serve & features

to training & inference

Platforms
bundle several stages as one managed service

Storage & formats

Iceberg / Delta / Hudi

Open table formats: ACID, schema evolution & time-travel on files.

The medallion architecture

A common way to organize a lakehouse. Same ELT idea: keep the raw, refine in layers.

Bronze

raw, untouched

Silver

cleaned & conformed

Gold

curated, business- & model-ready

What “good” looks like

A healthy AI data platform is judged on these six properties. They are what separate a reliable model factory from a data swamp.

Six properties of a healthy AI data platform

Each tile pairs the property with a schematic of what it looks like in practice.

Freshness

data reflects the real world now, not last quarter

Quality

filtered, deduped, validated - garbage in, garbage out

Lineage

you can trace every output back to its sources

Reproducibility

rerun the same pipeline, get the same dataset

Throughput

fast enough to keep the GPUs and users fed

Governance

access, privacy & retention are enforced, not hoped for

The pipeline feeds everything else

Curated data trains the models, fast storage keeps the GPUs busy, and the flywheel turns usage back into better data. The data pipeline is the substrate the rest of the stack runs on.