Applications
Data Pipeline for AI
FundamentalModels get the headlines, but a model is only as good as the data fed into it - and only as fast as the storage feeding the GPUs. This page starts with the plumbing in plain English, then goes deep on how AI data pipelines, lakes, and lakehouses actually work, and why high-performance storage is the hidden lever on cost and speed.
Start with the plumbing
A is plumbing for data: pipes that move it from where it's created (apps, sensors, the web) to where it's used (dashboards, models), cleaning and reshaping it along the way. Where the water is stored matters too - and there are three common reservoirs.
Three reservoirs - what actually sits inside each
A warehouse keeps modeled tables, a lake keeps raw files of every type, and a lakehouse keeps the same raw files with a table-format layer on top. Drawings are schematic.
Data warehouse
A tidy library
modeled tables only
Structured, schema-on-write tables. Data is cleaned and modeled before it lands.
Strong: Fast, reliable SQL & BI dashboards.
Watch out: Rigid; struggles with raw text, images, and exploratory AI work.
Data lake
A giant storage unit
raw files, any type
Raw files of any type (text, images, video, logs) dumped cheaply into object storage.
Strong: Holds everything; cheap; perfect for ML training corpora.
Watch out: No guarantees - easily becomes a “data swamp” with no order.
Lakehouse
A renovated warehouse on the lake
table format over raw files
A lake plus an open table format (Iceberg/Delta/Hudi) that adds database guarantees.
Strong: Warehouse reliability over lake-scale raw data - one system for BI and AI.
Watch out: Younger ecosystem; needs governance discipline to stay clean.
ETL vs ELT - the same steps, a different order
Switch modes and watch Transform and Load trade places.
The warehouse holds
ETL - Extract, Transform, Load
Clean and reshape data before loading it. Classic warehouse approach - you only store the polished result.
Great when the schema is known and storage is precious.
A normal pipeline vs. an AI pipeline
A traditional analytics pipeline ends at a clean table a human reads. An AI pipeline keeps going: the data has to become something a model can learn from or retrieve. That adds a whole new set of stages.
Same start, different finish
Both pipelines ingest and clean. The classic one stops at a table a human reads; the AI one keeps going and forks.
Shared start
Ingest
pull from apps & databases
Clean & join
fix types, merge sources
Classic analytics pipeline
Model & aggregate
build SQL tables
Dashboard
a human reads a chart
AI-centric pipeline (the extra stages)
Curation
select & weight sources by quality, not just volume
Deduplication
remove near-duplicate docs that waste compute & memorize
Tokenization
convert text to integer tokens, pack & shard for the GPUs
Embedding generation
encode chunks into vectors for retrieval (RAG)
Vector indexing
build ANN indexes so similar vectors are findable fast
Feature store
compute & serve the same features offline and online
Labeling / RLHF
human & AI feedback to align and rank model outputs
Same raw data, but it forks into a training corpus, a retrieval index, and a - each with its own quality bar.
Both forks start with the same unglamorous step: finding the data and bringing it in from wherever it lives today - old file shares, object buckets, collaboration apps - then keeping that copy in sync as the source changes. On VAST, that on-ramp is SyncEngine.
The data refinery: training data is manufactured, not found
Common Crawl is 9.5+ PB of raw web, ~250 TB and 2B+ pages in a single monthly snapshot - and over 80% of GPT-3's training tokens came from it. But you can't train on raw web; most of it is junk, boilerplate, and duplicates. Hugging Face's FineWeb distilled ~15 trillion high-quality tokens from 96+ snapshots - and beat bigger datasets by filtering and deduplicating harder, not crawling more. Watch the volume shrink at each stage.
The data refinery
Raw web in, training-ready tokens out. The stream's width is the data volume; each filter peels most of it away. Turn a filter off and see what gets through instead.
Survives the pipeline
3.46%
of the raw 250 TB snapshot
Training-ready output
15.00 T tokens
all filters on → FineWeb-scale ~15 T quality tokens
Turn filters off and the surviving volume balloons - but those extra “tokens” are duplicates and junk that hurt the model. FineWeb beat bigger datasets by filtering harder, not crawling more. Illustrative model; survival fractions approximate the real Common Crawl → FineWeb funnel.
Feeding the GPUs: compute is only as fast as the data
A training run is a race to keep thousands of GPUs busy. If storage can't stream training shards fast enough, GPUs sit idle and (MFU) - the fraction of theoretical compute you actually use - collapses. Idle GPUs at $3/hr each are pure burned cash. This is why high-performance storage is an AI-infrastructure decision, not a commodity one.
Keeping the GPUs fed
A GPU at $3/hr earns nothing while it waits on data. If storage can't supply bytes as fast as the fleet consumes them, utilization (MFU) drops and idle GPU-time burns cash. Pick a storage tier and watch the fleet.
fast but small, manual staging, no sharing
Pipe cross-section is proportional to bandwidth (thickness ∝ √GB/s). The colored bar at the outlet is the tier's ceiling.
GPU fleet (15/64 fed)
23% MFU
Burn rate while starved
$147/hr
49 idle GPUs at $3/hr each
Wasted since you arrived
$0.00
ticking in real time
Switch to the parallel all-flash filesystem and the ceiling jumps to ~1,000 GB/s - the fleet stays green even at full scale. Compute is only as fast as the data feeding it: high-performance storage is what turns expensive GPUs into useful FLOPs (the VAST value prop). Illustrative model; per-GPU demand and tier ceilings are approximate.
The other heavy I/O: checkpoints
Training also writes checkpoints - full snapshots so a crashed run can resume instead of starting over. With Adam mixed precision a checkpoint is roughly 16 bytes/param, so a 70B model is ~1 TB+. Frontier runs checkpoint frequently for fault tolerance, and the whole fleet stalls while writing - so fast write and fast read (for recovery) directly protect expensive GPU-hours.
Synchronous vs asynchronous checkpoints - same clock
Two checkpoints in the same window. Watch what the GPUs do while each one is written. Durations are illustrative units, not measurements.
At time 0: synchronous has stalled the GPUs for 0 units; asynchronous has paused them for 0.
Synchronous checkpoint
The classic approach: training pauses, copies the full model + optimizer state out of GPU memory, and waits for it to land on storage before resuming. Every GPU in the job sits idle for the whole write. On a large run that's thousands of GPUs burning money to do nothing - checkpoint a 70B every 30 minutes and the stall tax adds up fast.
Asynchronous checkpoint
The modern default ( DCP, NeMo, ). The state is snapshotted to fast CPU/GPU memory in a fraction of a second, then flushed to storage in the background while training keeps running. The GPU stall shrinks from minutes to near-zero, so teams can checkpoint more often (better fault tolerance) for less wasted compute. It leans even harder on high-bandwidth storage to absorb the burst.
Inference has a pipeline too - and a flywheel
Serving a model isn't a single forward pass. A modern request often retrieves context with RAG - embed the query, search a vector index, stitch results into the prompt - and looks up online features from a feature store. And every interaction feeds the most valuable asset in AI: the data flywheel.
The data flywheel
Every interaction becomes training data, and every better model brings more interactions. Play a few turns.
… and the loop closes: more usage → more data → a better model → more usage. Whoever spins this fastest wins, and it all runs on the data pipeline.
The technology landscape
These are the names you'll meet in any AI-data conversation. Open table formats (Iceberg/Delta/Hudi) deserve a special call-out: they add , schema evolution, and time-travel on plain files in object storage - “database guarantees without a database” - turning a messy lake into a reliable .
The landscape, placed on the pipeline
Each tool sits where it does its job. Orchestrators run across every stage and managed platforms bundle several. Tap a tool to see its role.
1Ingest & stream
data arrives
2Storage & formats
table format over files over buckets
3Process & transform
clean, join, curate
4Serve & features
to training & inference
Storage & formats
Iceberg / Delta / Hudi
Open table formats: ACID, schema evolution & time-travel on files.
The medallion architecture
A common way to organize a lakehouse. Same ELT idea: keep the raw, refine in layers.
Bronze
raw, untouched
Silver
cleaned & conformed
Gold
curated, business- & model-ready
What “good” looks like
A healthy AI data platform is judged on these six properties. They are what separate a reliable model factory from a data swamp.
Six properties of a healthy AI data platform
Each tile pairs the property with a schematic of what it looks like in practice.
Freshness
data reflects the real world now, not last quarter
Quality
filtered, deduped, validated - garbage in, garbage out
Lineage
you can trace every output back to its sources
Reproducibility
rerun the same pipeline, get the same dataset
Throughput
fast enough to keep the GPUs and users fed
Governance
access, privacy & retention are enforced, not hoped for
The pipeline feeds everything else
Curated data trains the models, fast storage keeps the GPUs busy, and the flywheel turns usage back into better data. The data pipeline is the substrate the rest of the stack runs on.