Systems & Infrastructure

Simulation & Digital Twins

Advanced

Everyone knows GPUs train and serve AI. The quieter third job is to simulate - to build photorealistic virtual worlds where robots, cars, and factories learn before they ever touch reality. This is the engine behind digital twins, synthetic data, and the world models now reshaping physical AI.

The third GPU workload class

We usually split AI into two jobs: training (teach the model) and inference (run it). Simulation is a distinct third class - and it leans on a different kind of GPU work: real-time physics and ray-traced rendering, not just the matmul that powers the other two. NVIDIA frames this as a three-computer model: one computer trains, one simulates, one infers on the machine itself.

Three computers, one loop

One computer trains, one simulates, one infers on the machine itself - and each hands its output to the next.

1/3
modeltested policyfield dataDGXin the data centerTrainDGXSimulateOmniverse / CosmosInferJetson AGX Thor

Field data from the machine lands in the data center, and DGX trains the next model.

Three words, in plain English

Before the stack, the vocabulary. These three terms come up in every conversation about physical AI - here is what they actually mean.

Three words, three pictures

Tap a picture for the plain-English definition. Drawings are schematic.

Digital twin

A living video-game copy of a real thing - a factory, a car, a city block - kept in sync with reality through sensor feeds. Change the twin, predict what happens to the real thing.

Why simulate at all

Real-world data is slow to collect, dangerous to gather (you cannot crash ten thousand real cars), and expensive to label by hand. A simulator runs faster than real time, spins up thousands of parallel instances, and labels every frame for free - because it generated the frame and already knows what is in it. The economics are not a little better; they are orders of magnitude better.

Real world vs simulation

Collecting and labeling data in the physical world is slow, dangerous, and expensive. A simulator runs faster than real time, in thousands of parallel instances, and labels itself. Race the two production lines to the same sample count.

501K

244×

cheaper

12K×

faster

1/44Press play to start the clock.

Real world

25 labelers · ~$0.22/sample

0 / 501K

elapsed

-

spent

$0.00

Simulation

2,000 streams (1 cell = 10) · ~$0.0009/sample

0 / 501K

elapsed

-

spent

$0.00

Time to finish (log scale)

1 s1 min1 hr1 day30 daysSimulation30 sReal world4.2 days12K× longer

Cost (log scale)

$1$100$10K$1MSimulation$451Real world$110.3K244× more

Both axes are logarithmic - each gridline is 10× the one before, so the simulation bars stay visible. The blue line is the race clock.

Illustrative model - unit rates are rough placeholders chosen to show the order-of-magnitude gap. Real economics vary by task, sensor suite, and GPU fleet.

Domain randomization: one scene, infinite data

A model trained on one perfect studio shot overfits to that studio. The fix is to deliberately randomize everything that does not matter - lighting, textures, camera angle, object pose - so the model is forced to learn the thing that does. Crank the variety high enough and the real world looks like just one more random variant. Drag the controls and watch a single base scene multiply into labeled training data.

One scene, infinite labeled training data

Start from a single base scene, then let the simulator randomize lighting, backgrounds, pose, and camera. Every variant comes with a perfectly accurate bounding box and class label for free - the model learns the object, not the studio it was shot in.

Base scene

bottlebase

Sliders reshape the base; each variant then re-rolls every parameter at random.

1 scene → 12 labeled samples

Auto-generated label · variant 1 (tap any image; example format)

{
  "image": "variant_001.png",
  "class": "bottle",
  "bbox_xywh": [0.34, 0.32, 0.27, 0.29],
  "rotation_deg": -37,
  "lighting_hue": 223
}

This is domain randomization: by training across wild variety, the model becomes robust to the real world it has never seen. OpenAI's robotic hand learned dexterous manipulation entirely in simulation this way, then transferred to real hardware.

The sim-to-real loop

Simulation is never perfect - there is always a gap between the virtual and the real, the gap. The trick is to make that gap shrink over time: deploy, watch where reality breaks the model, recreate those failures in the simulator, and retrain. The single most valuable output of the real world is its failures. The loop, not any one model, is the product.

The sim-to-real loop

Failures in reality are not setbacks - they are the highest-value training data there is. Capture them, replay them in simulation, and the model gets better every cycle. Play it, or step one phase at a time.

1/32

Train in sim. Millions of randomized episodes, faster than real time.

loop 1 of 8-real-world successTrain in simrandomized episodesDeploy to realship the policyCapture failurelog the edge caseFeed back to simrecreate as scenarios

Real-world success per deployment

sim-to-real gap

50%60%70%80%90%100%12345678loop (deployment) →perfect transferno deployment measured yet

Scenario library

0 captured

Empty - no real-world failures captured yet.

orange = just captured · violet = replayed in sim

Each completed loop closes the sim-to-real gap with sharply diminishing returns - the last few points are the hardest, which is why the loop never really stops.

Illustrative model - accuracy gains are a synthetic curve and the edge-case names are examples, to show the shape of the loop, not measured results.

NVIDIA Omniverse & the stack

Omniverse is the platform that ties all of this together. At its base sits - the shared 3D format that lets every tool agree on one scene. On top, the Kit engine renders and runs it; on top of that, domain apps for robotics, AVs, and synthetic data. Adoption is real: as of August 2025, 300,000+ downloads and 252+ enterprise deployments, with users including BMW, Toyota, Lucid, and TSMC.

The Omniverse stack, layer by layer

Each layer is built on the one below. Tap a layer or a piece to see what it does.

built on

built on

MayaBlenderRevitCATIA

authoring tools, one shared scene format

sim output

OpenUSD

the shared 3D source of truth

Universal Scene Description, the open standard from Pixar. The HTML of 3D: one format where CAD, physics, robot paths, IoT signals, and AI labels all coexist and interoperate across Maya, Blender, Revit, and CATIA.

In production · BMW

BMW simulates its entire 31-factory network as in Omniverse - reporting up to a 30% reduction in production-planning costs. A Replicator + Isaac Sim pipeline generates thousands of photorealistic, labeled training images "with one click," and BMW piloted an entire virtual factory 2+ years before series production began.

World models - the frontier

The next leap is not hand-building scenes but generating them. learn the dynamics of reality from video and then imagine new, interactive, physically-plausible worlds on demand - turning the simulator itself into a neural network. This is where simulation, generative AI, and robotics converge.

Three world models, one shape

What goes in, what the model is, and what comes out.

NVIDIA Cosmos

CES 2025

in

20 million hours of video

curated with NeMo Curator

model

World foundation models

4B to 14B params · Nano → Ultra · open model license

out

Physical AI builders

humanoids (1X, Agility, Figure AI) · AVs (XPENG, Uber with Waabi)

Curating the 20M hours, drawn to scale

CPU
3+ years
Blackwell
14 days

World foundation models for physical AI, unveiled at CES 2025. Models from 4B to 14B params (Nano → Ultra) under an open model license, trained on 20 million hours of video. NeMo Curator processed, curated, and labeled those 20M hours in 14 days on Blackwell - a job that would take 3+ years on CPU. Early adopters span humanoid robotics (1X, Agility, Figure AI) and AVs (XPENG, Uber with Waabi).

Google DeepMind Genie 3

August 2025

in

A text prompt

model

Genie 3

generates worlds in real time

out

An interactive world

24 fps, with memory of what it has already shown

Announced August 2025: generates real-time, interactive worlds from a text prompt at 24 fps, with memory of what it has already shown. Waymo reportedly adopted Genie 3 to build the Waymo World Model for AV development.

Wayve

in

Raw driving experience

not hand-coded rules

model

End-to-end embodied world model

simulation central to how it scales

out

An “AI driver”

learns to drive from experience

Builds end-to-end, embodied world models - an "AI driver" that learns to drive from raw experience rather than hand-coded rules, with simulation central to how it scales.

How much data does this actually take?

Simulation and training are, underneath the renders and the physics, a storage and data-movement problem. Follow the data for autonomous driving and it climbs by roughly three orders of magnitude at every step - from one car, to a fleet, to what a program keeps, to the simulated worlds that dwarf reality, to the corpus that trains the model.

The autonomous-driving data ladder

Every rung on one log axis - each gridline is 10× the one before. Tap a rung for its detail and source.

1 TB100 TB10 PB1 EB
  1. 01One vehicle, one day · 4–150 TB

    A production car streams ~4 TB/day from its cameras, LiDAR, and radar; a heavily-instrumented test rig hits 11–152 TB/day. A single hour of driving is 1–5 TB. - Intel / Tuxera ↗

  2. 02A test fleet, one day · 2–30 PB

    A ~200-vehicle fleet generates 2.2–30.4 PB of sensor data per day. The self-driving industry collectively ingests well over a petabyte every day. - Tuxera ↗

  3. 03What one program keeps · 200+ PB

    Mobileye stores 200+ petabytes of driving footage - roughly 16 million one-minute clips, about 25 years of continuous driving - kept live for training and replay. - Mobileye / Intel, CES 2022 ↗

  4. 04Then multiply by simulation · ~100×

    Waymo has driven ~200 million real autonomous miles - and 20+ billion miles in simulation. Roughly a hundred virtual miles for every real one, each one more data to generate, store, and replay. - Waymo ↗

  5. 05The training corpus · ~45 PB

    NVIDIA Cosmos was trained on 20 million hours of video - about 9,000 trillion tokens, an estimated ~45 PB raw and ~2 PB after curation. Curating it took 14 days on Blackwell instead of 3+ years on CPU. - NVIDIA, CES 2025 ↗

The ×100 arrow is the real-to-simulated miles multiplier, drawn from the 200+ PB rung to show the size of the step - it illustrates the multiplier, not a measured storage total.

Per-vehicle and fleet figures are widely-cited industry estimates; the Cosmos PB sizes are derived from NVIDIA's stated 20M-hour / 9,000T-token corpus. Mobileye reported the storage totals and Waymo the simulated-mile counts.

Robots learn from the same data pyramid

Embodied AI runs the identical loop - real demonstrations at the top, synthetic data generated in simulation underneath. A single released robot dataset is already the size of a small data lake, and simulation multiplies it further.

Robot data, drawn to scale

Two released real-robot datasets on one terabyte axis, and one synthetic run on one hours axis.

Real demonstrations · size on disk

Open X-Embodiment

~32 TB

1M+ trajectories

Over a million real-robot demonstrations pooled from 60 datasets across 22 robot types - the ImageNet moment for robot learning.

AgiBot World

~43.8 TB

~1M trajectories

Roughly 3,000 hours of humanoid manipulation from 100 robots - one released dataset that already weighs as much as a small data lake.

0 TB25 TB50 TB

Synthetic · time to generate

NVIDIA Isaac GR00T

780K sim trajectories in 11 hrs
Human teleoperation6,500 hrs · nine months
Simulation11 hrs

Same 780,000 trajectories on one hours axis - the simulation bar is a sliver, about 590× less wall-clock time (6,500 ÷ 11).

Simulation generated 780,000 synthetic trajectories - the equivalent of 6,500 hours (nine months) of human teleoperation - in just 11 hours. This is why synthetic data wins.

And none of it is cold storage

Every petabyte here has to stay live - read back constantly to curate, replay edge cases, generate synthetic variations, and retrain. The data isn't a byproduct of simulation; at this scale it is the product, and feeding it to the GPUs fast enough is the real bottleneck.

Simulation is a data factory

Every loop here produces enormous volumes of rendered frames, sensor logs, and labeled samples that have to be stored, curated, and fed back into training. A single autonomous-vehicle or robotics program routinely generates petabytes of sensor and rendered data per day - and all of it must stay live for curation, replay, and retraining. Simulation does not stand alone - it sits on the same GPU fleet and data platform as everything else, which is exactly the scale problem VAST is built for.