Tokenomics

How a Token's Cost Is Calculated

Fundamental

Every AI answer is paid for in GPU time. This page builds the cost of a token from scratch: one equation, checked against a real benchmark, then what happens inside a single request, what an hour of GPU is made of, and how serving cost turns into the price on an API page.

The one equation

A GPU is paid for by the hour, whether it is busy or not. While it runs a model it produces - pieces of words - at some rate per second. So the cost of a token is simply what the GPU costs per second, divided by how many tokens it makes per second. The industry quotes it , which gives one line of arithmetic:

Cost per million tokens

$ / M tokens=$ per GPU-hour × 1,000,000tokens/s per GPU × 3,600

The 3,600 turns seconds into hours; the 1,000,000 turns one token into a million. Flip it over and you get tokens per dollar = tokens/s per GPU × 3,600 ÷ $ per GPU-hour.

Check it against a real benchmark

DeepSeek R1 (4-bit weights) on a rack, 8K tokens in / 1K out per request, each user streaming at 162 tokens/s. In benchmarks (InferenceX by SemiAnalysis) each GPU handles 5,394 tokens/s and costs $2.31 per hour to own at hyperscale.

GPU cost per hour

$2.31

owner's all-in cost at hyperscale

Tokens per GPU per hour

5,394 × 3,600 = 19.4M

5,394 tokens/s in benchmarks

Cost per million tokens

$2.31 ÷ 19.4 = $0.119

the whole calculation

Tokens per dollar

19.4M ÷ $2.31 = 8.4M

the same number, flipped

The formula lands on $0.119 per million tokens - the same figure the benchmark publishes. A billion tokens costs about $119 of GPU time.

Always check the denominator. That benchmark counts input and output tokens together. In an 8K-in / 1K-out request only 1 token in 9 is output, so the same GPU writes about 600 output tokens/s; charge everything to output and the figure becomes about $1.07 per million output tokens. Both are correct - they answer different questions.

Play with the dials

The equation has two inputs, plus one that people forget: utilization. A GPU that makes tokens only a third of the day still bills all 24 hours, so every token it does make carries three times the cost. Start from the GB300 benchmark, jump to DeepSeek's own published day of serving, or see what an idle-heavy internal deployment pays. The number of GPUs per replica changes the hourly bill and the token output by the same factor, so it cancels out - which is why costs are quoted per GPU.

Token cost playground

Set what a GPU costs per hour and how many tokens it makes per second. Start from a real data point, then move one dial at a time.

DeepSeek R1 FP4, 8K in / 1K out, 162 tok/s per user, in benchmarks (InferenceX). Input and output tokens are counted together.

Where the hourly cost comes from

All-in owner's cost at hyperscaler volume (TCO model)

$ per GPU-hourAll-in: hardware, power, facility, staff$2.31
Tokens per second, per GPUSummed over every user sharing the GPU (log scale)5,394
GPUs per replicaGPUs that serve one copy of the model together72
UtilizationShare of paid hours spent making tokens at that rate100%

1 · Hourly bill

$166/hr

$2.31 × 72 GPUs

paid every hour, busy or not

2 · Tokens made

1.4B/hr

5,394 tok/s × 3,600 s × 72

at full utilization

3 · Cost per million

$0.119

$166 ÷ 1.4B × 1M

the GPU count cancels out

$ per million tokens

$0.119

Tokens per $1

8.4M

Cost of 1 billion tokens

$119

Idle hours still cost money

At 100% utilization every paid hour makes tokens. Slide utilization down to see what quiet hours do to the price of each token.

making tokensidle - still billed

Where a request's cost goes: input vs output

A single request runs in two phases. reads the whole prompt in one parallel pass: thousands of tokens share each trip through the model's weights, so the GPU's math units stay busy and each input token is cheap. then writes the answer one token at a time, and every step re-reads the weights and the conversation's from memory. That makes each output token several times more expensive in GPU time - about 3x in the example below - and list prices typically charge 3-6x (up to 8x) more for output than for input.

Step through one request and watch the meter. Then switch the request shape: long-prompt RAG questions and coding agents are dominated by input, while an agent's lets most of that input skip prefill entirely.

One request, traced

DeepSeek R1 on GB300 NVL72 at $2.31 per GPU-hour. Prefill and decode speeds are modeled from first principles (illustrative), not measured.

8,192 in / 1,024 out - The benchmark workload from the playground above: a long prompt and a medium answer, 8:1.

Prompt · 8,192 tokens

Answer · 1,024 tokens

Each square = 128 tokens
GPU time this request uses0.341 s total
prefill 0.220 s (65%)decode 0.121 s (35%)

GPU-seconds so far

0.0 ms

Cost so far

$0

$0 per million such requests

What the user waits

~0.0 s

answer streams at ~66 tok/s

1/6A request arrives with its prompt. Nothing has run yet.

One input token

27 µs of GPU time → $0.017 per million

~37,200 prompt tok/s per GPU at this prompt length

One output token

118 µs of GPU time → $0.076 per million

~8,450 tok/s per GPU, 128 conversations on a 32-GPU decode group - 4.4× an input token

The user waits ~16 s for the answer, but the request only uses 0.341 s of GPU time: the 32-GPU decode group works on 128 conversations at once. Decode speed is held fixed here; longer contexts slow it down, which the next page covers.

Why decode is memory-bound, how batching shares each step across users, and why longer contexts make output dearer are covered on Why Output Costs More: Batching & Speed.

From GPU-hour to list price

The dollars per GPU-hour in the equation is itself a stack of costs. For an owner, roughly 72% is : the servers, networking and storage, spread over their together with the that bought them. Facility space, cooling and power delivery add about 15%, staff and maintenance about 8%, and the electricity itself only about 6%.

A list price has to cover much more than a fully busy GPU. Real fleets sit partly idle at night, serve free-tier and internal traffic that nobody pays for, and run gateways, cache storage and support. In benchmarks at full load, GPU time is only a small fraction of a typical list price; the waterfall below adds the rest one bar at a time.

From GPU cost to list price

Every bar is tagged: benchmark-derived, an illustrative assumption you can move, or a list price.

Inside one GB300 NVL72 GPU-hour

$2.99/hr

Modeled owner's cost: 5-year life, 11% cost of capital, $0.08/kWh.

71%
15%

Capital $2.13 · 71%

depreciation + cost of money

Facility $0.44 · 15%

space, cooling, power delivery

Ops $0.24 · 8%

staff, maintenance, software

Electricity $0.18 · 6%

at the meter, including PUE

1,000 requests of 8,192 tokens in / 1,024 out

GPU cost: DeepSeek R1 FP4 on GB300 NVL72 at $2.31/GPU-hour, 5,394 tokens/s per GPU in benchmarks (InferenceX, input and output counted together).

$0.50 input / $2.15 output per million tokens. Same model as the benchmark, from an inference host.

UtilizationShare of paid GPU-hours doing work60%
Unbilled trafficFree tier, internal use, retries20%
Other serving costsOn top of GPU time+10%

Benchmark GPU cost vs price

17%

full load, before any overhead

Gross margin, this scenario

60%

before training and R&D

DeepSeek's published day

84.5%

theoretical gross margin, H800s, 2025-03-01

A real-world anchor: DeepSeek's published day

In March 2025 DeepSeek published 24 hours of its own serving numbers. On an average of 226.75 nodes of 8 H800 GPUs, priced at $2 per GPU-hour, the day cost $87,072. It processed 608 billion input tokens (56.3% of them prefix-cache hits) and wrote 168 billion output tokens - about 1,072 output tokens/s per GPU, or $0.52 per million output tokens. Charged at its R1 list prices, the same tokens would have earned $562,027: DeepSeek's "545% cost-profit margin", which is an 84.5% gross margin. DeepSeek noted that actual revenue was well below that, because of cheaper V3 pricing, free app and web use, and night-time discounts.

How capex, depreciation, power and financing build up the hourly figure for each GPU generation is on What an Hour of GPU Really Costs.

Energy per token

The same equation works for energy: joules per token = watts ÷ tokens per second. A GB300 GPU draws about 1.9 kW all-in (its share of CPUs, networking and cooling included); at 5,394 tokens/s that is about 0.35 J per token. Once power rather than GPUs is the scarce resource, operators flip it around and count . At good batch sizes electricity is only about 5-15% of a token's cost, and the batching that makes tokens cheap also makes them frugal: serving 64 users at once instead of one cuts energy per token about 40x on the same GPU.

Energy per token

Power draw is all-in per GPU (host CPUs, networking and cooling share included). PUE and electricity price are illustrative defaults.

Tokens per second, per GPULog scale5,394
PUEFacility power ÷ IT power1.20
Electricity price$ per kWh$0.08

Joules per token

0.35 J

1,900 W ÷ 5,394 tok/s

Electricity per 1M tokens

$0.0094

0.098 kWh at the IT load

Tokens/s per megawatt

2.4M

439 GPUs per facility MW

10K tok/s per MW100M

Electricity in one GB300 NVL72 GPU-hour ($2.31 at hyperscale): $0.182 (8%)

electricity hardware, facility, staff

The share does not depend on tokens/s: power and the rest of the bill are both paid per hour. Batching makes every part of a token cheaper at once.

What moves the number

Every tokenomics question comes back to the same fraction. These are the dials that move it, and where this site goes deeper on each.

Figures as of Sep 2026. Throughput and cost per million tokens come from InferenceX benchmarks (DeepSeek R1 FP4, 8K in / 1K out); $ per GPU-hour from the SemiAnalysis AI Cloud TCO model and a bottom-up owner's cost model; list prices from provider pricing pages; the real-world anchor from DeepSeek's published inference system overview (March 2025). Prefill and decode speeds in the request tracer and the Llama 3.3 70B energy points are first-principles estimates. Values marked illustrative are assumptions for teaching.