Tokenomics
How a Token's Cost Is Calculated
FundamentalEvery AI answer is paid for in GPU time. This page builds the cost of a token from scratch: one equation, checked against a real benchmark, then what happens inside a single request, what an hour of GPU is made of, and how serving cost turns into the price on an API page.
The one equation
A GPU is paid for by the hour, whether it is busy or not. While it runs a model it produces - pieces of words - at some rate per second. So the cost of a token is simply what the GPU costs per second, divided by how many tokens it makes per second. The industry quotes it , which gives one line of arithmetic:
Cost per million tokens
The 3,600 turns seconds into hours; the 1,000,000 turns one token into a million. Flip it over and you get tokens per dollar = tokens/s per GPU × 3,600 ÷ $ per GPU-hour.
Check it against a real benchmark
DeepSeek R1 (4-bit weights) on a rack, 8K tokens in / 1K out per request, each user streaming at 162 tokens/s. In benchmarks (InferenceX by SemiAnalysis) each GPU handles 5,394 tokens/s and costs $2.31 per hour to own at hyperscale.
GPU cost per hour
$2.31
owner's all-in cost at hyperscale
Tokens per GPU per hour
5,394 × 3,600 = 19.4M
5,394 tokens/s in benchmarks
Cost per million tokens
$2.31 ÷ 19.4 = $0.119
the whole calculation
Tokens per dollar
19.4M ÷ $2.31 = 8.4M
the same number, flipped
The formula lands on $0.119 per million tokens - the same figure the benchmark publishes. A billion tokens costs about $119 of GPU time.
Always check the denominator. That benchmark counts input and output tokens together. In an 8K-in / 1K-out request only 1 token in 9 is output, so the same GPU writes about 600 output tokens/s; charge everything to output and the figure becomes about $1.07 per million output tokens. Both are correct - they answer different questions.
Play with the dials
The equation has two inputs, plus one that people forget: utilization. A GPU that makes tokens only a third of the day still bills all 24 hours, so every token it does make carries three times the cost. Start from the GB300 benchmark, jump to DeepSeek's own published day of serving, or see what an idle-heavy internal deployment pays. The number of GPUs per replica changes the hourly bill and the token output by the same factor, so it cancels out - which is why costs are quoted per GPU.
Token cost playground
Set what a GPU costs per hour and how many tokens it makes per second. Start from a real data point, then move one dial at a time.
DeepSeek R1 FP4, 8K in / 1K out, 162 tok/s per user, in benchmarks (InferenceX). Input and output tokens are counted together.
Where the hourly cost comes from
All-in owner's cost at hyperscaler volume (TCO model)
1 · Hourly bill
$166/hr
$2.31 × 72 GPUs
paid every hour, busy or not
2 · Tokens made
1.4B/hr
5,394 tok/s × 3,600 s × 72
at full utilization
3 · Cost per million
$0.119
$166 ÷ 1.4B × 1M
the GPU count cancels out
$ per million tokens
$0.119
Tokens per $1
8.4M
Cost of 1 billion tokens
$119
Idle hours still cost money
At 100% utilization every paid hour makes tokens. Slide utilization down to see what quiet hours do to the price of each token.
Where a request's cost goes: input vs output
A single request runs in two phases. reads the whole prompt in one parallel pass: thousands of tokens share each trip through the model's weights, so the GPU's math units stay busy and each input token is cheap. then writes the answer one token at a time, and every step re-reads the weights and the conversation's from memory. That makes each output token several times more expensive in GPU time - about 3x in the example below - and list prices typically charge 3-6x (up to 8x) more for output than for input.
Step through one request and watch the meter. Then switch the request shape: long-prompt RAG questions and coding agents are dominated by input, while an agent's lets most of that input skip prefill entirely.
One request, traced
DeepSeek R1 on GB300 NVL72 at $2.31 per GPU-hour. Prefill and decode speeds are modeled from first principles (illustrative), not measured.
8,192 in / 1,024 out - The benchmark workload from the playground above: a long prompt and a medium answer, 8:1.
Prompt · 8,192 tokens
Answer · 1,024 tokens
GPU-seconds so far
0.0 ms
Cost so far
$0
$0 per million such requests
What the user waits
~0.0 s
answer streams at ~66 tok/s
One input token
27 µs of GPU time → $0.017 per million
~37,200 prompt tok/s per GPU at this prompt length
One output token
118 µs of GPU time → $0.076 per million
~8,450 tok/s per GPU, 128 conversations on a 32-GPU decode group - 4.4× an input token
The user waits ~16 s for the answer, but the request only uses 0.341 s of GPU time: the 32-GPU decode group works on 128 conversations at once. Decode speed is held fixed here; longer contexts slow it down, which the next page covers.
Why decode is memory-bound, how batching shares each step across users, and why longer contexts make output dearer are covered on Why Output Costs More: Batching & Speed.
From GPU-hour to list price
The dollars per GPU-hour in the equation is itself a stack of costs. For an owner, roughly 72% is : the servers, networking and storage, spread over their together with the that bought them. Facility space, cooling and power delivery add about 15%, staff and maintenance about 8%, and the electricity itself only about 6%.
A list price has to cover much more than a fully busy GPU. Real fleets sit partly idle at night, serve free-tier and internal traffic that nobody pays for, and run gateways, cache storage and support. In benchmarks at full load, GPU time is only a small fraction of a typical list price; the waterfall below adds the rest one bar at a time.
From GPU cost to list price
Every bar is tagged: benchmark-derived, an illustrative assumption you can move, or a list price.
Inside one GB300 NVL72 GPU-hour
$2.99/hr
Modeled owner's cost: 5-year life, 11% cost of capital, $0.08/kWh.
Capital $2.13 · 71%
depreciation + cost of money
Facility $0.44 · 15%
space, cooling, power delivery
Ops $0.24 · 8%
staff, maintenance, software
Electricity $0.18 · 6%
at the meter, including PUE
1,000 requests of 8,192 tokens in / 1,024 out
GPU cost: DeepSeek R1 FP4 on GB300 NVL72 at $2.31/GPU-hour, 5,394 tokens/s per GPU in benchmarks (InferenceX, input and output counted together).
$0.50 input / $2.15 output per million tokens. Same model as the benchmark, from an inference host.
GPU time at full load benchmark
9.216M tokens × $0.119 per million
Idle capacity illustrative
GPUs busy 60% of paid hours
Unbilled traffic illustrative
20% of tokens are free tier or internal
Other serving costs illustrative
+10%: gateway, cache storage, networking, support
Serving cost result
what 1,000 requests really cost
Gross margin result
60% of the list price
List price list price
$0.50 in / $2.15 out per million
Benchmark GPU cost vs price
17%
full load, before any overhead
Gross margin, this scenario
60%
before training and R&D
DeepSeek's published day
84.5%
theoretical gross margin, H800s, 2025-03-01
A real-world anchor: DeepSeek's published day
In March 2025 DeepSeek published 24 hours of its own serving numbers. On an average of 226.75 nodes of 8 H800 GPUs, priced at $2 per GPU-hour, the day cost $87,072. It processed 608 billion input tokens (56.3% of them prefix-cache hits) and wrote 168 billion output tokens - about 1,072 output tokens/s per GPU, or $0.52 per million output tokens. Charged at its R1 list prices, the same tokens would have earned $562,027: DeepSeek's "545% cost-profit margin", which is an 84.5% gross margin. DeepSeek noted that actual revenue was well below that, because of cheaper V3 pricing, free app and web use, and night-time discounts.
How capex, depreciation, power and financing build up the hourly figure for each GPU generation is on What an Hour of GPU Really Costs.
Energy per token
The same equation works for energy: joules per token = watts ÷ tokens per second. A GB300 GPU draws about 1.9 kW all-in (its share of CPUs, networking and cooling included); at 5,394 tokens/s that is about 0.35 J per token. Once power rather than GPUs is the scarce resource, operators flip it around and count . At good batch sizes electricity is only about 5-15% of a token's cost, and the batching that makes tokens cheap also makes them frugal: serving 64 users at once instead of one cuts energy per token about 40x on the same GPU.
Energy per token
Power draw is all-in per GPU (host CPUs, networking and cooling share included). PUE and electricity price are illustrative defaults.
Joules per token
0.35 J
1,900 W ÷ 5,394 tok/s
Electricity per 1M tokens
$0.0094
0.098 kWh at the IT load
Tokens/s per megawatt
2.4M
439 GPUs per facility MW
Electricity in one GB300 NVL72 GPU-hour ($2.31 at hyperscale): $0.182 (8%)
The share does not depend on tokens/s: power and the rest of the bill are both paid per hour. Batching makes every part of a token cheaper at once.
What moves the number
Every tokenomics question comes back to the same fraction. These are the dials that move it, and where this site goes deeper on each.
Tokens per second per GPU
The denominator. Batching many users, 4-bit weights, a smaller KV cache and speculative decoding all raise it - and batching is the biggest lever by far.
Why Output Costs More: Batching & Speed
Speed promised to each user
Faster streams mean smaller batches, so the same GPU makes fewer tokens. Moving along that tradeoff swings cost per token several-fold on the same hardware.
Why Output Costs More: Batching & Speed
Dollars per GPU-hour
The numerator. Owning at hyperscale, owning at smaller scale and renting differ about 2x; capital is roughly 72% of an owner's hour.
What an Hour of GPU Really Costs
Utilization
Idle hours are paid for too. Halving utilization doubles the cost of every token, which is why providers discount off-peak hours and batch jobs.
What an Hour of GPU Really Costs
Workload shape
Input tokens are cheap to process, output tokens are not, and cached input is cheaper still. The input:output mix of chat, RAG and agents changes the bill.
KV Cache
Figures as of Sep 2026. Throughput and cost per million tokens come from InferenceX benchmarks (DeepSeek R1 FP4, 8K in / 1K out); $ per GPU-hour from the SemiAnalysis AI Cloud TCO model and a bottom-up owner's cost model; list prices from provider pricing pages; the real-world anchor from DeepSeek's published inference system overview (March 2025). Prefill and decode speeds in the request tracer and the Llama 3.3 70B energy points are first-principles estimates. Values marked illustrative are assumptions for teaching.