Tokenomics

Why Output Costs More: Batching & Speed

Intermediate

API price lists charge 3-8× more for output tokens than for input tokens. The reason is physical: generating text is limited by how fast a GPU can read memory, not by how fast it can calculate. This page builds that idea from one decode step up to the batch size, context length and speed decisions that set the cost of every output token.

Two phases, two bottlenecks

Serving a request happens in two phases. reads the whole prompt at once: thousands of tokens flow through the model's weights together, so the GPU's math units stay busy. then writes the reply one token at a time, and every new token needs a full trip through the model. For that trip the GPU streams every weight, plus the of the conversation so far, out of its memory. The math per token is small; the reading is huge.

In numbers, for Llama 3.3 70B with weights on one B300 with 8 TB/s of memory bandwidth:

Decode ceiling

70 GB of weights per token

8 TB/s ÷ 70 GB

= 114 tokens/s at most

That is for one user at 100% bandwidth. With a realistic 60% and 1 ms of overhead per step: about 64 tokens/s.

Math actually used

~2 FLOPs per weight per token

141 GFLOP × 64 tokens/s

= ~9 TFLOP/s of ~4,500

About 0.2% of the tensor cores. At batch 1 you pay for bandwidth, not FLOPs.

Prefill, same GPU

8K prompt: ~151 GFLOP per token

4.5 PF FP8 × 50% achieved

= ~14,927 tokens/s

Roughly 235× the single-user output rate, from the same hardware.

Engineers call this : FLOPs done per byte moved. A B300 can do about 560 FP8 FLOPs in the time it reads one byte (4.5 PFLOP/s ÷ 8 TB/s), so any work below that ratio waits on memory. Prefill does thousands of FLOPs per byte of weights; single-user decode does about 2. Step through both below.

Inside one step

Pick a setup and a batch, then play the step. Decode moves a lot of bytes to do a little math; prefill moves the same weights once and does a lot of math.

Setup

Phase

Users in the batch

FP8 weights and KV, one B300, 4K tokens of context per user; 60% of peak bandwidth, 50% of peak FLOP/s; illustrative.

HBM

268 GB · 8 TB/s

Read this step

Weights 70 GB

KV cache: 1 slice, 671 MB

4.8 TB/s achieved

Tensor cores

4.5 PF FP8 · 13.5 PF FP4

0.21% of peak FLOP/s in use

tokens appear at the end of the step

One decode step

15.7 ms total

Reading HBM
14.7 ms
Math
0.07 ms
Fixed overhead
1.00 ms

Reads and math overlap, so a step lasts as long as the slower of the two, plus the fixed overhead.

1/4Stream the weights

All 70 GB of weights stream out of HBM at about 4.8 TB/s - the same read whether 1 user is waiting or 64.

Bytes read

71 GB

per step, this GPU

Step time

15.7 ms

memory-bound

Per user

64 tok/s

64 tok/s per GPU

Tensor cores busy

0.21%

of peak FLOP/s

The two phases are introduced on Inference & Serving; how the KV cache is built and sized is on KV Cache.

Batching: one weight read, many users

If every decode step reads all the weights anyway, the obvious move is to serve many users in the same step. Serving engines do exactly that with : one pass over the weights produces the next token for every sequence in the batch. The weight read is shared; only each user's own KV cache adds bytes. The GPU makes far more tokens per second in total, and each user's stream gets a little slower.

UsersStepPer userPer GPU$ per M output
115.7 ms64 tok/s64 tok/s$13.1
816.7 ms60 tok/s479 tok/s$1.74
3220.1 ms50 tok/s1,595 tok/s$0.52
6424.5 ms41 tok/s2,609 tok/s$0.32
12833.5 ms30 tok/s3,823 tok/s$0.22

Llama 3.3 70B, FP8 weights and KV, one B300 at $3/GPU-hr, 4K context per user; simplified model.

Going from 1 to 64 users cuts the cost per output token about 41×, from $13.1 to $0.32 per million, while each user slows about 36%, from 64 to 41 tokens/s. A single-user token costs far more than any market price: batching is the business. Past roughly 100 users the gains flatten, because each extra user adds nearly as many KV bytes as it saves in shared weight reads.

Batch cost playground

Pick a model and a GPU, then drag the batch. Every step reads the weights once for the whole batch plus each user's own KV cache, so watch the cost per token fall while each user's stream slows - until the KV cache no longer fits in memory.

Dense: every step reads all 70B weights.

268 GB usable HBM, 8 TB/s, 4.5 PF dense FP8, 13.5 PF FP4.

Weight precision

KV cache precision

GPUs sharing the model

Batch: users decoding together272 fit at this context64 users
Context held per userprompt + reply so far, in KV cache4K tokens
GPU cost per hourB300 (HGX): $2.26 hyperscaler TCO, $4.25 retail$3.00

Step time

24.5 ms

limited by memory reads

Speed per user

41 tok/s

64 tok/s when alone

Throughput per GPU

2,609 tok/s

64 × 41 tok/s

$ per M output tokens

$0.32

41× cheaper than 1 user

HBM per GPU

128 GB of 268 GB

Weights 70 GBReserve 15 GBKV cache 43 GB

64 users × 4K tokens of KV fit; up to 272 would.

Read from HBM every step

113 GB

Weights 70 GB - shared by all 64KV 43 GB - one slice per user

Per output token: 1.1 GB of weights + 671 MB of KV. Weights still dominate, so each extra user is nearly free.

Throughput vs interactivity

02k4k6k0204060tokens/s per user (interactivity) →tokens/s per GPU1 user8 users256 users64 users · $0.32/M

Each point on the line is one batch size; small dots mark 1, 8, 64 and 256 users. Up and left is cheap but slow; right is fast but expensive. The dashed red part needs more KV cache than the GPU holds.

Reality check: this model vs benchmarks

This is a simplified roofline model: bytes over bandwidth, FLOPs over peak, plus a fixed overhead per step. Set against InferenceX results it comes out optimistic:

Llama 3.3 70B FP8 on B200, 1K in / 1K out, ~50 tok/s per user

This model

4,203 output tok/s per GPU (decode only)

In benchmarks

6,757 tok/s per GPU counting input and output - roughly 3,400 of them output

model ≈ 1.2× the benchmark

DeepSeek R1 FP4 on GB300 NVL72, 8K in / 1K out, ~100 tok/s per user

This model

5,429 tok/s per decode GPU; 2,504 output tok/s per GPU once prefill GPUs count

In benchmarks

11,906 tok/s per GPU counting input and output across prefill and decode GPUs, with MTP on - about 1,300 of them output

model ≈ 1.9× the benchmark

Why real systems land lower: achieved bandwidth and FLOP efficiency are often 30-50%, not 50-60%; expert all-to-all traffic and KV hand-offs take time; hot experts and ragged sequence lengths make every step wait for the slowest GPU; latency targets cap the batch; and batches are rarely perfectly full. Use the playground to see the shape of the tradeoff, not to quote a price.

Why output is priced 3-8× input

Put the two phases side by side on the same GPU and the price gap falls out. Llama 3.3 70B on a B300 at $3/GPU-hr reads an 8K prompt for about $0.056 per million input tokens, but generates for $0.32 per million output tokens with 64 users holding 4K conversations: 5.7×. If each of those users holds the full 8K prompt, output rises to $0.45 and the gap to 8.1×. DeepSeek R1 on with 8K prompts and 1K replies works out to about $0.017 vs $0.073, or 4.3×. Production numbers point the same way: in DeepSeek's March 2025 inference system overview, an 8-GPU H800 node processed about 73,700 input tokens/s in prefill against about 14,800 output tokens/s in decode, roughly 5×.

List prices in September 2026 follow suit: 5× is the default at OpenAI, Anthropic and Google's Flash models, while most open-weight APIs charge 3-4×. Four things drive the premium:

  • Different bottleneck. Prefill tokens arrive thousands at a time and fill the tensor cores; decode tokens arrive one per user per step and wait on memory bandwidth.
  • Output holds memory while it is written. A sequence keeps its KV cache in HBM for every step until the reply ends, and that memory caps how many users a GPU can serve.
  • Speed is sold with output. A faster stream needs a smaller batch, which makes each output token dearer; prefill only has to meet a target.
  • Cached input is cheaper still. A prompt-cache hit skips prefill compute entirely, which is why cached input is often priced at 10% of regular input or less.

One GPU-hour, two prices

How many GPU-seconds does a million tokens take? Prompt tokens stream through in big parallel passes; output tokens come one step at a time. Same GPU, same hourly cost - very different bills.

Setup

Prompt lengthlonger prompts: more attention math per token8K tokens
Users decoding togetherfewer users = faster streams64

FP8 weights and KV on one B300 at $3/GPU-hr; replies of 1K tokens; illustrative model.

Input (prefill)

$0.056 per M tokens

67 GPU-seconds per million tokens · 14,927 tok/s per GPU · compute-bound

Output (decode)

$0.45 per M tokens

541 GPU-seconds per million tokens · 1,850 tok/s per GPU · 29 tok/s per user · memory-bound

Output ÷ input

8.1×this setup (model)

list prices 3-8.3×1×2×3×5×8×10×20×3×: DeepSeek V4 Pro, Grok 4.7, Qwen3.8-Max, GLM-5.34×: deepseek-flash, MiniMax M3, gpt-oss-120b hosts5×: OpenAI GPT-6, Claude, Gemini Flash, Kimi K36×: Gemini 3.1 Pro8.3×: Gemini 3.5 Flash-Lite
  • 3×DeepSeek V4 Pro, Grok 4.7, Qwen3.8-Max, GLM-5.3
  • 4×deepseek-flash, MiniMax M3, gpt-oss-120b hosts
  • 5×OpenAI GPT-6, Claude, Gemini Flash, Kimi K3
  • 6×Gemini 3.1 Pro
  • 8.3×Gemini 3.5 Flash-Lite

Violet dots: standard-tier output price ÷ input price, Sept 2026. Orange line: this model's raw GPU-time ratio.

Prefill of a 8K prompt spends 7% of its math on attention. Fewer users per GPU make output dearer and widen the gap; longer prompts make input dearer and narrow it.

Longer context shrinks the batch

The KV cache grows with every token a user holds. Llama 3.3 70B stores 160 KB per token in FP8, so a 4K conversation takes about 0.67 GB and a 128K document about 21 GB. A B300 has roughly 183 GB left for KV after the weights and buffers: room for about 272 users at 4K, 34 at 32K and only 8 at 128K. Each step still reads about the same number of bytes, but it now produces a token for 8 users instead of hundreds.

With every user at full context and the GPU filled, the cost per output token goes from about $0.33 per million at 8K to about $5.35 at 128K - about 16× dearer. Attention design matters here: DeepSeek R1's keeps a compressed latent of about 34 KB per token in FP8, under a quarter of Llama 70B's cache despite a model ten times larger. Price lists show the effect too: OpenAI charges 2× for input and 1.5× for output above 272K tokens on GPT-5.5, and Google does the same above 200K on Gemini 3.1 Pro.

Longer context, fewer users, dearer tokens

Each row fills the GPU with as many users as its KV cache allows (up to 512), every one holding the full context - a deliberately pessimistic case that shows the direction.

Setup

KV cache precision

FP8 weights on one B300 at $3/GPU-hr; 60% MBU; illustrative.

4K
272 users
19 tok/s per user
$0.16 /M
8K
136 users
19 tok/s per user
$0.33 /M
16K
68 users
19 tok/s per user
$0.66 /M
32K
34 users
19 tok/s per user
$1.31 /M
64K
17 users
19 tok/s per user
$2.63 /M
128K
8 users
19 tok/s per user
$5.35 /M

From 8K to 128K context the cost per output token rises about 16×. Each step takes about as long either way - the GPU reads a similar number of bytes each time - but it produces far fewer tokens.

Cost bars use a log scale so every row stays visible.

How attention designs shrink the cache is on Attention Mechanisms; moving cache that no longer fits out to host memory and storage is on KV Cache.

The other levers, quantified

Every lever works on one of two numbers: the bytes read per step, or the tokens produced per step. Here is what each one is worth on its own, then what they add up to together.

Speculative decoding and MTP

2.4× cheaper

A small drafter, or the model's own multi-token-prediction (MTP) heads, proposes several tokens and the big model checks them in one step. Drafting 3 tokens at 80% acceptance yields 2.95 tokens per step. For Llama 70B with 32 users that takes each stream from 50 to 118 tokens/s and cost from $0.52 to $0.22 per million. DeepSeek-V3's technical report measured 85-90% acceptance for its second predicted token and about 1.8× faster decoding.

Model, assuming steps 25% longer; the gain shrinks at large batch, where the extra checking math counts.

FP8 to NVFP4 weights

26% cheaper

4-bit weights with a shared FP8 scale per 16 values take about 0.56 bytes each instead of 1. Llama 70B at 64 users goes from $0.32 to $0.24 per million; a lone user speeds up from 64 to 107 tokens/s. The gain is smaller at big batch because the KV reads don't shrink.

Model; quality at 4 bits must be checked per model.

FP8 KV cache

50% cheaper at 32K

Halving the bytes per cached token doubles the users that fit. Llama 70B at 32K context: BF16 KV fits 17 users ($2.63 per million), FP8 fits 34 ($1.31). Each user's speed barely changes - the GPU just carries twice as many users per byte read.

Model; the win grows with context length.

Mixture of experts

18× fewer FLOPs

DeepSeek R1 holds 671B parameters but runs 37B per token, so each token needs about 18× less math than a dense 671B model would. The catch: all 671B must sit in memory, and a batch of just 128 tokens touches about 98% of the experts in each layer, so decode still reads nearly all the weights. The byte savings arrive only when experts are spread across many GPUs, each fed a big batch.

Architecture from the DeepSeek-V3 technical report; the expert-touch share is arithmetic.

Wide expert parallelism on NVL72

about 1.9× in benchmarks

Spreading DeepSeek R1's ~385 GB of weights over 32 GPUs instead of 8 cuts the weights each GPU reads per step from ~48 GB to ~12 GB and frees room for more users. The simple model says ~4×; in benchmarks at 73 tokens/s per user, GB300 NVL72 delivers 6,969 tokens/s per GPU against 3,748 on HGX B300.

InferenceX results for DeepSeek R1; needs NVLink for the expert all-to-all traffic.

Utilization

60% busy = 1.7× dearer

A GPU costs the same per hour busy or idle. At 60% average load every token carries 1.7× the cost it would on a fully loaded GPU; at 30% it is 3.3×. That is why providers offer batch APIs and off-peak discounts - they fill the troughs.

Arithmetic: cost scales with 1 ÷ utilization.

Stack the levers

Llama 3.3 70B on one B300 at $3/GPU-hr. Switch levers on and off; after each one the GPU is refilled with as many users as memory allows while each user still gets at least the speed floor.

Context per user

Speed floor per userminimum tokens/s each stream must get25 tok/s
  1. Start: 1 user, BF16 weights and KV

    $25.4/M

    starting point

    1 user · 33 tok/s each

  2. Batch users together

    One weight read per step serves every user in the batch.

    $0.95/M

    27× cheaper

    35 users · 25 tok/s each

  3. FP8 weights

    Half the bytes of BF16 per weight: faster steps and more room for KV.

    $0.38/M

    2.49× cheaper

    87 users · 25 tok/s each

  4. FP8 KV cache

    Half the KV bytes per token: about twice the users fit.

    $0.19/M

    2.00× cheaper

    174 users · 25 tok/s each

  5. NVFP4 weights

    4-bit values plus a shared scale, about 0.56 bytes per weight.

    $0.15/M

    1.26× cheaper

    220 users · 25 tok/s each

  6. Speculative decoding

    Draft 3 tokens, 80% accepted: 2.95 tokens per step; steps 25% longer and verification math 4× (assumed).

    $0.060/M

    2.54× cheaper

    318 users · 44 tok/s each

  7. Real-world utilization 60%

    Traffic has peaks and troughs; idle GPU-hours still cost money.

    $0.099/M

    1.67× dearer

    318 users · 44 tok/s each

With the levers shown on: $25.4 → $0.099 per million output tokens at 4K context, about 256× cheaper.

Bars sit on a log scale of $ per million output tokens; each lever's bar spans the cost before and after it. Simplified roofline model with 60% memory bandwidth, 50% of peak FLOP/s and 1 ms per step; illustrative.

Takeaway: every deployment picks a point on one curve

Put together, inference economics come down to one curve: tokens per second per GPU, which sets cost, against tokens per second per user, which sets how fast replies stream. Bigger batches slide along it toward cheap, slower tokens; long contexts pull the whole curve down; 4-bit weights, FP8 KV, and wide push it up and out. A provider's "fast" tier and its batch discount are two points on the same curve.

Figures as of Sep 2026. The cost model on this page is simplified: a decode step lasts as long as its bytes over achieved HBM bandwidth or its FLOPs over achieved peak, whichever is longer, plus a fixed overhead, with 60% bandwidth and 50% compute efficiency assumed; it counts output tokens on decode GPUs unless stated. Sources: GPU specs from NVIDIA product pages; model geometry from published model configs and the DeepSeek-V3 technical report; benchmark results from InferenceX by SemiAnalysis; production figures from DeepSeek's inference system overview (March 2025); list prices from provider pricing pages (Sep 2026); method after kipply's transformer inference arithmetic and the roofline model. New to the core equation? Start with How a Token's Cost Is Calculated.