Tokenomics

What an Hour of GPU Really Costs

Intermediate

Every cost-per-token figure starts with one number: what a GPU costs for an hour. This page builds that number from the purchase price, the years it is spread over, power, the building and staff - then shows why idle hours, depreciation choices and power limits move it more than the sticker price.

Own or rent: three prices for the same hour

"$ per GPU-hour" means different things depending on who pays. An owner spreads the purchase over several years and adds power, the data center and people. A renter pays one hourly price that already includes all of that plus the cloud's margin. The gap between the two is the price of flexibility - and it is large.

Owner's cost

What it costs to own and run a GPU for an hour once the hardware, power, building and staff are paid for, at hyperscaler scale. SemiAnalysis's AI Cloud TCO model puts it at $1.17 for an H100 and $2.31 for a GB300.

Retail tier

The same model's second price tier, roughly twice the hyperscaler figure. It reads as a rental price level for smaller buyers rather than anyone's cost to own.

On-demand rent

The hourly price anyone can rent at, with no commitment. It carries the cloud's margin, its idle capacity and the flexibility you are buying - for Blackwell, 2.5-4x the owner's cost.

Per GPU-hour, H100 to Vera Rubin

H100 (HGX)

Owner (hyperscaler)
$1.17
Retail tier
$2.00
Capex per GPU
$28k-$42k
Typical rent
$2.5-2.7 (spot index); $3.6-6.9 list

H200 (HGX)

Owner (hyperscaler)
$1.22
Retail tier
$2.90
Capex per GPU
$32k-$48k
Typical rent
$3.3-4.8 (index); $4.3-6.3 list

B200 (HGX)

Owner (hyperscaler)
$1.73
Retail tier
$3.70
Capex per GPU
$45k-$68k
Typical rent
$5.8-8.1 (index); $6.7-14 list

B300 (HGX)

Owner (hyperscaler)
$2.26
Retail tier
$4.25
Capex per GPU
$58k-$85k
Typical rent
$6.8 (index); $7.9-17.8 list

GB200 NVL72

Owner (hyperscaler)
$1.86
Retail tier
$4.00
Capex per GPU
$52k-$66k
Typical rent
$10.50-10.58 list

GB300 NVL72

Owner (hyperscaler)
$2.31
Retail tier
$5.00
Capex per GPU
$61k-$90k
Typical rent
$9.72 list; mostly quote-only

Vera Rubin NVL72

Owner (hyperscaler)
$3.57
Retail tier
-
Capex per GPU
$69k-$153k
Typical rent
not yet rentable

Owner's cost and retail tier: SemiAnalysis's AI Cloud model as shown on InferenceX (Vera Rubin from an InferenceX post, July 2026). is our low-to-high range for one GPU's share of its server or rack plus networking and storage; analyst estimates for a Vera Rubin rack differ about 2x. Rent: price indices (Silicon Data, Ornn) and published on-demand list prices from neoclouds and AWS.

Building the number line by line

An owner's hourly cost has four lines. Capital spreads the purchase price over its and adds a tied up in it; one number, the capital recovery factor (CRF), does both. Electricity is the GPU's share of power - including its CPUs, network and switches - times the cooling overhead () times the power price. Facility is data-center space and cooling, usually priced per kW per month, and ops covers staff, spares and support.

=Capitalcapex x CRF(WACC, years) / 8,760
+ElectricitykW x PUE x $/kWh
+FacilitykW x $/kW-month / 730
+Opscapex x ops% per year / 8,760

Worked through for a GB300 owned by a large neocloud: $69,000 x 0.2706 / 8,760 = $2.13 of capital, plus $0.18 of electricity (1.9 kW x 1.20 x $0.08), $0.44 of facility and $0.24 of ops - $2.99 per GPU-hour. Capital is 70-75% of the total in every scenario, and power plus the building about 20%. That is why the purchase price and the years it is spread over matter far more than the electricity bill.

TCO builder - $ per GPU-hour, line by line

Pick a GPU and a kind of buyer, then move any input. The three buyer presets load the scenarios behind the figures on this page; everything else is yours to change.

Buying it

Running it

Using it

Default 5,394: GB300 NVL72 on DeepSeek R1, 8K in / 1K out at 162 tok/s per user, in benchmarks (InferenceX).

Per GPU-hour

$2.99

every hour

Per useful hour

$4.27

at 70% busy

Per M tokens

$0.220

at that usage

Every hour of the year$2.99
Per useful hour (70% busy)$4.27
Capital 71%Electricity 6%Facility 15%Ops 8%Idle hours you still pay for

The math, with your numbers

  • Capital$69,000 x 0.2706 CRF / 8,760 h$2.13
  • Electricity1.90 kW x 1.20 PUE x $0.080$0.18
  • Facility1.90 kW x $170 / 730 h$0.44
  • Ops$69,000 x 3% / 8,760 h$0.24
  • Per useful hour$2.99 / 0.70 = $4.27
  • Per M tokens$4.27 x 1M / (5,394 x 3,600) = $0.220

CRF (capital recovery factor) turns the purchase price into an equal yearly payment that repays it plus a return on the money over its life. At 100% utilization the same GPU would cost $0.154 per M tokens.

Against published figures - GB300 NVL72, $ per GPU-hour

Your cost, every hour$2.99
Your cost per useful hour$4.27
SemiAnalysis TCO, hyperscaler$2.31
SemiAnalysis retail tier$5.00
Rental price indicesnone
On-demand list prices$9.72
$0$3$6$9$12

Indices are Silicon Data and Ornn; list prices run from neoclouds up to AWS on-demand, Sep 2026. A renter pays only for hours it uses, so compare rent with your cost per useful hour.

Scenario inputs are modeling assumptions: hyperscaler 6 years, 8% cost of capital, $0.06/kWh, $120/kW-month, ops 2% of capex a year; large neocloud 5 years, 11%, $0.08, $170, 3%; smaller buyer 4 years, 14%, $0.12, $220, 4%. PUE is lower for liquid-cooled NVL72 racks (1.15-1.30) than for air-cooled HGX servers (1.30-1.45). For H100 through GB300 the hyperscaler case lands 4-12% under SemiAnalysis's figures.

Depreciation: the biggest assumption after price

How many years a GPU is spread over changes its hourly cost more than any input except the purchase price. A GB300 owned by a large neocloud costs $2.72 an hour on a 6-year life but $4.08 on a 3-year life - 50% more for the same machine. Useful life is a company's accounting choice, and in 2025-26 it became one of the most argued-over numbers in AI infrastructure.

Same GPU, different life - $ per GPU-hour

Every other input stays at the chosen buyer's values; only the years the purchase is spread over, and the cost of the money, change. Drag the life slider to read the curves.

$ per GPU-hour$0$1$2$3$4$5$62345678useful life, years →1234GB300B200H100

GB300

$4.08

+50% vs 6 years

B200

$3.27

+50% vs 6 years

H100

$2.16

+47% vs 6 years

  1. 14 years - Nebius
  2. 25 years - Amazon (a subset of servers, from 2025)
  3. 35.5 years - Meta
  4. 46 years - Microsoft, Alphabet, Oracle, CoreWeave

GB300 for this buyer: $4.08 per hour on a 3-year life vs $2.72 on 6 years. Lives are server depreciation policies from company filings and press, 2025-26.

The case for shorter lives

  • NVIDIA now ships a new architecture every 12-24 months, so a GPU's economic life can end before its accounting life does.
  • Amazon became the first hyperscaler to shorten: from 2025 it cut a subset of servers from 6 to 5 years, citing the pace of AI/ML technology.
  • H100 rental prices fell about 80% from 2023 to October 2025 - a sign that older GPUs lose earning power fast once a new generation ships.

The case for longer lives

  • Microsoft and Alphabet moved from 4 to 6 years, Oracle from 5 to 6, CoreWeave from 5 to 6 (2023), and Meta to 5.5 years (2025).
  • CoreWeave said H100s coming off contract were rebooked at about 95% of their original price, and in August 2026 disclosed an A100 contract running into 2029.
  • H100 one-year rental prices rose about 40% from October 2025 to March 2026. Still, a 2020-era GPU under contract in 2029 does not prove any single GPU lasts nine years.

A practical reading: 3 years is a cautious buyer's assumption, 4-5 years is typical for neoclouds (Nebius uses 4) and 6 years is hyperscaler accounting. When comparing anyone's cost per token, ask which life they assumed.

Utilization: you pay for the idle hours too

An owned or reserved GPU costs the same every hour, busy or not. So the cost that matters is the cost per useful hour: total cost divided by utilization. A GB300 at $2.31 an hour really costs $2.72 per useful hour at 85% utilization, $3.30 at 70%, $4.62 at 50% and $7.70 at 30%. Cost per token moves the same way - halving utilization doubles it.

Live inference is the hard case. Traffic follows the working day, and fleets are sized for the peak plus headroom so replies stay fast, which leaves the night half-empty. Nodes that fail, ramp up or wait for a whole rack to be repaired add more idle time.

A day of traffic on a fleet sized for the peak

Illustrative demand for a chat service on one 72-GPU rack. The fleet is sized for the afternoon peak plus headroom; every hatched hour is paid for and makes no tokens.

00:0006:0012:0018:0024:00hour of day →GPUs busy, as a share of peak demand0%50%100%GPUs you pay for (72)
Live trafficIdle, still paid for

Utilization

56%

752 idle GPU-hours a day

Per useful hour

$4.09

from $2.31 per hour

Per M tokens

$0.211

$0.119 if always busy

Cost per million tokens as utilization falls

$0.00$0.25$0.50$0.75$1.0010%25%50%75%100%utilization →you: $0.211

85%: $0.140 50%: $0.238 30%: $0.397

Demand shape, fleet size and batch fill (90% of capacity) are illustrative. The GPU cost and 5,394 tokens/s per GPU are GB300 NVL72 on DeepSeek R1 (8K in / 1K out, 162 tok/s per user) in benchmarks. Live traffic alone fills 56% of the fleet.

"Utilization" means three things

The share of hours a GPU is allocated to work, the share of those hours producing useful output (), and the share of peak math the chips achieve (). Only the first two divide the hourly cost. Meta's Llama 3 405B training ran at 38-43% MFU on 16K H100s - normal, and already inside tokens-per-second figures.

Enterprise fleets run cold

Cast AI's 2026 Kubernetes report measured average GPU compute utilization of about 5% across roughly 23,000 clusters on AWS, Azure and Google Cloud before optimization. That measures chip activity rather than allocation, but either way most paid-for hours produce nothing.

Filling the night

On one day in early 2025 DeepSeek's inference fleet peaked at 278 nodes but averaged 227, handing nodes to research and training at night. Its API now charges half price off-peak, and the big model APIs offer batch processing at 50% off - discounts that pull work into the quiet hours.

Rent or buy: where the lines cross

On-demand Blackwell rents for about $7-10 per GPU-hour at neoclouds and $14-18 at AWS, against an owner's cost of about $2-3 - a 2.5-4x markup for the freedom to stop paying at any time. Committed contracts of one to five years cut rent by 25-60%. The quick rule: divide the owner's hourly cost by the rental price. The result is the share of hours a GPU must be busy before owning wins. For a B300 at $3.00 against $7.85 on demand it is about 38%; below that, renting is cheaper.

Rent or buy one GPU - cumulative cash

Owning costs the purchase up front, then power, space and staff every hour. On-demand rent is paid only for the hours used; a reserved contract is paid for every hour, at a discount.

Owned as a

Default price: Nebius on-demand, Sep 2026. Commitment discounts run up to 35-60% at large neoclouds.

cumulative $ per GPU$0$50k$100k$150k$200k$250k$300k0122436486072months →end of 5-yr life
Own: $70k up front + $0.84/hr to runOn-demand: $7.85/hr x 60% of hoursReserved: $5.10/hr x every hour

Owning beats on-demand after

25 months

Owning beats reserved after

22 months

Break-even usage

38%

$3.00 owner's cost / $7.85 rent

Cash view: ignores the cost of capital and any resale value, which flatters owning a little. Break-even usage uses the full owner's cost per hour, capital included: if you need the GPU for fewer hours than that, renting on demand is cheaper over its life.

The same arithmetic is the business model of a GPU cloud. One megawatt of all-in power runs about 472 GB300s. Rented at $5 an hour they bring in about $20.7M a year against an owner's cost of about $9.5M at hyperscale; at $2.50 an hour revenue falls to about $10.3M, close to break-even. SemiAnalysis notes that the four- to five-year offtake contracts behind neocloud financing typically lock in a project return in the teens.

Power: a tenth of the cost, all of the constraint

At $0.08/kWh and a PUE of 1.2, a megawatt of electricity costs about $0.84M a year - roughly a tenth of the $7.5-10.5M a year it takes to own the GPUs that megawatt feeds. Yet power is what limits how many GPUs can be switched on. Data-center vacancy in the main North American markets hit a record-low 1.4% in the first half of 2026 (0.2% in Northern Virginia), and hyperscalers have said publicly that they hold GPUs they cannot yet power up for lack of ready data-center space. Each generation also packs more power into a rack.

Power per server or rack

DGX H100 server (8 GPUs)~10.2 kW
DGX B300 server (8 GPUs)~14.5 kW
GB200 NVL72 rack~120 kW
GB300 NVL72 rack~142 kW
Vera Rubin NVL72 rack (est. 190-230 kW)~210 kW
Rubin Ultra rack (roadmap)~600 kW

Maximum or rated power per system from NVIDIA reference architectures and OEM specs; Vera Rubin is an estimate range and Rubin Ultra a roadmap figure for the ~600 kW class. An air-cooled 8-GPU server fits an ordinary data-center rack; an NVL72 rack needs liquid cooling and about ten times the power delivery.

When megawatts are the scarce input, the number that matters is . The cost of owning a megawatt's worth of GPUs barely changes between generations, but the tokens that megawatt produces rise steeply - which is how a newer, pricier GPU ends up with a far lower .

One megawatt, six GPUs

H100$7.5M$0.520.46M
H200$7.8M$0.221.15M
B200$8.9M$0.102.82M
B300$10.4M$0.113.12M
GB200$8.7M$0.083.39M
GB300$9.6M$0.074.1M
Owner's cost per MW-year Tokens/s per MW$ per M tokens

Owner's cost per MW-year: SemiAnalysis hyperscaler $/GPU-hour x GPUs per all-in MW x 8,760 hours. Tokens: DeepSeek R1, 8K in / 1K out, FP8, 50 tok/s per user, in benchmarks (InferenceX, input and output tokens counted). $ per M tokens is our arithmetic from the two columns at full utilization.

The rental market in brief

GPU rental prices fell hard and then came back. The H100 is the bellwether: scarce in 2023, oversupplied by late 2025 as Blackwell arrived, then pulled back up in 2026 by inference and agent demand, memory-limited Blackwell supply and data centers short of power. That rebound is also why 6-year depreciation looked less reckless in 2026 than it did a year earlier.

  1. Mid-2023~$8

    Peak scarcity for H100 rentals.

  2. Jun 2025-44%

    AWS cuts P5 (H100) on-demand from $12.29 to $6.88 per GPU-hour.

  3. Oct 2025$1.70

    Low point of SemiAnalysis's H100 one-year rental index - about 80% below 2023.

  4. Jan 2026$2.20

    Silicon Data's index is up about 10% in four weeks; AWS raises Capacity Block prices about 15%.

  5. Mar 2026$2.35

    One-year contracts are up about 40% from the low, and SemiAnalysis finds on-demand capacity sold out across GPU types.

  6. Jul 2026+20%

    AWS raises Capacity Block prices again, across B300, B200 and H100/H200 instances.

  7. Sep 2026$2.56-2.72

    H100 rental indices (Ornn, Silicon Data). Nebius lifts on-demand H100 from $3.85 to $4.50 from October 1.

Where you rent matters as much as when. The same H100 lists at $6.88 an hour on AWS on-demand, $3.85 at Nebius and $3.56 at Verda. Rack-scale NVL72 systems are rarely sold by the hour at all: GB200 lists at about $10.50 per GPU-hour at CoreWeave and on AWS Capacity Blocks, and GB300 is mostly quote-only. ICE and Ornn announced cash-settled GPU compute futures in May 2026, subject to regulatory approval.

Figures as of September 2026; capex and rental prices vary widely by deal. Sources: SemiAnalysis's AI Cloud TCO model and newsletter, InferenceX, analyst bills of materials, company filings (10-K and 10-Q), provider pricing pages (AWS, CoreWeave, Nebius, Lambda, Crusoe, Verda), rental indices (Silicon Data, Ornn), CBRE data-center reports, EIA electricity prices and the Cast AI 2026 Kubernetes report. The equation these hours feed is on How a Token's Cost Is Calculated.