AI Economics Wiki
GPU and Compute Economics
Compute economics turns accelerator capacity, throughput, and reliability into cost per successful customer outcome.
Snapshot
What it is
The analysis of how paid accelerator capacity becomes reliable, billable customer work. The decision unit is not the GPU-hour. It is cost per accepted output, successful task, or service-level-compliant request.
Why it matters
Three variables dominate — the all-in price of capacity, productive utilisation, and useful throughput. Utilisation alone can double unit cost without a single price changing. A commitment signed at the wrong utilisation converts a variable experiment into fixed burn.
What it is not
A hardware comparison. Peak FLOPS, memory bandwidth, and list price predict almost nothing about your cost per accepted output, which depends on model, precision, sequence length, batch density, interconnect, and software.
Key takeaways
Reservations are much less of a discount than founders assume. At August 2026 AWS rates, an H200 Capacity Block is only about 13% below on-demand, so break-even utilisation is roughly 87%.
Never merge scheduling losses and runtime failures into one utilisation number.
A newer accelerator costs more per hour. On current AWS rates a B200 is 2.38x an H100 per accelerator-hour — it improves economics only if it beats 2.38x on your accepted throughput.
Spot capacity is genuinely cheap and genuinely interruptible; classify workloads before you buy.
On this page10 sections
What does an accelerator-hour actually buy?#
An accelerator-hour is one GPU or other accelerator available to you for one hour. It is a capacity input, not an output. A working model connects five layers, and value leaks at every join:
contracted capacity → scheduled capacity → productive runtime → accepted output → customer revenue
Maintenance windows, fragmentation, queueing, model loading, failed jobs, retries, low batch density, and rejected outputs each remove a slice of the capacity you paid for. The purchasing model then decides which of those risks you carry:
| Purchasing model | You get | You carry | Fits |
|---|---|---|---|
On-demand | Start and stop freely | Highest hourly rate; availability risk at peak | Bursty, customer-facing, unproven |
Reserved / Capacity Blocks | Guaranteed capacity for a fixed window, paid upfront | Every idle hour in the window | Steady base load, scheduled training |
Committed use / Savings Plans | A lower rate for a term commitment | Term risk if the workload moves | Predictable multi-quarter demand |
Spot / preemptible | The lowest price by a wide margin | Preemption at any time | Checkpointable batch, evaluation, offline jobs |
Google Cloud states that Spot prices are dynamic and deliver 60% to 91% off the corresponding on-demand price for most machine types and GPUs, with smaller discounts on some accelerator families, and that Spot VMs can be preempted at any time (Google Cloud, GPU pricing). AWS charges the Capacity Block reservation fee upfront at purchase, prices it dynamically against supply and demand, and bills the operating system separately (AWS, Capacity Blocks pricing).
Why does this matter to founders?#
It decides whether gross margin scales. AI revenue can grow while margin falls if customers adopt longer contexts, premium models, more retries, or low-density real-time workloads. Per-seat pricing is especially exposed, because usage is unbounded while revenue is not — see Gross Margin.
Capacity is a product promise. A reservation can look expensive at average load and still be necessary for latency, regional residency, or a contractual service level. The right way to record that is an explicit risk value, not a fudge buried in the rate.
Optimisation trades against quality. Batching, quantisation, smaller-model routing, caching, and output caps all raise throughput. They can also reduce accuracy, lengthen tails, or change the failure distribution. Measure accepted outcomes, not benchmark tokens.
Investors will interrogate the bridge. Margin improvement that comes from durable engineering is worth more than margin improvement that comes from a one-off cloud credit or an optimistic allocation. A capacity model makes the difference auditable.
How do you calculate cost per accepted output?#
Start with what you paid for:
available accelerator-hours = accelerator count x hours in period
Then separate the two very different loss mechanisms:
productive accelerator-hours = available hours x scheduled utilisation x runtime success rate
Keep those two factors apart. A scheduling problem is solved by better packing and demand shaping; a runtime-failure problem is solved by engineering. One blended percentage hides which of the two you have.
accepted output = productive hours x throughput per accelerator-hour x acceptance rate
all-in cost = accelerator charges + host CPU/RAM + storage + networking
+ orchestration + observability + directly attributable operations
cost per accepted unit = all-in cost / accepted output
Reservation break-even#
If ancillary costs are identical under both options:
break-even utilisation = reserved capacity cost / (available hours x on-demand rate)
Below that point on-demand is cheaper on a pure cost basis; above it the reservation wins. The decision still needs a value for supply assurance, interruption risk, and engineering complexity — but that value should be written down separately, not smuggled into the rate.
Classify before you buy#
| Workload | Buy |
|---|---|
Steady and latency-sensitive | Reservation or committed capacity |
Bursty and customer-facing | On-demand, optionally over a reserved floor |
Checkpointable batch | Spot or spare capacity |
Rare experiments | On-demand |
Scheduled training runs | Time-bounded reservation or capacity block |
Worked example: does the reservation pay?#
An AI company is deciding how to buy eight NVIDIA H200 accelerators — one p5en.48xlarge — for a month. Rates are US East, verified 13 August 2026: AWS Capacity Block at $54.920 per instance-hour ($6.865 per accelerator-hour); on-demand at $63.296 per instance-hour ($7.912 per accelerator-hour) and spot at $27.209 ($3.401 per accelerator-hour) per the Vantage EC2 price tracker, which is a third-party source rather than an AWS rate card.
available accelerator-hours = 8 x 730 = 5,840
reserved accelerator cost = 5,840 x $6.865 = $40,091.60
host, storage, network, observability, on-call = $14,600
total monthly cost = $54,691.60
The workload produces 1.8 million accepted output tokens per productive accelerator-hour.
| Productive utilisation | Productive hours | Accepted output | Cost per million accepted tokens |
|---|---|---|---|
35% | 2,044 | 3,679.2M | $14.87 |
65% | 3,796 | 6,832.8M | $8.00 |
85% | 4,964 | 8,935.2M | $6.12 |
Nothing about the hardware price changed across those rows. Utilisation alone moves unit cost by 2.4x.
Now test the reservation against on-demand.
break-even utilisation = $40,091.60 / (5,840 x $7.912) = 86.8%
| Utilisation | Reserved total | On-demand total | Reservation saves |
|---|---|---|---|
35% | $54,691.60 | $30,772.13 | −$23,919.47 |
65% | $54,691.60 | $44,633.95 | −$10,057.65 |
85% | $54,691.60 | $53,875.17 | −$816.43 |
This is the finding that most capacity models get wrong. Because the current AWS Capacity Block rate is only 13.2% below on-demand for this instance, the reservation has to run at almost 87% productive utilisation before it wins on price alone. At a realistic 65% it is over $10,000 a month more expensive.
That does not make the reservation wrong — it makes the justification different. If on-demand p5en capacity cannot actually be obtained during peak demand, the reservation is buying supply assurance, and the right way to write it up is: "we are paying $10,058 a month for guaranteed access." That is a defensible sentence. "The reservation is cheaper" is not.
Acceptance rate compounds everything. Hold utilisation at 65% and vary the share of output the product accepts:
100% accepted → $8.00 per million accepted tokens
90% accepted → $8.89
75% accepted → $10.67
A 25-point drop in acceptance costs more than the entire reservation discount.
Caveat on this model. The on-demand and spot rates come from a third-party tracker, and spot pricing moves continuously; the Capacity Block rate is AWS's own and AWS notes prices are updated regularly, with the next scheduled change in October 2026. Re-pull all three before signing anything. The model also assumes ancillary cost is identical across purchasing options, which is true for storage and network and false for the engineering time that checkpointing for spot requires.
Key Facts
The reservation discount is thinner than founders expect
For p5en.48xlarge (8x H200) in US East, an AWS Capacity Block is $54.920 per instance-hour against roughly $63.296 on-demand — about 13% — putting break-even utilisation near 87% (13 August 2026). (AWS, Capacity Blocks pricing; )
Newer accelerators cost proportionally more per hour
AWS Capacity Block rates per accelerator-hour in US East are $5.191 for H100, $6.865 for H200, $12.355 for B200, and $14.04 for B300 — a B200 is 2.38x an H100 per hour and only improves unit economics if it clears 2.38x on your accepted throughput (13 August 2026).
AWS, Capacity Blocks pricingSpot is a real discount and a real risk
Google Cloud states Spot prices run 60% to 91% below the corresponding on-demand price for most machine types and GPUs, and that Spot VMs may be preempted at any time.
Google Cloud, GPU pricingCommitted-use discounts have a lock attached
On Google Cloud, GPUs qualify for resource-based committed-use discounts only if you create and attach a reservation, and that attached reservation cannot be modified or deleted for the duration of the commitment.
Google Cloud, GPU pricingEven the supplier's margin moves with the platform transition
NVIDIA reported fiscal 2026 revenue of $215.9 billion, up 65%, with GAAP gross margin of 71.1% — down year on year as the business shifted from Hopper systems to full Blackwell datacentre solutions. Hardware pricing is not a fixed backdrop.
NVIDIA, Q4 and fiscal 2026 results, February 2026What are the common mistakes?#
- Treating list price as all-in cost. Host CPU and memory, storage, egress, orchestration, observability, idle time, and on-call are part of the denominator's cost, and AWS bills the operating system on Capacity Blocks separately from the reservation fee.
- Reading device utilisation as productive utilisation. A GPU can be 100% busy producing output nobody accepts.
- Comparing hardware on peak specifications. Production throughput depends on model, precision, sequence length, batch size, interconnect, and serving software. Benchmark your own workload.
- Assuming a reservation is always cheaper. As above, at current AWS rates it usually is not, unless utilisation is very high or supply is genuinely scarce.
- Running one utilisation scenario. Model 35%, 65%, and 85% and include the cost of exiting or repurposing the commitment in each.
When does this model break?#
When capacity is shared. Cross-customer batching, shared caches, heterogeneous accelerators, and mixed training-plus-inference fleets defeat a single denominator. Use job-level telemetry and an explicit allocation rule, and label allocations as allocations — see Token Economics and Metering.
When quality moves. Throughput stops being a valid denominator the moment acceptance changes. A cheaper stack that produces more rejected answers has not improved economics; it has moved the cost somewhere the dashboard cannot see. Define the acceptance metric jointly with product and risk owners.
When supply or architecture shifts faster than the commitment. New hardware, model compression, provider price changes, and demand migration can strand capacity mid-term. AWS Capacity Block reservation fees are paid upfront and priced dynamically; Google Cloud commitment-attached reservations cannot be deleted for the term. Both mean the exit option you assumed may not exist.
Frequently asked questions
01Should an early-stage AI product own compute at all?
Usually not. Until you have a stable base load with a known utilisation, a metered API is cheaper in cash and dramatically cheaper in optionality — see Inference vs. Training Costs.
02How high does utilisation need to be to justify a reservation?
Compute it rather than assuming it. At current AWS H200 rates the arithmetic break-even is about 87%. Anything below that is a purchase of supply assurance, and should be written up and approved as one.
03Is spot capacity usable for customer-facing inference?
Only behind a reserved or on-demand floor that can absorb a preemption without breaching latency. Spot is excellent for evaluation runs, batch generation, offline enrichment, and checkpointed training.
04How should we compare accelerator generations?
On cost per accepted output for your workload, not on price per hour or on published throughput. Divide the per-accelerator-hour rate by the accepted output you measure yourself. A 2.38x price step needs more than a 2.38x throughput step to be worth taking.
05What belongs in the compute line of a board deck?
Available hours, scheduled utilisation, runtime success rate, acceptance rate, all-in cost, and cost per accepted unit — with the same five lines for the prior period. That set makes the margin bridge checkable.
Related concepts#
- Inference vs. Training Costs — separate recurring delivery cost from model-development investment.
- Token Economics and Metering — attribute compute consumption to customers and tasks.
- AI-as-a-Service — connect capacity to packaging and service levels.
- AI Unit Economics and Gross Margins — carry cost per accepted unit into cost of revenue.
- Hardware-as-a-Service — compare capacity ownership with a service contract.
- Usage-Based Pricing — charge for a controllable unit of consumption.
- Pricing Metric / Value Metric — avoid exposing a technical input as the customer meter.
- Economies of Scale — test whether serving cost actually falls at higher volume.
Sources#
- Amazon Web Services, Amazon EC2 Capacity Blocks for ML pricing, accessed 13 August 2026. Per-instance and per-accelerator reservation rates for P5, P5e, P5en, P6-B200, P6-B300 and Trainium instances; the upfront reservation fee; the separate operating-system charge; and the note that prices are updated regularly with the next change scheduled for October 2026.
- Vantage, p5en.48xlarge pricing and specs, price snapshot 13 August 2026. Third-party tracker used for the on-demand and spot comparison rates; not an AWS rate card, and should be re-checked against the console before any commitment.
- Google Cloud, GPU pricing, accessed 13 August 2026. Sustained-use and committed-use discount mechanics, the requirement to attach an unmodifiable reservation for GPU committed-use discounts, and the 60–91% Spot discount range with preemption at any time.
- NVIDIA, NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026, February 2026. Fiscal 2026 revenue of $215.9 billion, Data Center revenue growth, and GAAP gross margin of 71.1% alongside the Hopper-to-Blackwell transition.
Price-freshness note: Every compute rate on this page was checked on 13 August 2026 against the source named beside it. AWS states Capacity Block reservation prices are updated regularly and are next scheduled to change in October 2026, and spot prices move continuously — re-verify before using any figure in a board pack, a customer quote, or a capacity commitment.
Author
Dr. Sarah Zou
Independent economist · EconNova
Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.
About SarahTopics
Cite this page
Suggested citation
Zou, S. (2026). GPU and Compute Economics: Capacity, Utilisation, and Cost per Outcome. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/gpu-compute-economics
Open license
Reuse with attribution
This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.
Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.