AI Economics Wiki

Inference vs. Training Costs

Training creates or adapts model capability as a periodic investment; inference spends compute every single time that capability is used.

AI EconomicsUpdated Aug 13, 202611 min read

Snapshot

What it is

Training cost is the cost of creating or adapting model parameters — data preparation, training and fine-tuning runs, failed experiments, evaluation, and the engineering that makes the result reproducible. Inference cost is the cost of using a trained model on production requests: input, cached and output tokens, retrieval, tools, guardrails, retries, routing, and serving capacity.

Why it matters

Training is a periodic investment that can be amortized. Inference scales with every customer action, so it sits directly on gross margin. Confusing the two produces both kinds of error — refusing an investment that would pay back in months, and shipping a prototype whose per-task cost never clears its price.

What it is not

The token line on your provider invoice. Tokens are one input to one layer of a cost stack that also contains retrieval, tools, observability, human review, and idle reserved capacity.

Key takeaways

  • The decision unit is a successful customer task at the promised quality and latency, not a token and not an API call.

  • Output tokens are the expensive half of most rate cards — 5x input on Claude Haiku 4.5 and 6x input on gpt-5.6-sol as of August 2026.

  • Fine-tuning is justified by recurring net saving per successful task, not by comparing a training invoice with an API invoice.

  • Token counts are not comparable across models. A tokenizer change can eat most of a headline price cut.

What is the difference between training and inference cost?#

Training and inference are stages of the same production system, with very different cash profiles.

StageWhat it buysCost behaviourWho usually pays it
Pre-training
General capability from a large corpus
Very large, lumpy, one-time per model
Frontier labs; almost never a startup
Post-training / fine-tuning
Domain or task-specific behaviour
Project-shaped: data, compute, evaluation, iteration, maintenance
The startup, if it can show a recurring per-task saving
Inference
Applying the model to a live request
Variable with usage, or fixed if you reserve capacity
The startup, every day, forever

Three consequences follow. First, most startups buy pre-trained capability and own only the inference curve. Second, model-development choices change that inference curve: the Chinchilla result showed that under a fixed training-compute budget, a 70-billion-parameter model trained on substantially more data outperformed a much larger model — and a smaller model is cheaper to serve for its entire life (Hoffmann et al., 2022). Third, the boundary is no longer clean. Reasoning systems spend compute again at request time, so "the model" is really a policy for allocating compute across training, routing, and each live task.

Why does this distinction matter to founders?#

Pricing and gross margin. If inference is a meaningful variable cost, a flat subscription hides loss-making power users. If most of your cost is reserved capacity, pure per-token pass-through underprices the availability and latency guarantee you are actually selling. The price architecture should track the dominant cost driver without exposing customers to a meter they cannot predict — see Pricing Metric / Value Metric.

Build, buy, and model choice. Fine-tuning can shorten prompts, raise task success, or let a smaller model do the job. It is economically justified only when those recurring gains exceed the full cost of data, experiments, evaluation, deployment, and ongoing maintenance. The comparison is risk-adjusted lifecycle cost for equivalent outcomes, not one invoice against another.

Fundraising. Investors want to know whether AI cost falls with routing, caching, and volume, or grows with every user action. A credible deck separates five things: model-development investment; variable inference cost per successful task; fixed serving and reliability cost; support and human-review cost caused by model failures; and margin sensitivity to longer outputs, retries, and vendor price changes.

How do you calculate cost per successful task?#

For a period with N successful tasks:

total AI cost = amortized training + fixed serving + variable inference + expected failure cost

For API-priced models, the variable model layer is no longer a single rate:

variable model cost = (uncached input x input rate)
                    + (cache reads x cache-read rate)
                    + (cache writes x cache-write rate)
                    + (output tokens x output rate)

Then normalise to outcomes, not calls:

inference cost per successful task
  = (model + retrieval + tools + guardrails + retries + observability) / successful tasks

The denominator matters more than founders expect. A workflow that costs $0.01 per call, averages six calls, and succeeds 80% of the time costs at least $0.01 x 6 / 0.80 = $0.075 per successful task before tools and human review — 7.5x the naive per-call figure.

For owned or reserved capacity, add the idle you paid for:

serving cost per task = (reserved compute + orchestration + operations) / successful tasks
                      + burst cost per task

When does a fine-tune pay for itself?#

Let T be the incremental cost of fine-tuning and productionising a model, and s the recurring net saving per successful task against the best alternative:

break-even tasks = T / s

Use a net saving that already subtracts evaluation, retraining, monitoring, and any quality loss. Discount future savings if break-even takes years, and cap the horizon at the point where the base model will be superseded anyway. A break-even that lands beyond your model-refresh cycle is not a break-even.

A model-selection scorecard#

MeasureWhat to capture
Task success
Pass rate on representative production cases, not benchmarks
Cost
Fully loaded cost per successful task
Latency
Median and tail latency; tails drive churn
Capacity
Throughput and throttling behaviour at peak
Variance
Cost and quality spread across customer segments
Maintenance
Evaluation, prompt, data, and model-update burden
Concentration
Dependency on one provider, model, or accelerator

Choose on the efficient frontier: an option is dominated only if another is cheaper, faster, and at least as reliable.

Key Facts

01

Output is the expensive half of the meter

As of 13 August 2026, Claude Haiku 4.5 is $1 per million input tokens and $5 per million output tokens — a 5x ratio — while gpt-5.6-sol is $5 input and $30 output, a 6x ratio. Output length, not prompt length, is usually the first margin lever. (Anthropic pricing; )

OpenAI pricing
02

Caching changes the arithmetic before any model change does

Anthropic prices a cache read at 0.1x the base input rate, a five-minute cache write at 1.25x and a one-hour write at 2x — so caching pays for itself after one read on the five-minute cache and two reads on the one-hour cache (verified 13 August 2026).

Anthropic pricing
03

Asynchronous work costs half

Both providers discount batch processing by 50% on input and output — Claude Sonnet 5 falls from $2/$10 to $1/$5 per million tokens, gpt-5.6-sol from $5/$30 to $2.50/$15 (13 August 2026). Anything not latency-sensitive belongs there. (Anthropic pricing; )

OpenAI pricing
04

Long context is a separate price tier

On gpt-5.6-sol, input rises from $5 to $10 and output from $30 to $45 per million tokens above the short-context threshold; Gemini 3.1 Pro doubles input from $2 to $4 and raises output from $12 to $18 for prompts over 200k tokens (13 August 2026). Retrieval that stuffs context can silently move you into a higher rate band. (OpenAI pricing; )

Gemini API pricing
05

Compute-optimal training lowers serving cost too

Under a matched training-compute budget, a 70B model trained on far more data beat a 280B model — and remains cheaper to serve for its whole production life.

Hoffmann et al., "Training Compute-Optimal Large Language Models," 2022

Worked example: does the fine-tune clear its bar?#

An AI operations product handles 2,000,000 successful tasks per month. Each task averages 900 input tokens and 300 output tokens. It runs on Claude Haiku 4.5 at the rates published on 13 August 2026: $1.00 per million input tokens, $0.10 per million cache-read tokens, $5.00 per million output tokens.

Step 1 — per-task model cost, no caching

input  = 900 / 1,000,000 x $1.00 = $0.00090
output = 300 / 1,000,000 x $5.00 = $0.00150
total                            = $0.00240

Step 2 — build the monthly stack

Retries, guardrails, and evaluation traffic add 25%. Fixed deployment and observability cost $6,000. A $120,000 fine-tuning and productionisation programme is assessed over a 12-month horizon, so the decision model allocates $10,000 per month.

Cost layerMonthly cost
Base model calls (2,000,000 x $0.0024)
$4,800
Retries, guardrails, evaluation (+25%)
$1,200
Fixed serving and observability
$6,000
Training investment, 12-month allocation
$10,000
Total
$22,000

Fully loaded AI cost is $22,000 / 2,000,000 = $0.0110 per successful task.

Step 3 — check the cheap lever first

Route 60% of input through the prompt cache at $0.10 per million tokens:

input = (900 x 0.40 x $1.00 + 900 x 0.60 x $0.10) / 1,000,000 = $0.000414
per-task model cost = $0.000414 + $0.00150 = $0.001914   (down 20.3%)
monthly total = ($0.001914 x 2,000,000 x 1.25) + $6,000 + $10,000 = $20,785

Caching cuts the model bill by 20.3% but the total by only 5.5%, because fixed serving and the training allocation dominate. That gap is the whole point of separating the layers.

Step 4 — test the fine-tune

Suppose the fine-tuned model saves a net $0.004 per successful task through shorter prompts, fewer retries, and smaller-model routing:

break-even = $120,000 / $0.004 = 30,000,000 successful tasks
at 2,000,000 tasks/month = 15 months
at 5,000,000 tasks/month = 6 months

Fifteen months is longer than the 12-month evaluation horizon, so the project fails on cost savings alone. It needs additional value — higher conversion, better quality, lower latency, or strategic control over a capability — or it needs volume growth. At 5 million tasks a month the same project pays back in six months and the decision flips.

Caveat on this model. Straight-line amortisation of $120,000 / 12 implies the fine-tune retains full value for exactly twelve months and none afterwards. It does neither. Treat the allocation as a decision aid, not an accounting entry, and re-run it against the date you expect to migrate to a newer base model.

What are the common mistakes?#

  • Treating the token price as total cost. Retrieval, vector storage, tools, moderation, observability, retries, support, and human review routinely exceed the model bill in workflow products.
  • Dividing by requests instead of outcomes. Failed and repeated calls consume real resources. Reliability then masquerades as efficiency.
  • Calling all model work "training." Prompt engineering, retrieval, and routing are inference-system work with inference economics. Pre-training, fine-tuning, and evaluation have different cash profiles and different depreciation.
  • Ignoring output variance. Long-tail outputs drive both cost and latency. Report p50, p90, and p99 by customer segment, not an average.
  • Comparing token counts across models. Anthropic notes that Claude 4.7 and later models use a newer tokenizer producing roughly 30% more tokens for the same text, so a per-token price comparison against an earlier model overstates the saving. See Token Economics and Metering.

When does this framework break?#

When quality cannot be scored. Average expected cost is not a sufficient control for medical, financial, security, or safety-critical workflows where one failure dominates thousands of successes. Those need worst-case loss limits and human accountability, not a cost-per-task target.

When the meter is not a unit of value. A token is a provider billing primitive. Two systems with identical token counts can deliver different outcomes because tokenizers, reasoning behaviour, tool protocols, and caching rules differ.

When the curve moves under you. Better models, specialised accelerators, and software optimisation push unit cost down; richer modalities, agentic loops, and higher reliability requirements push it up. Provider rate cards also change on their own schedule. Use scenarios with explicit price and volume ranges, not a single forecast.

Frequently asked questions

01

Should we fine-tune or improve the prompt and retrieval first?

Prompt, retrieval, caching, and routing changes almost always win on payback because they cost engineering time rather than a capital-shaped project, and they survive a base-model change. Reach for fine-tuning when you have measured a persistent quality or cost gap that context engineering cannot close, and you can name the recurring per-task saving.

02

How should we amortise training cost in a board deck?

Show it as a separate line, not inside gross margin, and state the horizon and the assumed obsolescence date. Then show gross margin both including and excluding it. Investors read a blended number as a claim about steady-state economics, which it usually is not.

03

Is inference cost a fixed or variable cost?

It depends on how you buy it. Per-token API pricing is variable. Provisioned throughput and reserved accelerators are fixed until consumed productively. Most products run a mix, which is why utilisation belongs in the cost model — see GPU and Compute Economics.

04

Do falling token prices guarantee improving margins?

No. Per-token prices have fallen while contexts, output lengths, reasoning steps, and tool calls have grown. Track dollars per successful task over time; that is the series that determines margin, and it can rise while every rate card falls.

05

What single metric should we instrument first?

Fully loaded cost per successful task, segmented by customer and workflow, with attempts-per-success reported alongside it. Everything else in this page is a decomposition of that number.

Sources#

  1. Jordan Hoffmann et al., "Training Compute-Optimal Large Language Models", arXiv:2203.15556, March 2022. The Chinchilla result on allocating training compute between model size and data, and why a smaller compute-optimal model is cheaper to serve.
  2. Anthropic, Claude Platform pricing documentation, accessed 13 August 2026. Per-model input, cache-write, cache-read, and output rates; the 50% Batch API discount; cache multipliers of 1.25x, 2x, and 0.1x; and the note that Claude 4.7 and later use a tokenizer producing approximately 30% more tokens for the same text.
  3. OpenAI, API pricing, accessed 13 August 2026. Standard, Batch, Flex, and Fast-mode rates per million tokens, including short- and long-context bands and cache-write pricing on the GPT-5.6 family.
  4. Google, Gemini Developer API pricing, accessed 13 August 2026. Paid-tier input and output rates by model and service tier, and the higher rate band for prompts above 200k tokens on Gemini 3.1 Pro.
  5. Amazon Web Services, How tokens are counted in Amazon Bedrock, accessed 13 August 2026. Model-specific output-token burndown rates and the distinction between quota consumption and billable usage.

Price-freshness note: Every rate quoted on this page was checked against the provider's own documentation on 13 August 2026. AI rate cards change on the provider's schedule, not yours — re-verify before using any figure in a board pack, a customer quote, or a pricing decision.

Author

Dr. Sarah Zou

Independent economist · EconNova

Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.

About Sarah

Topics

AI economicstraining costinference costtokenscomputeunit economicspricinggross margin

Cite this page

Suggested citation

Zou, S. (2026). Inference vs. Training Costs: How AI Founders Model Compute Economics. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/inference-vs-training-costs

Open license

Reuse with attribution

This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.

Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.