AI Economics Wiki
Inference vs. Training Costs
Training creates or adapts model capability as a periodic investment; inference spends compute every single time that capability is used.
Snapshot
What it is
Training cost is the cost of creating or adapting model parameters — data preparation, training and fine-tuning runs, failed experiments, evaluation, and the engineering that makes the result reproducible. Inference cost is the cost of using a trained model on production requests: input, cached and output tokens, retrieval, tools, guardrails, retries, routing, and serving capacity.
Why it matters
Training is a periodic investment that can be amortized. Inference scales with every customer action, so it sits directly on gross margin. Confusing the two produces both kinds of error — refusing an investment that would pay back in months, and shipping a prototype whose per-task cost never clears its price.
What it is not
The token line on your provider invoice. Tokens are one input to one layer of a cost stack that also contains retrieval, tools, observability, human review, and idle reserved capacity.
Key takeaways
The decision unit is a successful customer task at the promised quality and latency, not a token and not an API call.
Output tokens are the expensive half of most rate cards — 5x input on Claude Haiku 4.5 and 6x input on gpt-5.6-sol as of August 2026.
Fine-tuning is justified by recurring net saving per successful task, not by comparing a training invoice with an API invoice.
Token counts are not comparable across models. A tokenizer change can eat most of a headline price cut.
On this page10 sections
What is the difference between training and inference cost?#
Training and inference are stages of the same production system, with very different cash profiles.
| Stage | What it buys | Cost behaviour | Who usually pays it |
|---|---|---|---|
Pre-training | General capability from a large corpus | Very large, lumpy, one-time per model | Frontier labs; almost never a startup |
Post-training / fine-tuning | Domain or task-specific behaviour | Project-shaped: data, compute, evaluation, iteration, maintenance | The startup, if it can show a recurring per-task saving |
Inference | Applying the model to a live request | Variable with usage, or fixed if you reserve capacity | The startup, every day, forever |
Three consequences follow. First, most startups buy pre-trained capability and own only the inference curve. Second, model-development choices change that inference curve: the Chinchilla result showed that under a fixed training-compute budget, a 70-billion-parameter model trained on substantially more data outperformed a much larger model — and a smaller model is cheaper to serve for its entire life (Hoffmann et al., 2022). Third, the boundary is no longer clean. Reasoning systems spend compute again at request time, so "the model" is really a policy for allocating compute across training, routing, and each live task.
Why does this distinction matter to founders?#
Pricing and gross margin. If inference is a meaningful variable cost, a flat subscription hides loss-making power users. If most of your cost is reserved capacity, pure per-token pass-through underprices the availability and latency guarantee you are actually selling. The price architecture should track the dominant cost driver without exposing customers to a meter they cannot predict — see Pricing Metric / Value Metric.
Build, buy, and model choice. Fine-tuning can shorten prompts, raise task success, or let a smaller model do the job. It is economically justified only when those recurring gains exceed the full cost of data, experiments, evaluation, deployment, and ongoing maintenance. The comparison is risk-adjusted lifecycle cost for equivalent outcomes, not one invoice against another.
Fundraising. Investors want to know whether AI cost falls with routing, caching, and volume, or grows with every user action. A credible deck separates five things: model-development investment; variable inference cost per successful task; fixed serving and reliability cost; support and human-review cost caused by model failures; and margin sensitivity to longer outputs, retries, and vendor price changes.
How do you calculate cost per successful task?#
For a period with N successful tasks:
total AI cost = amortized training + fixed serving + variable inference + expected failure cost
For API-priced models, the variable model layer is no longer a single rate:
variable model cost = (uncached input x input rate)
+ (cache reads x cache-read rate)
+ (cache writes x cache-write rate)
+ (output tokens x output rate)
Then normalise to outcomes, not calls:
inference cost per successful task
= (model + retrieval + tools + guardrails + retries + observability) / successful tasks
The denominator matters more than founders expect. A workflow that costs $0.01 per call, averages six calls, and succeeds 80% of the time costs at least $0.01 x 6 / 0.80 = $0.075 per successful task before tools and human review — 7.5x the naive per-call figure.
For owned or reserved capacity, add the idle you paid for:
serving cost per task = (reserved compute + orchestration + operations) / successful tasks
+ burst cost per task
When does a fine-tune pay for itself?#
Let T be the incremental cost of fine-tuning and productionising a model, and s the recurring net saving per successful task against the best alternative:
break-even tasks = T / s
Use a net saving that already subtracts evaluation, retraining, monitoring, and any quality loss. Discount future savings if break-even takes years, and cap the horizon at the point where the base model will be superseded anyway. A break-even that lands beyond your model-refresh cycle is not a break-even.
A model-selection scorecard#
| Measure | What to capture |
|---|---|
Task success | Pass rate on representative production cases, not benchmarks |
Cost | Fully loaded cost per successful task |
Latency | Median and tail latency; tails drive churn |
Capacity | Throughput and throttling behaviour at peak |
Variance | Cost and quality spread across customer segments |
Maintenance | Evaluation, prompt, data, and model-update burden |
Concentration | Dependency on one provider, model, or accelerator |
Choose on the efficient frontier: an option is dominated only if another is cheaper, faster, and at least as reliable.
Key Facts
Output is the expensive half of the meter
As of 13 August 2026, Claude Haiku 4.5 is $1 per million input tokens and $5 per million output tokens — a 5x ratio — while gpt-5.6-sol is $5 input and $30 output, a 6x ratio. Output length, not prompt length, is usually the first margin lever. (Anthropic pricing; )
OpenAI pricingCaching changes the arithmetic before any model change does
Anthropic prices a cache read at 0.1x the base input rate, a five-minute cache write at 1.25x and a one-hour write at 2x — so caching pays for itself after one read on the five-minute cache and two reads on the one-hour cache (verified 13 August 2026).
Anthropic pricingAsynchronous work costs half
Both providers discount batch processing by 50% on input and output — Claude Sonnet 5 falls from $2/$10 to $1/$5 per million tokens, gpt-5.6-sol from $5/$30 to $2.50/$15 (13 August 2026). Anything not latency-sensitive belongs there. (Anthropic pricing; )
OpenAI pricingLong context is a separate price tier
On gpt-5.6-sol, input rises from $5 to $10 and output from $30 to $45 per million tokens above the short-context threshold; Gemini 3.1 Pro doubles input from $2 to $4 and raises output from $12 to $18 for prompts over 200k tokens (13 August 2026). Retrieval that stuffs context can silently move you into a higher rate band. (OpenAI pricing; )
Gemini API pricingCompute-optimal training lowers serving cost too
Under a matched training-compute budget, a 70B model trained on far more data beat a 280B model — and remains cheaper to serve for its whole production life.
Hoffmann et al., "Training Compute-Optimal Large Language Models," 2022Worked example: does the fine-tune clear its bar?#
An AI operations product handles 2,000,000 successful tasks per month. Each task averages 900 input tokens and 300 output tokens. It runs on Claude Haiku 4.5 at the rates published on 13 August 2026: $1.00 per million input tokens, $0.10 per million cache-read tokens, $5.00 per million output tokens.
Step 1 — per-task model cost, no caching
input = 900 / 1,000,000 x $1.00 = $0.00090
output = 300 / 1,000,000 x $5.00 = $0.00150
total = $0.00240
Step 2 — build the monthly stack
Retries, guardrails, and evaluation traffic add 25%. Fixed deployment and observability cost $6,000. A $120,000 fine-tuning and productionisation programme is assessed over a 12-month horizon, so the decision model allocates $10,000 per month.
| Cost layer | Monthly cost |
|---|---|
Base model calls (2,000,000 x $0.0024) | $4,800 |
Retries, guardrails, evaluation (+25%) | $1,200 |
Fixed serving and observability | $6,000 |
Training investment, 12-month allocation | $10,000 |
Total | $22,000 |
Fully loaded AI cost is $22,000 / 2,000,000 = $0.0110 per successful task.
Step 3 — check the cheap lever first
Route 60% of input through the prompt cache at $0.10 per million tokens:
input = (900 x 0.40 x $1.00 + 900 x 0.60 x $0.10) / 1,000,000 = $0.000414
per-task model cost = $0.000414 + $0.00150 = $0.001914 (down 20.3%)
monthly total = ($0.001914 x 2,000,000 x 1.25) + $6,000 + $10,000 = $20,785
Caching cuts the model bill by 20.3% but the total by only 5.5%, because fixed serving and the training allocation dominate. That gap is the whole point of separating the layers.
Step 4 — test the fine-tune
Suppose the fine-tuned model saves a net $0.004 per successful task through shorter prompts, fewer retries, and smaller-model routing:
break-even = $120,000 / $0.004 = 30,000,000 successful tasks
at 2,000,000 tasks/month = 15 months
at 5,000,000 tasks/month = 6 months
Fifteen months is longer than the 12-month evaluation horizon, so the project fails on cost savings alone. It needs additional value — higher conversion, better quality, lower latency, or strategic control over a capability — or it needs volume growth. At 5 million tasks a month the same project pays back in six months and the decision flips.
Caveat on this model. Straight-line amortisation of $120,000 / 12 implies the fine-tune retains full value for exactly twelve months and none afterwards. It does neither. Treat the allocation as a decision aid, not an accounting entry, and re-run it against the date you expect to migrate to a newer base model.
What are the common mistakes?#
- Treating the token price as total cost. Retrieval, vector storage, tools, moderation, observability, retries, support, and human review routinely exceed the model bill in workflow products.
- Dividing by requests instead of outcomes. Failed and repeated calls consume real resources. Reliability then masquerades as efficiency.
- Calling all model work "training." Prompt engineering, retrieval, and routing are inference-system work with inference economics. Pre-training, fine-tuning, and evaluation have different cash profiles and different depreciation.
- Ignoring output variance. Long-tail outputs drive both cost and latency. Report p50, p90, and p99 by customer segment, not an average.
- Comparing token counts across models. Anthropic notes that Claude 4.7 and later models use a newer tokenizer producing roughly 30% more tokens for the same text, so a per-token price comparison against an earlier model overstates the saving. See Token Economics and Metering.
When does this framework break?#
When quality cannot be scored. Average expected cost is not a sufficient control for medical, financial, security, or safety-critical workflows where one failure dominates thousands of successes. Those need worst-case loss limits and human accountability, not a cost-per-task target.
When the meter is not a unit of value. A token is a provider billing primitive. Two systems with identical token counts can deliver different outcomes because tokenizers, reasoning behaviour, tool protocols, and caching rules differ.
When the curve moves under you. Better models, specialised accelerators, and software optimisation push unit cost down; richer modalities, agentic loops, and higher reliability requirements push it up. Provider rate cards also change on their own schedule. Use scenarios with explicit price and volume ranges, not a single forecast.
Frequently asked questions
01Should we fine-tune or improve the prompt and retrieval first?
Prompt, retrieval, caching, and routing changes almost always win on payback because they cost engineering time rather than a capital-shaped project, and they survive a base-model change. Reach for fine-tuning when you have measured a persistent quality or cost gap that context engineering cannot close, and you can name the recurring per-task saving.
02How should we amortise training cost in a board deck?
Show it as a separate line, not inside gross margin, and state the horizon and the assumed obsolescence date. Then show gross margin both including and excluding it. Investors read a blended number as a claim about steady-state economics, which it usually is not.
03Is inference cost a fixed or variable cost?
It depends on how you buy it. Per-token API pricing is variable. Provisioned throughput and reserved accelerators are fixed until consumed productively. Most products run a mix, which is why utilisation belongs in the cost model — see GPU and Compute Economics.
04Do falling token prices guarantee improving margins?
No. Per-token prices have fallen while contexts, output lengths, reasoning steps, and tool calls have grown. Track dollars per successful task over time; that is the series that determines margin, and it can rise while every rate card falls.
05What single metric should we instrument first?
Fully loaded cost per successful task, segmented by customer and workflow, with attempts-per-success reported alongside it. Everything else in this page is a decomposition of that number.
Related concepts#
- Token Economics and Metering — build the auditable usage ledger that produces the cost figures above.
- GPU and Compute Economics — model reserved capacity, utilisation, and cost per accepted output.
- AI Unit Economics and Gross Margins — connect cost per task to cost of revenue and contribution.
- AI-as-a-Service — turn the compute stack into a commercial offer with service levels.
- Copilots vs. Agents — understand why autonomy multiplies inference and control cost.
- Usage-Based Pricing — align price with a measurable value or cost driver.
- Credits and Drawdown Models — translate heterogeneous AI usage into a customer-facing balance.
- API-as-a-Product — design the contract, reliability, and developer experience around programmatic usage.
- Economies of Scale — test whether serving cost actually falls at higher volume.
- Data Moats — decide whether the capability a fine-tune buys is defensible.
Sources#
- Jordan Hoffmann et al., "Training Compute-Optimal Large Language Models", arXiv:2203.15556, March 2022. The Chinchilla result on allocating training compute between model size and data, and why a smaller compute-optimal model is cheaper to serve.
- Anthropic, Claude Platform pricing documentation, accessed 13 August 2026. Per-model input, cache-write, cache-read, and output rates; the 50% Batch API discount; cache multipliers of 1.25x, 2x, and 0.1x; and the note that Claude 4.7 and later use a tokenizer producing approximately 30% more tokens for the same text.
- OpenAI, API pricing, accessed 13 August 2026. Standard, Batch, Flex, and Fast-mode rates per million tokens, including short- and long-context bands and cache-write pricing on the GPT-5.6 family.
- Google, Gemini Developer API pricing, accessed 13 August 2026. Paid-tier input and output rates by model and service tier, and the higher rate band for prompts above 200k tokens on Gemini 3.1 Pro.
- Amazon Web Services, How tokens are counted in Amazon Bedrock, accessed 13 August 2026. Model-specific output-token burndown rates and the distinction between quota consumption and billable usage.
Price-freshness note: Every rate quoted on this page was checked against the provider's own documentation on 13 August 2026. AI rate cards change on the provider's schedule, not yours — re-verify before using any figure in a board pack, a customer quote, or a pricing decision.
Author
Dr. Sarah Zou
Independent economist · EconNova
Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.
About SarahTopics
Cite this page
Suggested citation
Zou, S. (2026). Inference vs. Training Costs: How AI Founders Model Compute Economics. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/inference-vs-training-costs
Open license
Reuse with attribution
This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.
Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.