AI Economics Wiki
Token Economics and Metering
Token economics connects provider usage, successful customer work, price, and margin through an auditable metering system.
Snapshot
What it is
Token economics translates model consumption into product cost, customer usage, price, and contribution margin. Metering is the measurement and attribution layer that records which customer, feature, task, and model produced that consumption.
Why it matters
Without event-level attribution, a single workflow can consume the margin of an entire plan and you will find out at the quarter close. With it, every optimisation can be ranked in dollars per successful task instead of tokens saved.
What it is not
A line on the provider invoice. An invoice tells you what you spent; it cannot tell you which customer, which feature, or which failed retry spent it.
Key takeaways
The minimum ledger separates uncached input, cache reads, cache writes, output, tools, retries, and successful tasks — and stores the rate that was in effect.
Cache writes are a real cost. Modelling caching as a pure discount overstates the saving.
Quota consumption is not billable usage. Bedrock can deduct 1,500 tokens from your quota for a request it bills at 1,100.
Do not expose raw tokens as the customer price merely because the provider bills that way.
On this page10 sections
What is a token, and why is it a bad customer meter?#
A token is a model-specific unit produced by a tokenizer. It may be part of a word, a whole word, punctuation, or a fragment of code. Token counts are not comparable across models, modalities, or providers — Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text than earlier models (Anthropic pricing). A headline per-token price cut of 25% against an older model is, on that arithmetic, no cut at all.
For metered inference the economic chain runs:
customer action → product task → one or more model and tool calls
→ provider usage → invoice cost → customer charge → contribution
Every arrow needs an identifier and a reconciliation rule. A provider usage field alone does not say which customer received value. A product event alone does not prove what was billed upstream.
Four quantities are routinely confused, and they are not interchangeable:
| Quantity | What it measures | Who uses it |
|---|---|---|
Usage quantity | Technical consumption at the call | Engineering, optimisation |
Pricing quantity | The amount a rate applies to | Finance, provider reconciliation |
Capacity quantity | Throughput or reservation burndown | Capacity planning, rate limits |
Customer billing quantity | What your contract charges for | Revenue, invoices, disputes |
Amazon Bedrock makes the gap explicit: output tokens are converted into quota usage through a model-specific burndown rate, so a request with 1,000 input and 100 output tokens can deplete 1,500 tokens of quota while being billed for 1,100 (AWS, How tokens are counted in Amazon Bedrock). A team that treats quota telemetry as cost telemetry will over-report spend by a third and mis-price the product accordingly.
Why does metering matter to founders?#
It protects gross margin. Falling token prices do not guarantee improving margins if contexts, outputs, retries, tool calls, or quality targets expand at the same time. Only per-customer, per-feature attribution shows which is happening — see Gross Margin.
It makes pricing a decision rather than a guess. A reliable cost ledger lets you offer allowances, credits, per-task pricing, or commitments with a known contribution floor. It also exposes where customer value and provider cost diverge, which is where better packaging lives.
It makes optimisation rankable. Prompt compression, caching, smaller-model routing, batching, output limits, and retry policy can each be scored in dollars per successful task. An optimisation that cuts tokens but lowers task success raises total cost — and only the ledger will say so.
It survives due diligence. Enterprise buyers want explainable bills and hard limits; finance needs to reconcile usage against invoices, discounts, and commitments; investors want evidence that AI margin is observed rather than extrapolated from a demo.
How do you build the usage-to-margin ledger?#
1. Define the canonical event#
Each model or tool call should record, at minimum:
- immutable event and task IDs, plus timestamp and billing period;
- customer, workspace, end user, and feature;
- provider, model, region, and endpoint;
- uncached input, cache-read, cache-write, output, and any other provider units;
- request status, retry number, and latency;
- tool and retrieval charges;
- the rate-card or contract version in effect;
- estimated cost at event time; and
- final task outcome, plus the customer-facing meter and billable quantity.
Do not log prompt content by default in order to obtain cost data. Token counts, request metadata, and task outcome can be recorded separately from the text.
2. Calculate provider cost per call#
model cost_i = uncached input_i x input rate
+ cache read_i x cache-read rate
+ cache write_i x cache-write rate
+ output_i x output rate
+ other provider units_i
task variable cost = Σ(model calls + tools + retrieval + guardrails + variable review)
The cache-write term is the one teams drop. Anthropic prices a five-minute cache write at 1.25x base input and a one-hour write at 2x, against a cache read at 0.1x; OpenAI prices cache writes at 1.25x input with cached reads at 0.1x (both as of 13 August 2026). Caching is a net win quickly — after one read on a five-minute cache, two on a one-hour cache — but a model that ignores writes will overstate the benefit on short-lived or low-reuse content.
3. Normalise to successful customer work#
cost per successful task = total task-attributed variable cost / successful tasks
Report alongside it: cost per attempt, attempts per success, input and output percentiles (p50/p90/p99), cache-hit rate, routing mix, escalation rate, and contribution by customer, feature, and cohort.
4. Translate into a customer meter#
A good customer meter is predictable, controllable, auditable, and correlated with value. Options run from raw provider units (for developer buyers) through successful tasks or documents, workflow actions, a subscription allowance plus overage, credits, or reserved capacity.
credits consumed = Σ(provider unit x internal conversion factor)
Publish the conversion logic, give a pre-run estimate where you can, and version changes visibly. Opaque credits are an untraceable price increase, and customers eventually price that risk in — see Credits and Drawdown Models.
5. Reconcile and control#
reconciliation variance = provider invoice cost − metered estimated provider cost
Investigate variance from discounts, commitments, minimums, taxes, credits, late events, untagged traffic, model aliases, rounding, and missing logs. Set budgets and alerts at customer, workspace, feature, and model level. Hard stops suit abuse and explicit prepaid limits; graceful degradation, an approval step, or a cheaper route usually suits customer-critical workflows. The FinOps Foundation's FOCUS specification gives a provider-neutral schema for this reconciliation, including non-monetary units such as credits and tokens (FOCUS specification).
Worked example: where the margin actually goes#
An AI product handles 100,000 customer tasks a month. Each averages 2,400 input and 600 output tokens. Forty percent of input is served from cache, each cached block is read about 20 times before it expires, and retries add 8% to all model usage. Rates are Claude Haiku 4.5 as published on 13 August 2026: $1.00 input, $0.10 cache read, $1.25 five-minute cache write, $5.00 output per million tokens.
Step 1 — monthly usage after retries
total input = 100,000 x 2,400 x 1.08 = 259.20M tokens
uncached = 259.20M x 0.60 = 155.52M
cache reads = 259.20M x 0.40 = 103.68M
cache writes = 103.68M / 20 = 5.184M
output = 100,000 x 600 x 1.08 = 64.80M
Step 2 — model cost
| Usage type | Tokens | Rate / MTok | Cost |
|---|---|---|---|
Uncached input | 155.52M | $1.00 | $155.52 |
Cache reads | 103.68M | $0.10 | $10.37 |
Cache writes | 5.18M | $1.25 | $6.48 |
Output | 64.80M | $5.00 | $324.00 |
Model total | $496.37 |
Step 3 — full variable cost
tools = 100,000 x 1.08 x $0.002 = $216.00
guardrails and observability = $180.00
total variable = $496.37 + $216 + $180 = $892.37
cost per task = $892.37 / 100,000 = $0.008924
At a price of $0.04 per task, revenue is $4,000 and contribution is $4,000 − $892.37 = $3,107.63, a 77.7% contribution margin.
Step 4 — one product change, eight points of margin
Output length doubles to 1,200 tokens. Nothing else changes:
output cost = 129.60M x $5.00 = $648.00
model total = $820.37
total variable = $1,216.37 ($0.012164 per task)
contribution margin = ($4,000 − $1,216.37) / $4,000 = 69.6%
An 8.1-point margin loss from a single behavioural change, with no price move and no provider price change. Note also that output is 65% of the model bill in the base case and 79% after the change — output length, not prompt length, is the first lever to reach for.
Step 5 — what caching is actually worth
Without caching, all 259.20M input tokens bill at $1.00: model cost would be $259.20 + $324.00 = $583.20. The cached configuration costs $496.37, so caching saves $86.83, or 14.9% of the model bill — meaningful, but well short of the "40% of input at a tenth of the price" intuition, because the 60% uncached remainder and the write cost both survive.
Credit translation. If one credit represents $0.025 of contracted value, a four-cent task consumes $0.04 / $0.025 = 1.6 credits. Set that conversion from your commercial price and target contribution, never from provider cost alone — otherwise every provider price change silently repriced your product.
Caveat on this model. The cache-write term assumes a 20:1 read-to-write ratio. That ratio is a property of your traffic shape, not of the cache, and it collapses for low-frequency tenants whose cached content expires before it is reused. Instrument reads and writes separately rather than assuming a ratio.
Key Facts
Quota is not the bill
Amazon Bedrock applies a model-specific output burndown rate to quota: a request with 1,000 input and 100 output tokens can deplete 1,500 tokens of quota while being billed for 1,100.
AWS, *How tokens are counted in Amazon Bedrock*Token counts are not comparable across model generations
Claude 4.7 and later use a newer tokenizer that produces approximately 30% more tokens for the same text than Claude Sonnet 4.6 and earlier (13 August 2026).
Anthropic pricingA single call can have four different prices
On Claude Haiku 4.5 the four rates are $1.00 input, $0.10 cache read, $1.25 or $2.00 cache write, and $5.00 output per million tokens — a 50x spread between the cheapest and most expensive unit in one request (13 August 2026).
Anthropic pricingLong context is a separate rate band, not a surcharge on the overflow
On gpt-5.6-sol, crossing into long context moves the whole request to $10.00 input and $45.00 output per million tokens from $5.00 and $30.00 (13 August 2026). Retrieval that over-stuffs context can double the bill for the entire call.
OpenAI pricingThe reconciliation schema is standardised
The FinOps Foundation ratified FOCUS v1.4 on 4 June 2026; the specification defines fixed column names and meanings for billing data across providers, and since v1.2 covers contracted pricing in non-monetary units such as credits and tokens.
FOCUS specificationWhat are the common mistakes?#
- Using one blended token rate. Input, cache reads, cache writes, output, tools, and reserved capacity have different prices and different elasticities.
- Dropping failed calls from the ledger. They consume real cost and they belong in the denominator's failure count, not in a rounding error.
- Confusing quota burndown with billable usage. See the Bedrock example above; the gap is model-specific and can exceed 30%.
- Passing tokens through as the customer meter. Non-technical buyers cannot forecast tokens, and per-token pricing quietly penalises languages that tokenize less efficiently.
- Setting credit conversions from cost. Credits should preserve your commercial price and expected contribution across features, not track a supplier's rate card.
When does token metering break?#
Multimodal and tool-heavy products. Images, audio, video, search, code execution, and human review have native units and native prices. Forcing them into notional tokens produces a tidy dashboard and a wrong number. Anthropic bills web search at $10 per 1,000 searches and code-execution containers by the hour; those never appear in a token count.
Shared and asynchronous work. Cache entries, batched requests, background evaluations, and platform-wide retrieval serve several customers at once. Define an allocation rule, label the result as an allocation rather than a measurement, and keep the unallocated remainder visible.
Meters as incentives. Per-token pricing encourages customers to withhold useful context. Per-outcome pricing invites disputes about causality. Review the meter as a behaviour-shaping mechanism, not only an accounting device — see Pricing Metric / Value Metric.
Frequently asked questions
01Should we bill customers in tokens?
Only if the buyer is a developer who controls prompt and output length directly. Everyone else needs a meter they can predict and influence: tasks, documents, cases, or an allowance with a published overage rate.
02How exact does event-level cost estimation need to be?
Close enough that period-close variance against the provider invoice is explainable, and consistent enough that trends are trustworthy. Chasing cent-level accuracy on estimates is less valuable than storing the rate version so old periods can be recomputed correctly.
03How do we handle a mid-period rate change?
Store the contract version on the event and compute weighted-average cost by model and rate. Restating a whole period at the new rate is the most common source of an unexplainable margin bridge.
04Should we hard-stop customers who blow through an allowance?
Hard stops for abuse and explicit prepaid limits; degradation, approval, or a cheaper route for workflows the customer depends on. An unexpected hard stop on a production workflow costs more in churn than the overage ever saved.
05What is the one report to build first?
Cost per successful task by customer and feature, with attempts per success next to it. Everything else on this page is a decomposition of that report — see AI Unit Economics and Gross Margins.
Related concepts#
- Inference vs. Training Costs — separate usage cost from model-development investment.
- AI Unit Economics and Gross Margins — carry the ledger into cost of revenue and contribution.
- GPU and Compute Economics — attribute reserved capacity that no token count will explain.
- AI-as-a-Service — connect the ledger to offer design and service levels.
- Copilots vs. Agents — meter attempts and completions when the system loops.
- Usage-Based Pricing — choose and govern a consumption metric.
- Credits and Drawdown Models — package heterogeneous usage into a prepaid balance.
- Pricing Metric / Value Metric — keep the customer meter aligned with value.
- API as a Product — expose trustworthy usage, limits, and billing to developers.
- Gross Margin — the statement line this ledger ultimately feeds.
Sources#
- Amazon Web Services, How tokens are counted in Amazon Bedrock, accessed 13 August 2026. Model-specific output burndown rates, the quota calculation
InputTokenCount + CacheWriteInputTokens + (OutputTokenCount x burndown rate), and the statement that billing follows actual token usage rather than quota consumption. - Anthropic, Claude Platform pricing, accessed 13 August 2026. Per-model input, cache-write, cache-read and output rates; the 1.25x, 2x and 0.1x cache multipliers; web-search and code-execution tool pricing; and the note that Claude 4.7 and later use a tokenizer producing roughly 30% more tokens for the same text.
- OpenAI, API pricing, accessed 13 August 2026. Separate input, cached-input, cache-write and output rates, and the short- versus long-context rate bands on the GPT-5.6 family.
- Amazon Web Services, CountTokens API reference, accessed 13 August 2026. Pre-flight token estimation, used to build a cost estimate before a request is sent.
- FinOps Foundation, FOCUS specification, v1.4 ratified 4 June 2026. A provider-neutral billing schema with fixed column names and meanings, including support for contracted pricing in non-monetary units such as credits and tokens.
Price-freshness note: Every rate on this page was checked against the provider's own documentation on 13 August 2026. Provider rate cards change on the provider's schedule, not yours — re-verify before using any figure in a board pack, a customer quote, or a pricing decision.
Author
Dr. Sarah Zou
Independent economist · EconNova
Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.
About SarahTopics
Cite this page
Suggested citation
Zou, S. (2026). Token Economics and Metering: How AI Products Turn Usage Into Revenue. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/token-economics
Open license
Reuse with attribution
This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.
Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.