AI Economics Wiki
Copilots vs. Agents
A copilot advises a responsible human; an agent pursues a goal through actions inside a defined authority boundary.
Snapshot
What it is
A copilot produces analysis, drafts, or candidate actions while a human keeps the decision and usually triggers the consequential step. An agent receives a goal, plans or selects steps, calls tools, observes results, and executes with less moment-to-moment human direction.
Why it matters
Autonomy is a cost decision before it is a product decision. Agents multiply model calls, tool calls, retries, and control infrastructure, and they move liability from the user to you. The economics only work when the human time you remove is worth more than the inference, oversight, and expected-failure cost you add.
What it is not
A chat window versus a background job. The distinction is the allocation of decision rights, execution rights, and accountability — not the interface, and not the model.
Key takeaways
Choose the minimum autonomy that produces real customer value, then earn the next rung.
Agent economics live or die on the escalation rate, not the token price. In the worked example below, an agent that escalates more than 41% of cases is worse than the copilot it replaced.
Every incremental permission needs a limit, an approval condition, a rollback path, and an audit record.
Autonomy expands the attack surface. NIST measured a jump from 11% to 81% hijacking success once attackers adapted to the specific agent under test.
On this page10 sections
What is the difference between a copilot and an agent?#
OpenAI's builder guidance defines agents as systems that use a model to control workflow execution, call tools to gather context and take action, and operate inside defined guardrails — explicitly excluding single-turn chatbots and classifiers, which do not control the workflow (OpenAI, A practical guide to building agents). Anthropic draws the same line between workflows, where the path through tools is fixed in code, and agents, where the model itself decides the path (Anthropic, Building effective agents).
For product decisions, a binary label is too coarse. Use an authority ladder:
| Level | What the product does | What the human does | Cost and risk profile |
|---|---|---|---|
Assist | Retrieves, summarises, drafts | Directs every task, uses the output | One call, low variance, no execution risk |
Recommend | Ranks options, explains a proposal | Approves or chooses | Adds retrieval and evaluation cost |
Prepare | Assembles a transaction or action plan | Reviews before execution | Multi-call; failure is wasted work |
Act with approval | Executes after a checkpoint | Approves consequential actions | Adds identity, permissions, audit |
Act within bounds | Executes routine actions under policy and budget limits | Handles exceptions, audits samples | Adds rollback, limits, exception queue |
Delegate | Pursues an open objective across tools | Sets goals, monitors, intervenes | Unbounded loops; tail risk dominates |
"Copilot" normally covers the first three rungs. "Agent" normally begins at rung four, where the system executes through tools. The rung is context-specific rather than absolute: scheduling a meeting and moving money can run through the identical software loop while demanding entirely different authority.
Why does this matter to founders?#
The buyer changes. A copilot sells productivity to the person doing the work. An agent sells completed throughput to whoever owns the process and its service level. That moves the pitch from minutes saved toward queue depth, cycle time, exception rate, and who is accountable when it goes wrong — see Positioning.
The cost stack changes shape. Agents plan, call tools, inspect results, recover from errors, and carry state, so a single completed case can consume ten times the tokens of a copilot suggestion. The model call is rarely the expensive part; the reliable system around it — identity, permissions, logging, evaluation, guardrails, rollback, and a staffed exception queue — usually is. See Inference vs. Training Costs for how those layers stack.
Liability moves to you. The more the product can spend, publish, delete, approve, or modify, the more the contract has to say about authorisation, reversibility, auditability, and recourse. NIST's Center for AI Standards and Innovation showed why this is not theoretical: agent hijacking — indirect prompt injection that smuggles instructions into data the agent reads — succeeded far more often once attackers tailored attacks to the specific system (NIST CAISI, 2025).
Procurement expands. A copilot can enter through one team's budget. An agent with write access pulls in security, legal, IT, and the process owner. Budget for a longer cycle and a heavier contract.
How do you decide how much autonomy to ship?#
Run the proposal through three gates in order. Failing any one of them means the answer is a lower rung, not a better prompt.
1. Value gate#
Define the output operationally: which job is completed, how often it occurs, and what human minutes, delay, or leakage disappear.
gross task value = labour saved + delay avoided + incremental contribution − displaced human value
Do not count time saved if the user spends it checking the AI. If success cannot be observed without asking the model whether it succeeded, you do not yet have a measurable outcome.
2. Control gate#
Rank actions by exposure before you rank them by frequency:
risk exposure = action probability x impact x irreversibility x propagation
This is a prioritisation heuristic, not a calibrated loss model — the four terms are not independent and the product has no natural units. Use it to sort which permissions deserve investment, never as an expected-loss figure in a board pack.
For every tool permission, specify seven things: the allowed resource and action; a spend, quantity, or audience limit; the evidence required; the approval condition; the rollback path; an immutable audit record; and a named exception owner. NIST's AI Risk Management Framework is a workable scaffold for mapping, measuring, and managing these across the lifecycle (NIST AI RMF).
3. Economics gate#
For N completed cases:
net value = human time removed + incremental contribution
− AI system cost − human oversight − expected failure cost − change-management cost
expected failure cost = Σ(probability of failure type i x impact of failure type i)
Compare four designs on the same workflow — human baseline, copilot with review, bounded agent with an exception queue, and full delegation — and pick the best risk-adjusted result. The winner is frequently rung three or four, not rung six.
Worked example: does the agent beat the copilot?#
An operations team resolves 20,000 invoice exceptions a month. Fully loaded analyst cost is $32 per hour. Model rates are Claude Sonnet 5 as published on 13 August 2026: $2 per million input tokens, $0.20 per million cache-read tokens, $10 per million output tokens.
Baseline. 12 minutes per case: $32 x 12 / 60 = $6.40 per case, or $128,000 per month.
Copilot. The model drafts a resolution; the analyst decides in 5 minutes.
model = 6,000/1,000,000 x $2 + 800/1,000,000 x $10 = $0.012 + $0.008 = $0.020
tools = $0.008
labour = $32 x 5 / 60 = $2.6667
per case = $2.6947
At 20,000 cases that is $53,893 of variable cost plus $8,000 of fixed platform cost: $61,893 per month, or $3.09 per case.
Bounded agent. The agent resolves 82% of cases end to end and escalates 18%. A run averages 40,000 input tokens (70% served from cache), 5,000 output tokens, 1.15 attempts per completed case, and $0.04 of tool calls.
input = (12,000 x $2 + 28,000 x $0.20) / 1,000,000 = $0.024 + $0.0056 = $0.0296
output = 5,000/1,000,000 x $10 = $0.0500
tools + guardrails = $0.0500
per run = $0.1296
x 1.15 attempts = $0.1490
escalation labour = 0.18 x ($32 x 9 / 60) = $0.8640
expected reversal = 0.82 x 0.2% x $40 = $0.0656
variable per case = $1.0786
| Design | Variable / case | Monthly variable | Monthly fixed | Monthly total | vs. baseline |
|---|---|---|---|---|---|
Human baseline | $6.4000 | $128,000 | — | $128,000 | — |
Copilot + review | $2.6947 | $53,893 | $8,000 | $61,893 | −51.6% |
Bounded agent | $1.0786 | $21,573 | $19,000 | $40,573 | −68.3% |
The agent wins by $21,321 a month — but almost none of that comes from tokens. Model and tool cost is $0.15 of a $1.08 variable case; 80% of the agent's variable cost is the labour it failed to remove. Solve for the escalation rate at which the agent and the copilot cost the same:
20,000 x ($0.1490 + 4.80e + $0.08(1−e)) + $19,000 = $61,893
e = 40.6%
Above roughly 41% escalation the agent is the more expensive product, even though it is the more impressive demo. That single number — not the rate card — is what belongs on the roadmap.
Caveat on this model. The reversal term assumes failures are independent and individually cheap. They are neither. One mis-scoped permission can produce thousands of correlated wrong actions in minutes, and no expected-value term captures that. Treat the table as a cost comparison for the ordinary case and handle the tail with hard limits.
What are the common mistakes?#
- Calling any LLM feature an agent. If the control flow is fixed in code, it is a workflow with a model inside it, and it should be costed and governed as one.
- Equating autonomy with value. A reliable recommendation the user trusts often beats a fragile end-to-end automation the user re-checks.
- Automating an unstable process. If the humans disagree about the policy, the agent cannot infer a durable target and every evaluation is contested.
- Granting broad permissions for convenience. Least privilege, narrow scopes, and explicit spend and audience caps cost one sprint and save one incident.
- Reporting attempts instead of completions. An agent that "handles" a case it later escalates has consumed cost twice — see Token Economics and Metering.
When does this framework break?#
When the loss is catastrophic or correlated. Expected-value arithmetic assumes many small independent outcomes. A credential leak, a mass send, a trade, or a destructive database command breaks that assumption. Those paths need hard technical limits and independent approval, sometimes no autonomous execution at all.
When success is subjective or lagged. In hiring, clinical, legal, or strategic work, a fast proxy metric rewards the wrong behaviour long before anyone notices. Keep the human decision-maker and design the system to surface evidence and uncertainty rather than a verdict.
When the security assumption is stale. NIST's red-teaming showed that a model robust against known attacks can be fragile against attacks written for it, and that repeated attempts materially raise attacker success. Defences validated once, against a public benchmark, are not a control.
Key Facts
Adaptive attacks change the risk picture entirely
Against the same agent, the strongest baseline hijacking attack succeeded 11% of the time; attacks developed specifically for that system succeeded 81% of the time.
NIST CAISI, "Strengthening AI Agent Hijacking Evaluations," January 2025Single-attempt evaluations understate exposure
Allowing 25 attempts raised average attack success across five injection tasks from 57% to 80%, which is the realistic setting whenever an attacker can retry cheaply.
NIST CAISI, January 2025Output is where agent token cost concentrates
On Claude Sonnet 5, output tokens cost $10 per million against $2 for input and $0.20 for a cache read — a 5x output multiple and a 50x spread between cached input and output (rates as of 13 August 2026). Long agent traces are an output problem first.
Anthropic pricingServer-side tools are metered separately from tokens
Anthropic bills web search at $10 per 1,000 searches on top of token cost; OpenAI bills web search at $10 per 1,000 calls and file search at $2.50 per 1,000 calls (13 August 2026). A tool-heavy agent has a bill your token ledger will not show. (Anthropic pricing; )
OpenAI pricingFrequently asked questions
01Should we ship a copilot first and add autonomy later?
Almost always. The copilot generates the labelled decisions, the failure taxonomy, and the escalation baseline that an agent's business case depends on. Shipping the agent first means guessing the escalation rate, which is the variable the economics turn on.
02How should we price an agent differently from a copilot?
A copilot maps reasonably onto seats because value scales with the number of people assisted. An agent removes the seat, so seat pricing shrinks your revenue exactly as you succeed. Price the completed unit of work or the throughput commitment instead — see Usage-Based Pricing and Outcome-Based Pricing.
03What is the single metric to instrument first?
Fully loaded cost per completed case, reported next to the escalation rate and attempts per completion. Those three numbers decide whether the next rung of autonomy is an investment or a subsidy.
04Does a human-in-the-loop checkpoint solve the security problem?
It caps the blast radius, it does not close the hole. Reviewers approve plausible-looking actions at high volume. Pair approval with narrow scopes, spend caps, and rollback so that an approved wrong action stays small and reversible.
05Is an agent a moat?
Not by itself. The defensible parts are the permission integrations, the evaluation set, the exception-handling process, and the accumulated record of what "correct" looks like in that customer's workflow — see Data Moats and Switching Costs.
Related concepts#
- Inference vs. Training Costs — model the extra inference, tool, and control cost that autonomy creates.
- AI Unit Economics and Gross Margins — carry cost per completed case into cost of revenue.
- Token Economics and Metering — instrument attempts, completions, and tool spend.
- AI-as-a-Service — package autonomy with service levels and exception handling.
- Managed Services — compare software autonomy with a human-backed delivery layer.
- Usage-Based Pricing — meter actions or completed work when seats stop tracking value.
- Product-Market Fit — prove the workflow and buyer before raising autonomy.
- Switching Costs — understand the dependence that write access creates.
Sources#
- OpenAI, A practical guide to building agents, accessed 13 August 2026. Definition of an agent as a system that uses a model to control workflow execution within guardrails, and the exclusion of single-turn chatbots and classifiers.
- Anthropic, Building effective agents, accessed 13 August 2026. The workflow-versus-agent distinction, and the argument for the simplest architecture that solves the task.
- Center for AI Standards and Innovation (NIST), Technical Blog: Strengthening AI Agent Hijacking Evaluations, 17 January 2025, updated 19 December 2025. The 11%-to-81% adaptive-attack result, the 57%-to-80% multi-attempt result, and the case for task-level rather than aggregate risk reporting.
- NIST, AI Risk Management Framework, accessed 13 August 2026. Lifecycle structure for governing, mapping, measuring, and managing AI risk.
- Anthropic, Claude Platform pricing, accessed 13 August 2026. Claude Sonnet 5 input, cache-read, and output rates used in the worked example, and web-search tool pricing.
- OpenAI, API pricing, accessed 13 August 2026. Web search and file search tool rates quoted in Key Facts.
Price-freshness note: Every rate on this page was checked against the provider's own documentation on 13 August 2026. Provider rate cards change on the provider's schedule, not yours — re-verify before using any figure in a board pack, a customer quote, or a pricing decision.
Author
Dr. Sarah Zou
Independent economist · EconNova
Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.
About SarahTopics
Cite this page
Suggested citation
Zou, S. (2026). Copilots vs. Agents: Choosing the Right AI Product Boundary. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/copilots-vs-agents
Open license
Reuse with attribution
This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.
Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.