Essay
Token Price Is Not Task Cost
A model's posted rate is only one input into the cost of useful AI work. The decision unit is the accepted task, after retries, tools, verification, review, and delay.
A token price is a tariff on one production input. It is not the unit cost of the work a business receives.
That distinction matters because an AI system rarely consists of one prompt, one response, and an automatic acceptance. It may retrieve documents, call tools, retry after an error, escalate to another model, run a validator, wait in a queue, and ask a person to review the result. A cheaper model can therefore produce a more expensive task if it needs more attempts or more supervision. A higher-priced model can be cheaper if it resolves the task once.
The finance question is not “What is the cheapest token?” It is: What does one accepted, useful task cost under the operating conditions we will actually run?
The price card is an input tariff
Reported facts, checked 30 August 2026. The major API price cards already signal that there is no single token price. OpenAI’s pricing table separates ordinary input, cached input, cache writes, output, and short- versus long-context rates. Anthropic’s pricing documentation adds cache operations, batch processing, long-context conditions, and separately metered tools. Google’s Gemini pricing varies by model and service tier and separately lists context caching, storage, and grounding charges.
Even the token denominator is not perfectly portable. OpenAI’s official token guide notes that models can tokenize the same text differently and can generate different quantities of output or reasoning tokens. A comparison based only on visible words or a headline input rate can therefore be misleading before workflow effects enter the calculation.
Inference. Price cards are necessary for estimating cost, but they are closer to electricity tariffs than to the cost of a finished product. They price measured consumption. They do not tell the buyer how much consumption, rework, assurance, or waiting is required to produce an acceptable outcome.
Choose the denominator before choosing the model
“Cost per task” is ambiguous until the task states are defined:
- A submitted task entered the workflow.
- An attempted task generated at least one model call.
- A completed task returned something rather than timing out or failing.
- An accepted task passed the workflow’s quality and control threshold.
- A useful task created the intended downstream benefit.
Those are not interchangeable. A workflow can show a low cost per completion while pushing defective outputs into manual rework. It can show a high technical success rate while users ignore the result. It can also be reliable but economically irrelevant because the benefit is smaller than the cost to serve.
A practical management equation is:
Cost per accepted task = (model usage + tools and runtime + retries and fallbacks + automated verification + human review + attributable operations) / accepted tasks
Retries should appear once in the accounting—normally as the additional model, tool, and verification consumption they cause—but should also be tagged as a causal driver. Latency is best kept as a separate operating measure until the business can defend a cash or opportunity-cost conversion.
The cost stack behind an accepted task
1. Inference consumption
Count input, cached input, cache writes, output, and any billed reasoning or modality-specific tokens at the applicable model and service-tier rates. Use actual API usage records, not prompt-length estimates, when available. Long contexts, premium processing, batch discounts, and regional processing can change the applicable rate.
2. Retries, fallbacks, and routing
An initial call may fail technically, fail a validator, or produce a low-confidence answer. The workflow might retry with a revised prompt, add retrieved evidence, or route to a stronger model. Every branch consumes more tokens and time. The useful comparison is the total cost of the resolution path, including unsuccessful paths, allocated across accepted outputs.
3. Tools and runtime
Search, retrieval, file handling, code execution, browser actions, and hosted containers can carry per-call, per-token, storage, or runtime charges. They may also create their own failure and latency profiles. A model-only invoice view misses these costs even when they are essential to task quality.
4. Verification and control
Verification can be deterministic—a schema check, reconciliation, policy rule, or calculation—or probabilistic, such as a second-model critique. Higher-consequence work may also require human approval. NIST’s Generative AI Profile recommends defining human-AI oversight responsibilities and making the robustness of evaluations proportionate to identified risks.
Inference. Assurance is not overhead that can be assumed away when comparing models. It is part of the production process. The right level depends on the cost of an error, reversibility, regulation, and who bears the loss.
5. Human review and exception handling
Human minutes can dominate a workflow whose raw inference costs are measured in cents. Record review time, escalation time, and downstream correction time separately. A task that needs three minutes of expert review is not “fully automated,” even if the model call itself costs almost nothing.
6. Delay, capacity, and fixed operations
Latency can affect user abandonment, employee throughput, service-level commitments, or the number of concurrent jobs the system can support. Observability, evaluation suites, security controls, and workflow maintenance also cost money. Their financial-statement classification will depend on the business model and accounting policy; their economic burden does not disappear.
Worked hypothetical: cents of inference, dollars of review
Illustrative estimate—not a market benchmark. Consider a document-analysis workflow with 10,000 submitted tasks. Assume:
- each first attempt uses 12,000 input tokens at $1 per million and 2,000 output tokens at $6 per million;
- 18% of tasks receive one retry with the same token profile;
- 25% use one external tool costing $0.01;
- every attempt receives a $0.003 automated check;
- 8% of submitted tasks receive three minutes of human review at a hypothetical loaded rate of $45 per hour; and
- 9,300 tasks ultimately pass the acceptance threshold.
| Cost item | Illustrative calculation | Cost |
|---|---|---|
| First-attempt model usage | 10,000 × $0.024 | $240.00 |
| Retry model usage | 1,800 × $0.024 | $43.20 |
| Tool calls | 2,500 × $0.01 | $25.00 |
| Automated checks | 11,800 × $0.003 | $35.40 |
| Human review | 800 × 3/60 hours × $45 | $1,800.00 |
| Total | $2,143.60 |
The first-attempt model cost is $0.024 per submitted task. The modeled cost under these assumptions is about $0.214 per submitted task and $0.230 per accepted task. On these assumptions, cost per accepted task is roughly 9.6 times the first-pass inference figure.
That multiple is not a general claim. Change the review rate, retry rate, task mix, or acceptance threshold and it changes immediately. The point is structural: the apparently small input price does not control the answer when another line item dominates.
Open question. Would a stronger and more expensive model reduce retries or human review enough to lower the $0.230? The price card cannot answer. Only a controlled comparison on representative tasks, using the same acceptance standard, can.
A measurement design for finance and engineering
The minimum useful cost ledger is task-level. For each submitted task, record:
- workflow and task category;
- model calls, token categories, service tiers, and price version;
- tools, runtime, retrieval, and storage charges;
- retry and fallback reasons;
- automated verification cost and result;
- human review minutes and disposition;
- end-to-end latency; and
- completion, acceptance, and, where observable, downstream usefulness.
Aggregate results by task class rather than hiding variation in one average. Report the cost per submitted, completed, and accepted task together. Show tail latency and exception rates. Keep the acceptance rubric fixed when comparing models or architectures.
For a build-versus-buy or model-selection decision, run the same task set through each candidate route. A financially honest scorecard includes acceptance rate, direct cost per accepted task, human minutes per accepted task, and the distribution of failures. If one route appears cheaper only because it rejects more hard cases or applies a weaker control threshold, it is not a like-for-like result.
Cost is not value
Cost per accepted task is the correct production denominator, but it is not the investment case. The other side is the value of the accepted work: labor genuinely avoided, capacity released, cycle time improved, loss prevented, or revenue enabled. These benefits should be measured against a credible baseline and adjusted for adoption and displacement, not asserted from model capability.
Open questions for a deployment review:
- What exactly qualifies as accepted, and who owns that definition?
- Which errors can be reversed cheaply, and which require mandatory review?
- How often do retries, fallbacks, or people rescue the first answer?
- Which costs sit outside the model invoice?
- Does the output change a decision or merely create more material to inspect?
Token prices matter. But the economically relevant unit begins where the price card ends: the complete path from request to accepted work.
Source register
Corrections
No corrections recorded.
Report a correction via LinkedIn with the note title and the supporting source.