Sources#
- Agent swarms and the new model economics
- Build more natural voice experiences with GPT‑Live‑1 in the API
- Building Prod with Jev and LangGraph
- Claude models explained: choosing the best model for your use case
- Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI
- Fable's judgement
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Gemini 3.5 Flash-Lite Model Card
- HarnessTax: How Much Does the Harness Matter for Coding Agents?
- How is the Bun Rewrite in Rust Going?
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- Introducing System One Models & Jev
- Recursive Self Improvement for Coding Agents
- Running a Software Factory Efficiently at Uber Scale
- September 2026 Ramp AI Index: Cracks in the AI thesis, part 2
- Startup ARR is less secure than ever, new research shows
- The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- The plunging price of thought
- The price is wrong: AI cost calculation has to consider task completion rates, not just token costs
- The state of AI in 2026: On the road to ROI
Summary#
Anthropic's published answer to "which model should I use for this workload?" (vendor-claim, July 2026). The default recommendation inverts the intuitive one: start with the most intelligent generally-available model and use the effort level to dial performance and cost down — not start cheap and escalate. The load-bearing claim is that the two cost metrics diverge: price-per-token is higher for stronger models, but cost-per-task is often lower, "because more capable models often take fewer turns and less thinking time to get most tasks right."
A second, non-economic argument does independent work: starting with a smaller model makes it harder to distinguish model failures from setup failures. A weak model failing your harness tells you nothing about whether the harness is wrong. Start at a capability level where failure is informative, then descend.
Anthropic documents both directions — start-smart-and-descend, and start-cheap-and-ascend-until-quality-clears — but leads with the former.
Why the two cost metrics diverge#
A task's cost is (tokens per turn) × (turns) × (price per token). Model class moves all three terms, and not in the same direction:
- Turns. A stronger model gets more tasks right on the first attempt; retries, re-planning, and repair loops are where a cheap model's savings go.
- Thinking tokens. At a fixed effort setting a stronger model reaches the answer with less deliberation.
- Price. Higher for stronger models — the only term that favors the cheap choice.
The same distinction decides how the trend in AI prices is measured. Earlier price-decline estimates (a16z's 10×/yr, Epoch's 2025 9–900×/yr) tracked price per token for models that clear a benchmark threshold. Epoch's 2026 report (empirical) switched to realized cost per question because reasoning models get their scores by spending more tokens on cheaper-per-token models. On that basis the cost of a fixed score falls ~47% per quarter. The slowest benchmark is SWE-bench Verified at 27.5%, the one closest to this page's agentic workloads. See The Price of Fixed Capability.
The claim is that the first two usually dominate. Note the failure mode this argument does not cover: a stronger model that over-deliberates rather than converging inverts the same arithmetic — see Unproductive Self-Verification, where Opus 5 performed worse at higher effort by re-verifying answers it had already verified. The cost-per-task argument holds only where extra capability buys convergence rather than rumination.
Effort as the second axis#
Model class and effort level are separate dials, and they overlap: "higher-class models at lower efforts can sometimes be more efficient than smaller models." So the selection space is a 2-D grid, not a ladder — Opus-at-low-effort and Sonnet-at-high-effort are different points that may cost the same. The two published curves (quality-vs-cost, quality-vs-latency) are labelled "illustrative and not plotted from benchmark data" — the shape of the tradeoff is asserted, not measured, which is unusual candor for a vendor selection guide and also means nothing here is a number you can plan against.
This is the same inference-budget-is-capability thesis as Large-Scale Test-Time Compute, surfaced as a product knob rather than a research finding.
The model classes#
Anthropic's framing: the classes do not specialize by domain ("we don't recommend one model class for finance and another for science"). Every class is trained for coding, agentic tasks, and knowledge work. The only axis is how hard a problem the class can reliably carry, and what that costs.
| Class | Position | Notes |
|---|---|---|
| Mythos / Fable | Most capable; frontier across domains; coding, long-running agents, previously-unsolved problems | Two packages of the same underlying model: Mythos for Project Glasswing organizations doing dual-use cyber/bio work, Fable with the extra safeguards that make it public-safe. Both require limited data retention |
| Opus | Reasoning-intensive enterprise tasks | Benchmarks cited: GDPval-AA (knowledge work), Terminal-Bench 2.1 (agentic coding) |
| Sonnet | Everyday tasks; balance of performance/cost/speed | Called out for high-volume sub-agents in multi-agent orchestration |
| Haiku | Lowest cost, fastest | High-frequency workloads where latency and cost dominate |
Opus vs. Fable: the benchmark-blind difference#
The most interesting claim in the piece. Opus and Fable have "similar benchmark scores," yet: "in real-world situations, larger models such as Fable tend to have more wisdom, creativity, and writing skills." The stated decision rule is therefore not benchmark-driven at all —
If your evals or internal testing show Opus struggling on some tasks, then Fable is the answer. If Opus already clears the quality bar, then its speed and price profile may make it the better choice.
A vendor stating that its own benchmark scores fail to separate two adjacent classes is a first-party construct-validity admission — the same gap Measuring Beyond Accuracy Saturation documents from the research side, where statistically-indistinguishable agents still differ sharply on reliability, efficiency, and cost.
The four selection questions#
- How hard is this task? Multi-step, long-running, or previously-unsolved → higher class.
- What are the latency needs? High-frequency customer-facing → Sonnet.
- What are the access constraints? Mythos is Glasswing-only; not every org exposes every class to every role.
- What are the unit economics? At high production volume, lower classes may be right if evals show the tasks complete satisfactorily.
Note that question 4 defers to evals, and question 1 defers to judgment. The framework is a scaffold for a measurement, not a substitute for one.
The advisor strategy#
The hybrid that avoids choosing: a faster, cheaper worker model calls a more intelligent advisor model to check its plan and evaluate its work — "the executor model is coached only when needed."
The one quantified result: on SWE-bench Pro, Sonnet 5 with a Fable 5 advisor lands within 10% of Fable 5's own score at 63% of the price of running Fable 5 for the whole task.
Two things make this more than a cost trick:
- It is Optimizer–Evaluator Decoupling shipped as a first-party recommendation — the thing that grades the work is structurally not the thing that produced it. Anthropic arrives at the same invariant from the cost side that the agent-quality literature arrives at from the Goodhart side.
- It is a partial answer to the standing question of where a cheap model crosses over into an expensive one (Claude Sonnet 5): the crossover isn't a point on the effort dial, it's a different topology. Selective coaching beats both "run the cheap model harder" and "run the expensive model throughout."
Caveat: 63% of the price is measured against running Fable 5 for everything, not against Sonnet 5 alone — the advisor is a markup on the cheap path, and the 10% quality gap is a real one.
The distinction, seen from the vendor side#
DeepMind's Gemini 3.5 Flash-Lite card (2026-07-21, vendor-claim) is what this page's argument looks like when a vendor publishes only half of it. Its benchmark table opens with price rows — input and output $/1M for itself, its predecessor, GPT-5.4 mini and Claude Haiku 4.5 — and its own generation-over-generation move is Anthropic's shape exactly: price-per-token up 67% on output, capability up a full tier (SWE-Bench Pro 38.3 → 54.2, OSWorld-Verified 54.3 → 74.0). If the cost-per-task argument holds, the more expensive model is the cheaper one to run, and this is the first cross-vendor instance of the pattern in the corpus.
It is also the cleanest demonstration of why the argument stays unverified. The card reports no tokens-per-task, no turn counts, and no thinking budget for any of the four models — so the one term that would decide it (does the stronger model converge in fewer turns and fewer thinking tokens?) is absent from the only artifact that prints the price. Both vendors publish the price axis and neither publishes the token axis; Anthropic at least labels its curves "illustrative", while DeepMind's numbers look precise and are missing the same multiplier. See Inference Efficiency as Capability for the per-dollar arithmetic and Compute-Controlled Benchmarking for why a price row is not a compute budget.
Four mixes, one quality bar, 8× the cost (Cursor, 2026)#
Every source above prints a price and omits the task. Cursor's swarm post (Wilson Lin, 2026-07-20, case-study) is the first in the corpus to do the reverse: hold the task, the harness and the time budget fixed, vary only which model plays which role, and publish the dollars.
The task is a from-scratch SQLite implementation in Rust graded on a held-out suite (Parallel Agent Orchestration). Four configurations under a four-hour budget: GPT-5.5 as both planner and worker; Grok 4.5 as both; Opus 4.8 planning with Composer 2.5 working; Fable 5 planning with Composer 2.5 working. Only the endpoints of the cost range are given in prose — $1,339 for the Opus 4.8 hybrid to $10,565 for GPT-5.5 alone (the per-run bars live in a chart this ingest did not capture).
Cursor's summary: "every model mix produced similar quality while the costs varied enormously." Quality at the cutoff sat between 73% and 85% across all four, and every configuration eventually passed the whole suite. So this is a rare thing — an 8× cost spread at a common quality bar, on real infrastructure, outside Anthropic.
Three findings underneath it:
- Tokens and dollars split differently, and the split is the whole argument. Workers carry at least 69% of the tokens in every run and over 90% in most. But in the Opus/Composer mix the planner produced a small fraction of the tokens and roughly two-thirds of the cost, with the worker fleet handling the vast majority of tokens for the remaining third. Frontier price is affordable precisely because frontier judgment is needed rarely: "few moments in a large task genuinely require frontier intelligence… once a frontier planner has collapsed the ambiguity into a detailed, explicit instruction, less expensive models simply have to follow it."
- The worker line is where the money is. GPT-5.5 doing both jobs spent $9,373 on workers alone; Opus-planning-with-Composer-working spent $411 on the entire worker fleet — a 23× gap on the same task at the same grade. This is the strongest evidence in the corpus for the guidance's own one-line hedge ("Sonnet for high-volume sub-agents in multi-agent orchestration"), which the guidance never develops.
- Cost-per-task held inside the role and reversed at the system level. The Fable 5 planner billed slightly less than the Opus 4.8 planner despite roughly twice the per-token price, because it emitted far fewer planning tokens — this page's thesis, confirmed, at the planner role. And the Fable run came out substantially more expensive overall, because its workers burned several times as many tokens. A better planner externalized cost downstream. Per-role cost-per-task can be right about every role and wrong about the bill.
That last point is the one to carry. The vendor guidance reasons about "a workload" as if it had one model in it; the moment a pipeline has two roles, a model's cost includes the work it causes other models to do, and no single-role measurement can see that term. It is the cost-side twin of the cross-stage coupling Client-Side Agent Optimization's combo abstraction describes for quality.
Discount appropriately. Vendor-authored, on the vendor's own harness, with the vendor's own model (Composer 2.5) as the worker in both cheap configurations — and the comparison across mixes is confounded with a simultaneous harness rebuild insofar as absolute figures go (the old harness's runs are not costed here). The solo Opus 4.8 and Fable 5 runs that would round out the frontier-cost picture were "graded only informally," so Cursor explicitly draws no quality conclusion from them.
The layer above the model menu (Writer, 2026)#
Every source above varies the model. Writer's harness-swap paper (arXiv 2607.06906, 2026-07-08, empirical) holds the model constant and varies the orchestration layer instead — 22 locked tasks, six models across five vendors, one LLM-judge panel, one pinned price table, and a single variable: a conventional production agent loop (frozen 2026-06-07) versus Writer's Agent Harness. Blended: cost per task −41% ($0.21 → $0.12), tokens −38% (14.2k → 8.8k), wall-clock −44%, quality at parity (0.78 → 0.81, reported as a wash at n = 22).
The result that bears on this page is the comparison of levers: under the baseline, moving from the most expensive model (Palmyra X6, $0.25/task) to the cheapest (Qwen 3.6, $0.16) saves 36%; keeping any model and swapping the harness saves 33–61%, with no exceptions across six models. On this workload the orchestration layer moved the bill more than the entire spread of the model menu did. Writer's framing: "teams comparing $/Mtok across vendors are comparing p; the bill is p × τ, and τ belongs to the harness."
It also supplies, for six models at once, the token axis this page keeps recording as missing — per-task tokens, turns' worth of cost, and latency, all under one price table. And on that axis the start-smart default does not hold here. Dividing the paper's per-model baseline quality means by its per-model baseline cost (arithmetic on two verified tables, not a figure the paper prints):
| Model, baseline arm | mean quality | $/task | quality per $ |
|---|---|---|---|
| Palmyra X6 | .789 | $0.25 | 3.16 |
| Claude Sonnet 4.6 | .785 | $0.24 | 3.27 |
| GLM 5.1 | .752 | $0.21 | 3.58 |
| Gemini 3.1 | .765 | $0.19 | 4.03 |
| Gemini Flash 3.5 | .740 | $0.18 | 4.11 |
| Qwen 3.6 | .710 | $0.16 | 4.44 |
The ordering is monotone and points the wrong way for "start with the most capable model": the two strongest models are the two worst per dollar, and the capability spread that buys the premium is eight points of a capability-mean score. Three caveats before this is treated as a refutation — the arms are not matched on quality (this is a ratio, not an iso-quality comparison), the workload is an enterprise-assistant task set rather than the long-horizon agentic work where extra capability is supposed to earn its price by converging in fewer turns, and quality is LLM-judged at n = 22. But it is the first place in the corpus where per-model cost and per-model quality are published together on non-Anthropic infrastructure, and the correlation between price and cost-per-task is positive, not negative.
Discount for a total conflict of interest: all 33 authors are Writer employees, the last author is co-founder and CTO, the harness under test is Writer's, the baseline is Writer's own superseded loop, and Palmyra X6 is Writer's model. The paper self-discloses and its design is auditable (frozen baseline, locked prompts, identical judges and price tables, candidate failures scored not excluded). Full treatment, including the mechanism inventory and the "harness leverage" finding, at Orchestration Sets Token Economics.
The production bench: Databricks on its own codebase (2026)#
Writer's paper is a vendor benchmarking its own product. Its production counterpart landed five days later and, on the model axis, points the other way. Databricks reports, via The Register (Thomas Claburn, 2026-07-13, case-study, secondary reporting of Databricks' benchmark blog post and CTO Matei Zaharia's social posts), an internal coding benchmark built from real engineering tasks its own staff performed against its multi-million-line codebase — built, Zaharia says, because models are tuned to public benchmarks like SWE-Bench, which the article notes OpenAI has called "broken."
| Model | $ / task | Task success |
|---|---|---|
| Opus 4.8 | $1.94 | 87% |
| Sonnet 5 | $2.09 | 81% |
| GLM 5.2 (Z.ai, open weight) | $1.28 | "statistically tied with Opus 4.8 on quality" |
"Cheaper per-token does not imply cheaper per-task. For example, Sonnet 5 costs less per token than Opus 4.8 but used more tokens, resulting in higher cost and lower quality." — Zaharia
The Anthropic pair is this page's thesis, measured by a third party. Sonnet 5 is "around 1.7x cheaper" per token — consistent with the published list prices ($3/$15 vs $5/$25 per Mtok = 1.67×, arithmetic here, not a figure The Register prints) — and 8% dearer per task, at six points less success. Both non-price terms move against the cheap model at once: more tokens per attempt and more attempts. This is the first instance in the corpus of the exact claim Anthropic makes about its own models being confirmed on someone else's production codebase by someone else's measurement. One ambiguity: the article never says whether $/task is per attempted or per completed task. If attempted, dividing through by the success rates widens the gap to roughly $2.58 vs $2.23 per completion (again arithmetic, not a published figure).
The GLM arm breaks the tidy version of the rule. The open-weight model lands in the top capability tier, statistically tied with Opus 4.8 on quality, at $1.28/task — 34% below Opus. (No per-token price is given for it; the article's framing is that "open weight models like Z.ai's GLM 5.2 are competitive with frontier models," and open-weight serving is the cheap end of the menu.) So the cheapest arm on offer is also the cheapest per task, at parity. "Cheaper per token implies dearer per task" is therefore not a law about price tiers; it is a claim about tokens-to-completion. Sonnet 5 lost because it burned more tokens and finished less often; GLM 5.2 does neither. Read strictly, the start-smart default would have selected the second-cheapest option available on this workload. (This figure exists in the raw only because the ingest pass rebuilt the article body from curl'd HTML — WebFetch silently dropped it.)
Two benches, one contradiction, worth keeping visible. Writer finds quality-per-dollar falling monotonically with model strength on an enterprise-assistant task set; Databricks finds the stronger Anthropic model cheaper per task on a long-horizon coding task set. Neither is Anthropic. The reconciliation the two support jointly — not proven by either — is that the convergence term this page rests on (fewer turns, fewer retries) only dominates where there are many turns to save, which is the coding regime and not the assistant regime. Note the asymmetry in what each can be trusted for: Writer is empirical with a total COI and a full methods section; Databricks is case-study, relayed second-hand, with no task counts, no variance, no n and no confidence intervals behind "statistically tied." The agreement between them is on the negative claim (listed per-token price is a bad predictor of the bill), which is the claim both sets of numbers actually support.
A third-party generalization the article cites, unread here. The Register points at an academic result (arXiv 2603.23971, March 2026): in about a third of the model comparisons the authors ran, the model with the lower listed price ended up costing more — "Gemini 3 Flash's listed price is 80 percent cheaper than GPT-5.4's, yet its actual cost across all tasks is 38 percent higher." That is the broadest statement of this page's thesis anywhere in the corpus and it is currently a pointer, not evidence: the wiki has not read the paper, the figure survives in the raw only because of the same HTML repair, and "about a third" is also the rate at which the inversion does not fire. Worth ingesting.
Interest to declare. Zaharia says these results are why Databricks built Omnigent, a wrapper for combining and swapping coding agents — so the harness half of the write-up (see Orchestration Sets Token Economics) is adjacent to a product, though no Databricks harness or model is in the comparison and Databricks sells neither of the models it prices.
Where benchmarks stop helping#
Anthropic's own guidance says public benchmarks are "helpful directional guides" that break down exactly where the choice gets expensive: at the Opus/Fable tier the models "solve almost all of the questions on the test" (saturation). The recommended replacement is a curated set of problems drawn from production, including tasks where current tooling falls short, with success criteria the team defines — Production-Sourced Evaluation as vendor advice, and the operational form of Evals as Product Spec.
The circularity is worth naming: the selection framework's two hardest questions (is this task hard? do the unit economics work?) both resolve to "build an eval," and the vendor's own benchmarks are declared insufficient for the tier where the decision matters most.
Campaign cost is not cost-to-shipped#
The largest published cost-per-task figure in the corpus is the ~$165,000 of API tokens for the Bun Zig→Rust port, set against a stated counterfactual of three engineers for a year. An independent audit of the same project (Lockwood, 2026-07-27, case-study) shows why that comparison is not apples-to-apples, and the correction generalizes well beyond Bun:
- The token figure is bounded at green, not at shipped. $165k covers the port through the May 14 merge to
main. It excludes CI (a continuously-running Buildkite cluster), the employee time spent monitoring and re-prompting ~50 workflows, and the post-merge stabilization tail — which was still visibly running eleven weeks after the last release tag, with the agent PR queue growing from 1,277 to ~2,475 in eighteen days. - The counterfactual is fully loaded; the measured side is not. "Three engineer-years" carries salary, overhead, review, and CI. "$165k" carries tokens. Comparing them favors the agent path by construction unless the same boundary is drawn on both sides.
- The extrapolation is not the correction. Lockwood's ~$800k comes from assuming the project still burns $10k/day — an assumed rate, not an observation, over a boundary he does not define. His direction holds; the number is speculation and should not be repeated as a measurement.
So the practical rule when reading any cost-per-task claim about an agent campaign: ask which cost line the number is drawn at. Cost-to-first-green is the cheapest honest boundary to report and the one least likely to be the number a buyer cares about. This is the accounting analogue of Verification as the New Bottleneck — the expensive part of the work is downstream of the part that is easy to price.
Tension: the strongest model is not always the best component#
The "start smart" default is stated for a workload, implicitly a single agent. It sits badly against the empirical multi-role result in Client-Side Agent Optimization: on HotpotQA, Opus 4.6 was the worst planner of 81 combinations (it answered from parametric knowledge instead of delegating to the solver's search tools), and a Ministral 3 8B planner paired with an Opus solver scored 74.27% vs. 31.71% for Opus-as-both. AgentOpt also measured 13–32× cost gaps between equally-accurate pipeline combinations.
Weighting by evidence: AgentOpt is empirical and multi-role; Anthropic's guidance is vendor-claim and single-role. They are not strictly contradictory — "start smart" is defensible as a first configuration precisely because it makes failures diagnostic — but "start with the most intelligent model" is not safe advice for per-role assignment inside a pipeline, and nothing in the vendor guidance flags that boundary. Anthropic does gesture at it once, obliquely: Sonnet is recommended for "high-volume sub-agents in multi-agent orchestration."
The other vendor's answer: don't publish a rule, ship a default#
Anthropic's response to "which model and how much effort?" is a published decision procedure for the user (start strong, dial down; four selection questions; build an eval). OpenAI's, as of the ChatGPT Work launch, is the opposite: be opinionated in the product and hide the axes. Akshay Nathan (Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI, practitioner-opinion):
"We want this default to be the best possible. Like, we wanna be opinionated about the default… We have for power users options under the hood. One could argue that there might be too many right now, and we're working on simplifying it… but the default should be good enough."
The mechanics of the collapse are explicit. There are "32 options" across model classes and reasoning levels; the shipped control is a one-dimensional slider — "reduce it to one dimension even though there's multiple dimensions… speed and efficiency on one side, quality and thoroughness on the other." That is this page's two axes (model class × effort) projected onto a single user-facing scalar, with the projection chosen by the vendor.
Three things worth separating:
- The escalation trigger is the same. Nathan's advice for when to leave the default — "if you're not seeing either the efficiency on the cost side or the quality on the intelligence side" — is the same two-sided test Anthropic's framework encodes. Neither vendor claims a rule for which direction to move first; both say measure your own workload.
- The exposure philosophy diverges, and only one side is falsifiable. A published rule can be wrong in public (as Client-Side Agent Optimization shows "start with the strongest model" is, per-role). A tuned default can be wrong silently, and the user has no way to know the projection is costing them. Nathan's own hedge — "there's a preference on, for you as an individual, how do you like to collaborate with the models" — concedes the projection is not user-invariant.
- Neither publishes tokens-per-task. The gap named above for DeepMind and Anthropic holds here too: no turn counts, no thinking budgets, no cost-per-task figures anywhere in the account. The default is asserted to be best "for most use cases" with nothing behind it.
The practitioner counter-current from the same episode: Vibhu reports instructing nearly every long-running task to "use sub-agents where possible," for wall-clock and to "offload to a lot of smaller, cheaper models" — the cost-per-token intuition this page argues against, running in the wild at the sub-agent layer, where Client-Side Agent Optimization suggests it may actually be right.
Who picks the model (Willison, July 2026)#
Every source above answers which model or harness is cheaper. Simon Willison (2026-07-03, practitioner-opinion) changes the question to who decides, and the answer is the model. Relaying a tip from Jesse Vincent, he prompted Claude Code:
"For all coding tasks use your judgement to decide an appropriate lower power model and run that in a subagent"
No routing table, no threshold, no per-task rule — the selection decision itself is delegated. That is a control-plane choice rather than a pricing one, and it is the general form the Claude Code team recommends: see Harness Shrinkage as Models Improve, where replacing a hard rule with "use your judgement" is the same move Anthropic made inside its own system prompt.
The policy the model then wrote into its own memory file is this page's conclusion restated unprompted — cost/efficiency, "implementation work rarely needs the top-tier model; judgment, review, and synthesis stay with the main loop," with Sonnet named for substantive implementation and Haiku for trivial edits. That is Cursor's planner/worker split reached by one developer in one sentence. It is also the arrangement Client-Side Agent Optimization identifies the failure mode for — a strong model that answers instead of delegating — which is precisely the failure this prompt asks the strong model to police in itself.
Nothing here is a measurement: "so far it seems to be working well" and a Fable allowance "shrinking less quickly than before," with no baseline, no task set, no dollars, and no check on whether the self-chosen tier was the right one. Record it as the practice in the wild — the second instance of the counter-current noted just above, differing in the one interesting place. Vibhu names the tier; Willison delegates the naming.
A third unit: per-minute presence (OpenAI, September 2026)#
Everything above prices work in tokens and compares it against tasks. GPT-Live-1's API launch (OpenAI, 2026-09-10, vendor-claim) introduces a unit neither metric covers: $0.05 per minute for a full-duplex voice layer, with the backend model "paired" behind it and billed separately in the ordinary per-token way.
Why that matters to this page rather than only to the voice pages:
- The two halves of one agent now bill in different units. The interaction half is bought by time present — it costs the same whether the user talks continuously or sits in silence, and whether the backend reasons hard or not at all. The background half is bought by work done. So the cost-per-task framing applies cleanly to only one half of the system; the other half's cost is set by conversation duration, a variable no model choice controls.
- It inverts the effort lever. This page's second axis is effort: dial reasoning down and the token bill falls. On a voice agent, dialling effort down shortens the backend bill and lengthens nothing, but dialling it up adds wall-clock time the user spends waiting — which the per-minute meter charges for. Reasoning effort therefore has a cost in both units, in the same direction, with the second one invisible to any token accounting. That is the first instance in this corpus where thinking longer is billed twice.
- It is a clean case for "the expensive component is the one with the hard constraint." OpenAI meters the part that cannot stall (Live-Path Minimalism) and commoditises the part that can, to the point of naming third-party models as acceptable backends. Cost-per-task reasoning applies to the commoditised half; the other half is priced like a leased line.
No cost-per-task measurement exists for any of this — the post publishes a price and no usage data, so nothing here tests the page's central claim. It is recorded as a unit the framework has to accommodate, not as evidence about model selection.
And five days later a third party measures the unit on a workload (Google, 2026-09-15). Google's Gemini 3.8 Live launch (vendor-claim) reproduces an Artificial Analysis chart titled Cost per Hour of Input Audio, computed on a Big Bench Audio subset: Gemini 3.8 Live $0.84, Gemini 3.1 Flash Live minimal $1.50 / high $1.75, Gemini 3.8 Live Extended Thinking (high) $3.50, Grok Voice Think Fast 2.0 $4.80, GPT-Live-1 + Astra (medium) $5.83. Three things this adds that a price list cannot.
- It is a measured bill on a stated workload, which is the artifact this page keeps saying nobody publishes. Not a rate to be multiplied by an unknown token count — an actual cost of running a benchmark subset, normalised by the one quantity the user controls. Its defect is the mirror image of the usual one: the denominator is disclosed and the numerator's composition is not. Nothing states whether backend tokens, output audio, or thinking time are inside it, and the methodology footnote printed on the chart points at Google's page rather than the evaluator's.
- The effort dial is now visible in the per-time unit, and it costs 4.2×. Within one vendor's own line, on a fixed hour of input audio, moving from 3.8 Live to Extended Thinking multiplies the bill $0.84 → $3.50 while the composite quality index moves 76.0 → 82.6 (+6.6 points) and the agentic component moves 30.1 → 68.6 (+38.5). Because the denominator is the user's talking time and not the model's work, the whole 4.2× is thinking per minute of conversation. That is the section above's "thinking longer is billed twice" becoming measurable rather than structural — and it is a cost-per-task result only if you accept an hour of input audio as a task, which is the same move as accepting a benchmark item as one.
- One arithmetic check, flagged as an inference and not a finding. $5.83 per audio hour is $0.097/minute for GPT-Live-1 + Astra, against OpenAI's published $0.05/minute for the front-end layer alone. The two are consistent with roughly half the realised cost of a delegating voice agent sitting behind the seam on this workload — but only if AA's hour of input audio maps to a billed minute (it need not: a benchmark subset is not a conversation, and wall-clock is not audio duration once the model is thinking), and only if the numerator contains what the guess assumes. Recorded because it is the first time anything in this corpus lets the two halves of a priced split be compared at all, and deliberately not used as a number anywhere else in the wiki.
Google itself publishes no price for either model, so none of this is purchasable as a rate — see Interaction / Background Model Split for why the priced-boundary shape still has exactly one vendor instance.
The buyer side, at market scale (Ramp, September 2026)#
Every measurement above is a bench or a production run inside one organization. Ramp's September 2026 AI Index (Ara Kharazian, empirical, corporate-card and bill-pay records for ~70,000 US businesses; instrument and COI at Ramp) is the first source in the corpus that reports what the buying population does with the choice this page frames — and the population is moving away from the strongest model.
- Frontier share of tokens is falling. Opus, Fable and Sol held 52.5% of tokens in the week of August 2 2026 and 44.7% in the week of August 30, while the standard tier (GPT-5.6 Terra, Claude's Sonnet series) rose 25.7% → 35.8%. The letter's stated explanation is buyers "imposing company-wide defaults that reduce usage of frontier models, saying standard models are still highly performant and also more cost effective."
- The blended price of a token collapsed. The effective price per million tokens fell 41% to $0.68 from a $1.15 peak in March 2026. It is a blended measure — it mixes announced list-price cuts (both labs made them in the month before publication) with the tier shift above, and separates neither.
- The two labs' effective prices have diverged, which the letter never mentions. In the recovered daily series the blended price of an OpenAI token fell $0.681 (2026-03-01) → $0.339 (2026-09-02) while Anthropic's fell $1.512 → $0.891. Anthropic's effective token ends the window at 2.6× OpenAI's and has been roughly flat since July ($0.888 on 2026-07-01), so nearly all of the recent decline in the blended index is OpenAI-side. (Vault arithmetic on the chart dataset. The CSV endpoint is CDN-cached and its tail lags the prose, so the levels here are the series' and the current value is the prose's $0.68.)
What this is and is not evidence for. It is not a cost-per-task measurement: Ramp counts dollars and tokens and never tasks, turns or completions, so nothing here tests whether a stronger model finishes in fewer turns. What it supplies is revealed preference at population scale, and the rule the population is using is not this page's. Buyers are applying a satisficing test — "still highly performant" and cheaper per token — which is the price-per-token reasoning the start-with-the-strongest default was written against. Two readings survive and this instrument cannot separate them: either the market is making exactly the error named at the top of this page (optimizing the visible unit while tokens-to-completion, the invisible one, gets worse), or the standard tier has crossed the quality bar for most production work and the vendor default is out of date. The Databricks bench above is the nearest thing to an adjudication and it splits the same way — Opus cheaper per task than Sonnet inside Anthropic's line, an open-weight model cheaper than both at tied quality.
One consequence for everything above. Every dollar figure on this page — Cursor's $1,339 against $10,565, Writer's per-task costs, Databricks' $1.94 against $2.09 — was measured against a price table that has since moved ~41% at the blended level, unevenly across vendors and tiers. None of those comparisons is invalidated, because each is internal to one study run against one day's prices; none of them can be carried forward as a current dollar amount.
What the buyer wants the meter to count (a16z via TechCrunch, September 2026)#
Everything above argues about which denominator the seller should reason in. Julie Bort's TechCrunch piece (2026-09-03, practitioner-opinion — a news article reporting a VC post that was not fetched) is the first source here that asks the buyer which denominator they want on the invoice, and the answer runs the same direction as this page's thesis from the opposite side of the table: in a16z's survey of 50 technical AI buyers, more than half want AI fees tied to the work produced or other outcomes rather than to usage like the number of tokens consumed.
a16z partners Tugce Erten and Sarah Wang put the argument in product terms rather than cost terms — pricing "around the recognizable work" (reports processed, tickets closed, leads generated) is what proves the product's worth and makes it "economically valuable to both sides," while token metering is a SaaS-era seats-and-storage habit applied to a product whose value is not proportional to its consumption.
Why this is not the same claim as this page's, despite sounding like it. This page argues that cost-per-task is the right unit for a model-selection decision, because tokens-to-completion varies with capability. a16z argues that outcome-per-fee is the right unit for a commercial one, because token counts are not legible as value to the person signing. They agree that the token is the wrong unit and disagree about who should be exposed to its variance: the cost-per-task framing tells a buyer to look past the sticker price to the bill; outcome pricing tells the buyer not to look at the bill at all and moves the whole token-count risk onto the seller — precisely the risk McKinsey's Chui describes below, where per-token prices fall while tokens consumed rise faster. An application vendor that grants this demand is writing the gap between price-per-token and cost-per-task directly into its gross margin.
It is also, notably, the inverse of the revealed preference measured above. Ramp's ~70,000 businesses are managing their own token bill down by imposing company-wide model defaults — buying tokens more carefully. a16z's 50 buyers want to stop buying tokens. These are not contradictory: the first population is buying model APIs, the second is buying applications built on them, and the natural equilibrium of both statements is that token-metered risk concentrates in the middle layer. Neither survey measures a task, and n=50 with no published sampling frame is the weakest population instrument on this page — carried as a demand signal about billing units, not as evidence about cost.
When the bill constrains the work (McKinsey, August 2026)#
Every measurement above prices a task. The state of AI in 2026 (McKinsey / QuantumBlack, 2026-08-25, empirical but self-reported, 1,719 respondents in 97 nations) supplies the organizational consequence at population scale: about 20% of respondents report that AI-related operating costs, including token costs, have constrained their organization's AI use (Exhibit 8) — 12% to 25% by industry, broadly flat across company sizes, and about one in ten on each of chatbots, agents and coding agents taken separately.
The shape of the constraint is the part that belongs on this page. Software coding agents are where it concentrates, and it concentrates upward: among AI high performers (the 6% attributing at least 5% of EBIT to AI) 18% report cost-constrained coding-agent use, against 6% of all others — the only tool class in Exhibit 14 where the leading cohort reports more constraint than the rest (chatbots 10 vs 9, other agents 8 vs 9, specialized tools 6 vs 6). Michael Chui states the mechanism in the article's own commentary:
"They've discovered that AI isn't 'too cheap to meter' when it comes to agentic software development and complex reasoning tasks that benefit most from frontier models: Even as per-token costs have declined, the number of tokens consumed and generated has increased even faster."
That is this page's thesis restated as a budget observation by a third party with no model to sell: a falling price per token is not a falling bill, because the unit of consumption is a task whose token count is set by the work, not by the price list. What the survey cannot do is test any version of the cost-per-task rule — it measures no tasks, turns, completions or model choices, and 28% of respondents reporting more than 10% of ICT budget on AI is a spend level, not a denominator. Treat it as the demand-side context in which the model-selection question is now being asked, not as evidence in it.
The full stack, measured end to end (Uber, 2026-08)#
Every source above prices one lever — model, harness, cache, or the buyer's demand for a different meter. Uber's own account of its "software factory" (Medisetty, 2026-08-27, case-study, first-party engineering blog) is the first in the corpus to publish a cost equation and price every term of it at once, on production traffic: >70% of PRs agent-attributed, 3,600+ agent skills, 30K+ skill executions/day.
The equation. Total Spend = Users × Sessions/User × Turns/Session × Requests/Turn × Tokens/Request × Price/Token. Grouped by effect: {Users, Sessions/User} grow (adoption); {Turns/Session, Requests/Turn, Tokens/Request} and {Price/Token} shrink (optimization targets) — the same six-way decomposition this page's contributors each price one or two terms of. With one model held fixed February–July, cost per 1,000 requests fell 34% from its peak and cost per session fell 52% from its June peak — both measured at each metric's trough, both having ticked back up somewhat by August, a caveat the source states about its own chart and this page repeats rather than smooths over.
Pareto model selection, on a named production point. uReview (AI code review for all PRs) is benchmarked on real PRs with known bugs, graded by difficulty, scored on precision/recall/F1 plus cost/review, latency, timeouts and noise. Ten configurations plot cost-per-review (log scale) against F1: two closed-frontier points near $2.2–2.4/review at F1 0.47–0.48; a third closed-frontier point at ~$1.05, F1 ≈ 0.39; the production choice — a genericized "Frontier Model C" at $0.47/review, F1 0.50, marked "what we run"; two open-weight points at $0.28 (F1 0.48) and $0.06 (F1 0.31, cheapest and worst). The shape is this page's thesis in one chart: the production point is neither the cheapest nor the highest-F1 option, it is the one the team chose off the Pareto frontier. Model names are genericized, so this source corroborates the shape of Pareto-driven selection but cannot be cross-referenced name-for-name against any model in the table above.
A lever this page hasn't priced: which model plays the subagent role, as a default rather than a rule. Uber's "subagent default setting has proven to be the most impactful lever" for tokens/request — subagents default to a weaker, cheaper model (manual override allowed) because their tasks are well-defined and don't need frontier reasoning, while the primary model retains decomposition and evaluation. No dollar figure attaches to it, but it is the production-scale version of Parallel Agent Orchestration's Sonnet-for-subagents guidance, run as an org default rather than one practitioner's self-delegated prompt (see that page's own new section on Uber's four-layer taxonomy).
Tokens/request, the rest of the lever list. Automatic compaction at 400K tokens even on 1M-context models; reasoning effort defaulted to Medium (output/thinking tokens billed at a multiple of input); prompt-cache TTL chosen by idle-gap distribution rather than a fixed default (full treatment on Prompt-Cache Economics); MCP tool schemas replaced by CLI-resolved calls and on-demand tool search (100+ pre-loaded tools cost ~50–70K tokens of schema per session, reduced to "near zero" via search+CLI); code-mode batching, measured on 5 identical SQL queries in the same Claude Code session — 55–71% token reduction on small result sets purely from removing polling/schema overhead, ~100% on a wide 50-row result (1,431,594 → 900 tokens), and >90% claimed for bulk N-turn-to-one-script workflows; SaaS MCP servers (a workspace suite's 49 tools costing ~22K tokens of schema) routed through the same gateway-plus-CLI-plus-code-mode-skill pattern.
Requests/turn: context grounding as a search-tax reduction. An AI Context Graph (24M nodes, 80M edges, 86 node types, 117 edge types, 30+ source systems) turned a table-usage question that took an ungrounded agent 20 minutes, 2 subagents and 3 errors — and a wrong answer — into a 38-second grounded correct one. A single paired example, not a distribution, but the largest single before/after gap on any lever in this source.
Visibility as its own lever. Live per-harness/per-user cost counters, a shared harness spend pool with Slack nudges at 50/80/100% of expected spend, and a session-analysis dashboard that flags 16 anti-pattern categories (suboptimal model routing, context bloat from persisted large MCP payloads, cache-expiration rebuilds, prompt-init overhead) with a per-finding dollar estimate — the operational form of Verification as the New Bottleneck applied to spend rather than quality: built into the runtime rather than a report someone has to request.
Discount appropriately. case-study, first-party, unaudited — no ablation isolates any single lever's dollar contribution to the 34%/52% headline, several figures are single illustrative examples rather than distributions (the context-graph 38s/20m pair, the code-mode SQL-query table run once "in the same session"), and both headline percentages are measured at a trough that had already partially reversed by publication.
A fifth unit: which harness, holding the model fixed (Arena.ai, 2026-09)#
HarnessTax runs the harness-as-cost-lever question as a controlled grid rather than a single swap: 7 models × 3 harnesses (Claude Code, Codex CLI, Pi) × 2 benchmarks, bootstrapped CIs on both cost and success. The headline is the same shape as Writer's harness swap above, sharper: Claude Code costs a geometric-mean 2.0× Pi and 1.6× Codex CLI on SWE-bench Lite, for a success-rate difference that stays within ±2%. Claude Fable 5 is the cleanest single point — 97.8% (Claude Code, $1.329) vs. 96.7% (Pi, $0.666) — same turn count (15.3 vs. 15.4), double the bill. The source traces the multiplier to the first model call: Claude Code's mean first-call context runs ~13.7× Pi's (27.0k vs. 1,972 tokens), driven by a 23-tool, 77,000-character tool schema against Pi's 4 tools and 2,873 characters. Unlike the Writer and Databricks sources this page already carries, HarnessTax has no vendor stake in any of the three harnesses it grades.
A sixth unit: cost per workflow, priced by a vendor that sells the cheap end (TypeSafe, 2026-09)#
TypeSafe AI's Jev launch (2026-09-15, vendor-claim) prices a whole decision graph — four code-defined workflows, each a chain of model readings and hand-written branches — and plots accuracy against cost per workflow run. Its non-LLM decision model sits at ~$0.0004/workflow and 68% agreement with a two-LLM reference; the frontier line then runs through GPT luna ($0.0035, 67%), terra ($0.03, 68%) and sol ($0.085, ~74%), with Opus 5 at ~$0.18/~73% (chart values, approximate). Two things carry over to this page. First, this is the start-smart default's mirror image: for decision-shaped calls, the vendor's case is that the cheapest adequate model wins by two orders of magnitude and the last ~6pp of agreement costs ~200×. Second, how the model is called moved cost as much as which model: every LLM run inside the fixed workflow cost about half what the same LLM cost doing the logic in chain-of-thought, and agreed with the reference more — the orchestration-over-model-menu result, here from a party selling neither harness. Every lever (workflows, reference, the wrapper the LLMs were called through) is the vendor's; see Typed Decision Verifiers for the discount.
The cascade form: "cheap by default, frontier on exception" (LangChain, 2026-09-25). LangChain's Jev integration post (vendor-claim, integration partner) states the routing rule this sixth unit implies, borrowing Jaya Gupta's "Great Unbundling of Intelligence": the shift is from "frontier by default and optimize later" to "cheap by default, frontier on exception." Its one concrete instance is a confidence-thresholded cascade — Browserbase rebuilt Stagehand's act() so Jev picks the browser action and element, and anything under 0.7 confidence goes to an LLM; median latency fell 1.97 s → 0.46 s (Browserbase's early testing, secondhand). That is FrugalGPT-style per-call routing, the lever Crystallizing Agent Work into Workflows calls orthogonal to its own, and it is the opposite default to the start-smart rule above: the cheap model is the default and the frontier model is the exception handler. Two things the post does not publish decide whether it lowers cost per task rather than cost per call — the fallback rate (what share of calls clear 0.7) and task success before vs after. A cascade whose cheap stage is confidently wrong ships the error without ever paying for the escalation, which is exactly the precision failure the shared board's retry-gate experiment measured on Jev (216 of 403 routed items already correct). Latency is the only axis reported.
Connections#
-
The Price of Fixed Capability — this page's unit distinction applied to a price trend: once reasoning models exist, a per-token series stops tracking what a fixed score costs, and the per-task series falls ~13×/yr at the frontier against Ramp's ~3×/yr realized blended per-token price
-
Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves — the fifth cost unit above: harness choice, isolated from the model, at 2.0× (SWE-bench Lite, Claude Code vs. Pi) and traced to a ~13.7× first-call-context gap
-
Shared-Budget Compute Allocation — the batching move priced. Amortizing cost by putting several tasks in one request under one budget is not free: at N=20 questions under a shared budget every one of seven reasoning models scores worse than the same total tokens split evenly, by 5.0 points, because the model spends by prompt order rather than by value. The penalty is an allocation failure, not a token-accounting one, and it reverses at small batch sizes (+2.6 at N=5) — so "how many tasks per call" is a cost-per-task parameter with a measured optimum, not a pure saving
-
GDPval Benchmark — what the GDPval-AA row in the model-selection table above is built on, and where the per-model split it implies gets its evidence. The primary paper (ingested 2026-09-10) puts numbers on it: win rate by deliverable file type, with Claude Opus 4.1 leading pdf (45%), xlsx (43%), pptx (45%) and "other" (48%), and GPT-5 high leading pure text (23% vs Claude's 14%) — the split the vendor advice recommends, measured. "Pick the model per task type" is the benchmark's own finding before it is vendor advice. It also supplies the cost-side result this page has been missing a shape for: pricing expert review into the loop turns GPT-5's 474× naive cost advantage over an unaided professional into 1.63×, and turns GPT-4o's into 0.53× — cost per task is not a property of the model alone but of the model's win rate against the quality bar, since every failed attempt is charged twice
-
AI Product Economics Maturation — where the billing-unit question is priced from the seller's P&L: consumption-based pricing at 42% and outcome-based at 23% of ~305 AI builders, blending 1.7 models each, against the >50% buyer-side demand recorded above — and the overrun ranking (token spend first, one workflow going $0.10 → $1.50+ per run) that says what an outcome-denominated seller is absorbing
-
Firm AI-Spend Intensity and Headcount Growth — the home of the payment-rail instrument behind the buyer-side section above, where the same September index is read for adoption and per-employee spend, and where its revision behavior (dollar series revise upward for months after first print) is documented
-
Crystallizing Agent Work into Workflows — the cost frame taken to its limit: for a solved, recurring problem the per-task cost falls to zero rather than to a cheaper model's price, because inference is removed instead of made cheaper (0→45% deterministic execution and >70% per-incident cost reduction over eight months in production, with no controlled counterfactual)
-
Orchestration Sets Token Economics — the lever one layer up, and the one this page never varies. Holding the model constant and swapping only the orchestration layer cuts cost per task 41% and tokens 38% across six models with no exceptions, which is a larger spread than the model menu itself (36%). It is also the first source to publish per-model cost and per-model quality together off Anthropic infrastructure — and on that data the quality-per-dollar ordering runs opposite to the start-smart default. Discount for a total vendor COI (33 Writer authors benchmarking Writer's harness against Writer's own frozen predecessor). It now also hosts the harness half of the Databricks bench whose model half is above — three shipped third-party harnesses on a real codebase, same success rate at "2x less cost" for the minimal one, and a 3.13× per-task context spread — which is the same lever measured by a party that sells none of the harnesses in it
-
Prompt-Cache Economics — the token axis this page keeps calling missing, actually published — by a third party, on a $98.96 end-to-end budget reconciled against Anthropic's invoice to within 1%. It also breaks the page's framing in a useful way: per-token price is not a static constant to multiply by a token count, it is a function of cache state, prefix size, and call count, which is exactly the term every cost-aware routing paper holds fixed. And the cost-per-task thesis runs backwards in at least one measured case — query-aware prompt compression cuts tokens 3× and raises the bill 40.1% over sending nothing compressed, because the cache-write tax on the busted prefix exceeds the read savings
-
Tool-Output Pruning — the two axes disagreeing inside a single system, which is the cleanest form of this page's problem. On SWE-Bench Verified, SWE-Pruner Pro posts the largest input-token reduction on one backbone (−13.5%) while running the highest API-call count of any method on both (111.8 vs 94.8; 139.8 vs 131.9) — pruning shortens each call and lengthens the trajectory, and the authors explicitly refuse to collapse the two into one efficiency number because they move in opposite directions across backbones. The reverse case is on the same table: on MiMo-V2-Flash every pruner raised per-trajectory input tokens (+6.6% to +14.9%) and every pruner improved the resolve rate, so the technique earns its keep on quality while losing on cost. Neither result is expressible in a price-times-tokens model
-
Context Lifecycle Management — the cost axis this page says nobody publishes, from the context side: pruning tokens breaks the provider prefix cache, so token reduction and billed cost can move in opposite directions. Self-GC prices the commit (
CommitBenefit ≈ N_future·(C−C′) − L_cache_break − L_GC) and reports a 0.3 expected-pruning break-even — the corpus's first published threshold for "is this context cut worth the cache break?" — alongside a measured 10–15% production input-token reduction that it explicitly declines to call a billed-cost saving -
Client-Side Agent Optimization — the empirical counterweight: model selection evaluated at the level of full pipeline combinations rather than per-workload, where the strongest model can be the worst component. Cursor's four mixes are the same abstraction run on a production build, and they add the cost-side coupling: a planner's bill includes the tokens it causes its workers to spend
-
Parallel Agent Orchestration — the harness these cost figures were measured inside, and why the comparison is credible at all: matched task, matched models, matched time budget, held-out oracle
-
Cursor — the vendor publishing the figures, and the conflicts of interest to net out of them
-
Optimizer–Evaluator Decoupling — the advisor strategy is that invariant reached from the cost side; the advisor is an independent grader that also happens to be cheaper than running it as the executor
-
Production-Sourced Evaluation — the vendor's own recommendation once benchmarks saturate: curate the eval from production traffic
-
Evals as Product Spec — what the selection framework defers to when its two hardest questions come due
-
Large-Scale Test-Time Compute — effort level is the inference-budget thesis productized as a dial; "how capable is the model?" is ill-posed without naming the budget, and the guidance concedes this by treating class and effort as one grid
-
Measuring Beyond Accuracy Saturation — the research-side statement of the Opus-vs-Fable problem: benchmark scores that no longer separate models, while other measurable axes still do
-
Unproductive Self-Verification — the failure mode that inverts the cost-per-task argument: more capability spent on rumination rather than convergence
-
Inference Efficiency as Capability — the supply-side twin: the same price-vs-capability trade seen from the model builder, where Gemini 3.5 Flash-Lite's +67% output price buys +78% relative on one benchmark and +2% on another
-
Compute-Controlled Benchmarking — the evaluation-side statement: a published price is a rate, and no vendor publishes the tokens-per-task that turns it into a bill
-
Claude Code Best Practices — where this guidance is applied per session (effort defaults, context budget) and where the sub-agent-overhead question that this page's cheap-fan-out counter-current bears on lives
-
Shared Harness, Differentiated Surfaces — the same choice made the other way: OpenAI collapses model class × effort onto a one-dimensional slider and hides the rest behind an opinionated default, rather than publishing a selection rule
-
Harness Shrinkage as Models Improve — the third answer to the same question, and the only one that isn't a rule at all: hand the selection to the model. Replacing a hard rule with "use your judgement" is the move Anthropic ran on its own system prompt, and a practitioner has now pointed it at model routing — a published decision procedure, a hidden vendor default, and a delegated judgement are three different places the choice can live
-
Anthropic — publisher of the guidance
-
Open Weights as Competitive Strategy — this page's arithmetic scaled up to a national argument: Ng treats the price of intelligence as an input cost compounding through an entire downstream industry, so builders paying 3× lose the application layer regardless of who holds the frontier. The Databricks bench above is the closest thing the vault has to evidence for it — an open-weight model cheapest at tied quality — and also the clearest reason to hold the claim loosely, since this page's whole finding is that per-token price is the wrong denominator for exactly this kind of comparison
-
Claude Fable 5, Claude Opus 5, Claude Sonnet 5, Mythos Model — the classes being chosen between
-
Dynamic Workflows: An Algebra for Agents — the argument at its largest published scale: ~$165k of tokens (5.9B uncached input, 690M output, 72B cached reads) against a stated counterfactual of three engineer-years, on a task the team says it would otherwise not have attempted — and, per the independent audit, a figure bounded at cost-to-green rather than cost-to-shipped
-
Verification as the New Bottleneck — the accounting analogue: the work that is easy to price (getting to green) is not the work that dominates the bill (getting to shipped)
-
Agent-Authored Harness Optimization — the campaign-cost question asked of an optimization campaign itself: ~$680 and ~1B tokens spent to move one benchmark run from $79 to $49.8, with no stated payback boundary — the same cost-to-green-vs-cost-to-shipped ambiguity as the Bun figure, one level up
-
Knowledge-Centric Self-Improvement — a self-improvement comparison run entirely in dollars rather than tokens, for this page's reason ("so they reflect what cached and uncached tokens actually cost and remain comparable across methods with different cache profiles"), with meta-loop tokens charged to the baselines that spend them. The result is the rarer shape: solve rate up and cost down against every arm (SWE-bench Pro $208 vs DGM's $713), so cost-per-task fell without any accuracy trade to argue about
-
Standardize the Infrastructure, Not the Tools — the org-level precondition for this accounting: Shopify's central LLM proxy is what makes every request visible to one meter, which a fleet of separately-billed tools is not
-
Deterministic Engineering for Agent Code Review — a third-party harness where tokens, wall clock and quality move the same direction at once, reproducing Writer's harness-is-the-bigger-lever result in a different task domain. Holding the backend model fixed and swapping only the review system takes Claude-4.6-Opus from 5,664K tokens / 13m06s / 11.57% SEM-F1 to 385K / 1m23s / 25.10% — 14.7× and 9.5× reductions alongside a 2.17× quality gain, measured by a party selling neither baseline being compared. The clean read doesn't survive the other baseline, though: against Codex the token saving nearly disappears (422K vs 525K) because Codex's own loop already terminates early, so part of the saving is a property of that baseline's exploration behavior rather than of the constrained design generally — and like every other source on this page, the paper prices itself in tokens and seconds and never once in dollars
-
Agent Quality Flywheel — the inverse of this page's complaint, and worth keeping for that reason. Every source above prices itself in tokens and never in dollars; Shopify's Sidekick account prices itself only in dollars and never in a unit — "could easily cost an estimated $27M per year based on average token costs" against "closer to $1M: a 96% reduction in serving cost" (
case-study, first-party). Both figures are counterfactual annual totals, the modals are the article's own, no per-request or per-task denominator appears anywhere, and the recurring cost of running the loop that produced the saving (a daily full-parameter fine-tune, a frontier critic panel, replay passes, an expert-annotation contract) is on neither side of the ledger. So a 96% "cost reduction" is a ratio of two modelled aggregates — which is exactly as unusable for a buying decision as a token count, one direction over -
Interaction / Background Model Split — the productised voice split bills its two halves in different units: per-minute presence on the interaction model, per-token work on the background model
-
Live-Path Minimalism — why the metered half is the one that cannot stall
-
Gemini 3.8 Live — the second frontier voice launch of the month, which publishes no price at all and is costed only by a third party, per hour of input audio
-
Artificial Analysis — the evaluator publishing measured cost on a workload rather than a rate
-
Build Instead of Buy Under Agentic Coding — the purchase the falling implementation cost displaces, and the reason the token bill is now an IT-budget line rather than an experiment line: the build decision is settled in tokens
-
Interactivity Benchmarks — where the voice cost chart sits alongside the quality boards it has to be read against, and why the $/audio-hour denominator is not a task
-
Typed Decision Verifiers — the sixth unit above: cost per workflow run, where a non-LLM decision model owns the Pareto frontier ~200× below frontier LLMs at ~5–6pp less agreement (vendor-run eval)
Open Questions#
- Does "cost-per-task is lower for more intelligent models" survive measurement on non-Anthropic production traffic? The claim is stated without data and the published curves are explicitly illustrative. Still open, with a near-miss (2026-07-30): DeepMind's Gemini 3.5 Flash-Lite card is a non-Anthropic instance of the shape — +67% output price, a full agentic tier of capability — but reports no tokens-per-task, so it supplies the premise and not the measurement. Partially answered (2026-08-03), and it splits: Cursor's four model mixes are the measurement — non-Anthropic infrastructure, a real four-hour workload, matched time budgets, matched quality, published dollars. Within the planner role the claim holds: the more expensive Fable 5 planner billed slightly less than Opus 4.8 at roughly twice the per-token price, because it emitted far fewer planning tokens. At the level of the whole run it fails: the same Fable configuration came out substantially more expensive because its workers burned several times the tokens, and the most expensive run of all was the strongest model used throughout ($10,565 versus $1,339). So the thesis appears to be a claim about a role, not about a system, and no source yet measures it on a single-agent workload outside Anthropic. Sharpened, with the first counter-datum (2026-08-03): Writer's harness swap publishes per-model cost and per-model quality for six models on non-Anthropic infrastructure under one pinned price table, and cost per task rises monotonically with model strength on that workload — quality per dollar is worst for the two strongest models (Palmyra X6 3.16, Sonnet 4.6 3.27) and best for the cheapest (Qwen 3.6 4.44). It is a controlled bench rather than production traffic, the arms are not iso-quality, and the capability spread is only eight points, so it does not close the question — but the sign is wrong for the vendor guidance and the paper's own conclusion is that the model menu is the smaller lever anyway. Closest yet, and it splits again (2026-08-04): Databricks' internal coding bench — real engineering tasks on its own multi-million-line codebase, measured by neither Anthropic nor a model vendor — puts Opus 4.8 at $1.94/task and 87% success against Sonnet 5 at $2.09 and 81%, on tokens ~1.7× cheaper. Within Anthropic's own line the claim therefore holds, in the long-horizon coding regime where its mechanism should be strongest, and this is the first time it holds on a third party's real codebase. Against the wider menu it fails: open-weight GLM 5.2 is statistically tied with Opus 4.8 on quality at $1.28/task. So the surviving form of the rule is about tokens-to-completion, not about price tier — a cheaper model is dearer per task when it burns more tokens and finishes less often, which is contingent, not structural. Still not closed:
case-studysecondary reporting, production-derived tasks rather than production traffic, and no n, variance or per-arm methodology behind "statistically tied." The article's own cited generalization (arXiv 2603.23971 — a third of comparisons invert; Gemini 3 Flash 80% cheaper listed, 38% dearer in practice) is the paper most likely to settle this and is not yet ingested. Related evidence (2026-09-22), not an answer: September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 is the first population-scale reading of the decision rather than the cost — frontier-tier token share 52.5% to 44.7% in five weeks, firms reporting company-wide defaults away from frontier models because standard ones are "still highly performant and also more cost effective." It measures no tasks, turns or completions, so it cannot test the claim; it does establish that the buying population is applying the price-per-token rule this bullet's claim argues against, on a market where the blended effective price fell 41% over six months. - Does the advisor strategy's result (within 10% of the advisor's score at 63% of its price) generalize beyond SWE-bench Pro and the Sonnet-5/Fable-5 pairing — and where is the crossover at which advisor calls cost more than they save?
- Is "start with the strongest model" safe inside multi-role pipelines, given AgentOpt's finding that the strongest model was the worst planner? Anthropic's Sonnet-for-sub-agents note hints at a boundary it never states. Partially answered (2026-08-03): Cursor's production swarm says the safe form of the rule is positional — strongest model as planner, cheapest capable model as worker — and that running the strongest model in every role is the single most expensive way to reach the same grade. It also dissolves the apparent conflict with AgentOpt: Opus was the worst HotpotQA planner because it answered from parametric knowledge instead of delegating, and Cursor's architecture makes that impossible ("a planner never implements"). The failure is a property of harnesses that let a planner execute, not of strong models in the planner seat. Still unsettled: whether the ordering survives on tasks where the worker's job is judgment-heavy rather than instruction-following, which is the regime Cursor's own framing exempts.
Sources#
- The plunging price of thought — Emberson & Roodman (Epoch AI), "The plunging price of thought", 2026-09-22 (
empirical). Cited here for the per-token-versus-per-task critique of earlier price-trend estimates (Previous work) and the SWE-bench Verified row of Table 1. Full treatment on The Price of Fixed Capability - GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan et al. (19 authors, OpenAI), arXiv 2510.04374 v1, 2025-10-05, 29pp (
empirical). Cited here for A.2.4 Figure 12 (win rate by deliverable file extension, viewed at compile time) and §3.2 with Table 2 (naive versus try-1×/try-n× speed and cost ratios, re-verified againstpdftotext -layoutat compile time). Note the scope limit: cost estimates were obtained for OpenAI models only. Full treatment and COI on GDPval Benchmark - Claude models explained: choosing the best model for your use case — Anthropic, July 2026 (
vendor-claim): the start-smart default, class taxonomy, four selection questions, advisor strategy, saturation → custom evals - Gemini 3.5 Flash-Lite Model Card — Google DeepMind, 2026-07-21 (
vendor-claim): price rows inside the benchmark table, a +67% output-price rise across one generation of the efficiency tier, and no token or turn counts to price a task with - Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI — Latent Space, 2026-07-28 (
practitioner-opinion): OpenAI's opposite exposure philosophy — "32 options" collapsed onto a one-dimensional slider behind an opinionated default, with the same escalation trigger and the same absence of tokens-per-task - How is the Bun Rewrite in Rust Going? — Tom Lockwood, lockwood.dev, 2026-07-27 (
case-study, independent): the outside view of the Bun port's public artifacts, and the argument that a campaign's headline token cost omits CI, employee time, and the post-merge tail - The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI — Sayed Ali et al. (33 authors, all Writer, Inc.), arXiv 2607.06906, 2026-07-08 (
empirical, total vendor COI — Writer's harness, Writer's baseline, Writer's model in the panel, last author is co-founder/CTO): §6.2 and Table 4 for the per-model baseline→harness cost figures and the 36%-model-menu-versus-33–61%-harness comparison, Table 6 for the per-model capability means used in the quality-per-dollar column above, §7.1 for the "$/Mtok compares p; the bill is p × τ" framing. Table 2 is cell-collapsed and Table 7 row-shifted in the raw parse — neither is cited here; the model roster comes from §5.3 prose. Full treatment and parse warnings at Orchestration Sets Token Economics - The price is wrong: AI cost calculation has to consider task completion rates, not just token costs — Thomas Claburn, "The price is wrong," The Register, 2026-07-13 (
case-study, secondary reporting: a news article relaying Databricks' benchmark blog post and Matei Zaharia's social posts; the primary at databricks.com is not in the corpus). Supplies the $1.94 / $2.09 / $1.28 per-task figures, the 87% / 81% success rates, the "around 1.7x cheaper" per-token ratio, the SWE-Bench-is-tuned-for motivation, and the cited arXiv 2603.23971 result. 644 words, no tables. Ingest hazard worth remembering: WebFetch silently dropped the GLM 5.2 figure and the Gemini 3 Flash example, and both exist in the raw only because the body was rebuilt from curl'd HTML — the two figures most load-bearing against the vendor guidance are the two that nearly vanished - Fable's judgement — Simon Willison, "Fable's judgement," 2026-07-03 (
practitioner-opinion, 460 words): the self-delegated model-routing prompt, the auto-saved memory file's stated rationale (implementation to a cheaper model, judgment/review/synthesis in the main loop), and the unquantified outcome. The tip is second-hand from Jesse Vincent and the underlying judgement-over-rules advice second-hand from Cat Wu and Thariq Shihipar; no measurement of any kind - Agent swarms and the new model economics — Wilson Lin, cursor.com, 2026-07-20 (
case-study, vendor-authored): "Results across model mixes" and "Model economics" — the four planner/worker configurations, the $1,339–$10,565 total range at matched quality, the ≥69%-of-tokens/one-third-of-cost worker split, the $9,373 → $411 worker-spend comparison, and the Fable-versus-Opus planner inversion. Footnote 1 records that the solo Opus 4.8 and Fable 5 runs (hatched bars in the cost chart) were graded only informally, so no quality claim attaches to them - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim): the $0.05/minute front-end voice layer with separately-billed backend — a per-minute unit rather than a per-token one. A price list, not a measurement; no usage or cost-per-task data of any kind - Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google, 2026-09-15 (
vendor-claim): the Artificial Analysis cost-per-hour-of-input-audio chart, read from a labelled bar chart viewed during compile. A measured cost on a benchmark subset, not a price; the post publishes no pricing of its own - The state of AI in 2026: On the road to ROI — Dan Tinkoff, Lieven Van der Veken & Michael Chui with Tara Balakrishnan, The state of AI in 2026: On the road to ROI (McKinsey / QuantumBlack, 2026-08-25,
empirical, self-reported; online survey, 1,719 participants in 97 nations, fielded May 4 - June 8 2026, GDP-weighted). Cited here for Exhibit 8 (20% cost-constrained, by industry), Exhibit 14 (18% vs 6% coding-agent constraint by cohort), the 28%-of-ICT-budget figure and Chui's tokenomics commentary. It measures no tasks, turns or completions and therefore tests none of this page's claims. COI: McKinsey sells AI transformation consulting; the instrument is self-report from its own panel - September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 — Ara Kharazian, September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 (Ramp, 2026-09-09,
empirical): the prose supplies the 41% decline to $0.68 per million tokens, the $1.15 March peak, and 45% frontier token share; the weekly tier series and the per-lab daily price series are recovered Datawrapper datasets (no static chart images exist on the page). The per-lab divergence is this vault's arithmetic on those datasets, not Ramp's claim, and the price CSV's cached tail runs a few days behind the prose. COI: Ramp measures its own VC-forward-skewed corporate-card base; full instrument notes at Firm AI-Spend Intensity and Headcount Growth and Ramp - Running a Software Factory Efficiently at Uber Scale — Uday Kiran Medisetty, Uber Engineering blog, 2026-08-27 (
case-study, first-party): the six-term cost equation, the uReview Pareto bench (Figure 5, genericized models, hand-transcribed from the article's images — WebFetch cannot see images), the subagent-default-model lever, the tokens/request lever list (compaction cap, reasoning-effort default, MCP tool search/CLI resolution, code-mode's 5-query token table, SaaS MCP routing), the AI Context Graph grounding example, and the spend-visibility tooling. Not PDF-derived — no docling table-parse hazard; images were recovered from the page's base64-encodedimg srcattributes and viewed directly. Full per-source notes at Sources - HarnessTax: How Much Does the Harness Matter for Coding Agents? — Pan, Yang, Arabzadeh, Chiang, Stoica & Zaharia, Arena.ai blog, 2026-09-16/18 (
empirical): the SWE-bench Lite 21-combination cost/success grid and the Figure 3 first-call-context table. Full treatment, including the Terminal-Bench 2.0 capture caveat, at Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves - Introducing System One Models & Jev — Diogo Almeida, TypeSafe AI, 2026-09-15,
vendor-claim: the workflow-eval accuracy-vs-cost chart (fig1; log-scale values read off the image, approximate). Full treatment on Typed Decision Verifiers - Building Prod with Jev and LangGraph — Sydney Runkle & Hunter Lovell, LangChain blog, 2026-09-25,
vendor-claim(Jev integration partner): the "cheap by default, frontier on exception" framing (after Jaya Gupta) and Browserbase's Stagehandact()cascade — 0.7 confidence fallback, median 1.97 s → 0.46 s, reported secondhand with no fallback rate or task-success figure
Cited by 48
- Parallel Agent Orchestration×4
This is not the concurrency axis this page otherwise tracks — it is N kinds of agent, each with a…
- The Price of Fixed Capability×4
Cost Per Task Over Cost Per Token — the unit distinction this measurement depends on. Every earlier…
- Prompt-Cache Economics×4
Cost Per Task Over Cost Per Token — the missing token axis, measured: total billed spend reconciled…
- GDPval Benchmark×3
Cost Per Task Over Cost Per Token — vendor model-selection guidance cites the GDPval-AA derivative…
- Inference Efficiency as Capability×3
Cost Per Task Over Cost Per Token — the buyer-side statement of the last caveat: price per token is…
- Open Questions Backlog×3
Cost Per Task Over Cost Per Token (72d) — Does the advisor strategy's result (within 10% of the…
- Standardize the Infrastructure, Not the Tools×3
The confound the same edition supplies. Effective blended cost fell 41% to $0.68 per million tokens…
- Agent Quality Flywheel×2
Cost Per Task Over Cost Per Token — where the $27M → $1M projection lands: a percentage published…
- Build Instead of Buy Under Agentic Coding×2
Cost Per Task Over Cost Per Token — what the build decision is denominated in once it is made, and…
- Claude Fable 5×2
Anthropic's model-selection guide gives a rule that is explicitly not benchmark-driven (Cost Per…
- Claude Opus 4.8×2
Cost Per Task Over Cost Per Token — 4.8 is the first Anthropic model given a per-task price by an…
- Claude Sonnet 5×2
Cost Per Task Over Cost Per Token — Sonnet is the class Anthropic names for "high-volume sub-agents…
- Client-Side Agent Optimization×2
Cost Per Task Over Cost Per Token — the vendor-side counterpart, and a direct tension: Anthropic…
- Compute-Controlled Benchmarking×2
And it is still not compute control. Price per token is a rate; the number a buyer needs is rate ×…
- Context Lifecycle Management×2
Figure 6 shows the mechanic directly: a stable prefix-cache hit runs the length of the session, the…
- Cursor×2
Four model mixes at matched quality, 8× apart in cost — Cost Per Task Over Cost Per Token, Client…
- Dynamic Workflows: An Algebra for Agents×2
What survives as a genuine challenge is the cost accounting. The ~$165,000 is the API token cost of…
- GLM (Z.AI)×2
Cost Per Task Over Cost Per Token — GLM-5.2 is that page's counter-arm: the cheap end of the menu…
- Interaction / Background Model Split×2
Each half is metered separately, in a different unit. GPT-Live-1 sells at $0.05 per minute for "the…
- Knowledge-Centric Self-Improvement×2
Cost Per Task Over Cost Per Token — the paper reports dollars rather than tokens for exactly this…
- Live-Path Minimalism×2
Cost Per Task Over Cost Per Token — the productised split prices its two halves in different units:…
- Mythos Model×2
Anthropic's July 2026 selection guide states the packaging cleanly: the Mythos class "ships in two…
- Open Weights as Competitive Strategy×2
Cost Per Task Over Cost Per Token — the firm-level version of Ng's input-cost argument, with an…
- Orchestration Sets Token Economics×2
register databricks cost per task — Thomas Claburn, "The price is wrong," The Register, 2026-07-13,…
- Production-Sourced Evaluation×2
The method arriving from the fourth direction — not a benchmark vendor, not a product loop, but a…
- Tool-Output Pruning×2
Cost Per Task Over Cost Per Token — the clearest instance yet of the two axes disagreeing inside…
- Agent-Authored Harness Optimization
Cost Per Task Over Cost Per Token — the campaign's economics: ~$680 and ~1B tokens to move a…
- AI Product Economics Maturation
Cost Per Task Over Cost Per Token — the same billing-unit argument one layer down, at the model API…
- Artificial Analysis
Cost per hour of input audio · Interactivity Benchmarks, Cost Per Task Over Cost Per Token · a…
- Claude Code Best Practices
When does subagent overhead exceed the benefit of context isolation? Partially answered 2026-08-03…
- Crystallizing Agent Work into Workflows
Cost Per Task Over Cost Per Token — the cost frame this operationalizes and then escapes: the…
- Deterministic Engineering for Agent Code Review
Cost Per Task Over Cost Per Token — a rare instance where both terms move the same way at once, on…
- Evals as Product Spec
Cost Per Task Over Cost Per Token — where the spec becomes a procurement decision: Anthropic's…
- Firm AI-Spend Intensity and Headcount Growth
Cost Per Task Over Cost Per Token — where the index's price and token-tier series are read as a…
- Gemini 3.8 Live
Cost Per Task Over Cost Per Token — cost per hour of input audio as a billing denominator
- GPT-Live
Artificial Analysis cost per hour of input audio: $5.83, the most expensive system on the chart —…
- Harness Shrinkage as Models Improve
Willison's own test points the delegation at model routing: "For all coding tasks use your…
- Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves
Cost Per Task Over Cost Per Token — the cost-axis instance this page supplies: a fourth data point…
- Interactivity Benchmarks
Cost per hour of input audio, on a Big Bench Audio subset (Artificial Analysis) — Gemini 3.8 Live…
- Jev
Cost Per Task Over Cost Per Token — "cheap by default, frontier on exception": Jev as the cheap…
- Measuring Beyond Accuracy Saturation
Cost Per Task Over Cost Per Token — the vendor conceding the same point about its own products:…
- Agent Systems & Harness Engineering
Cost Per Task Over Cost Per Token — Anthropic's inverted model-selection default: start with the…
- Optimizer–Evaluator Decoupling
Cost Per Task Over Cost Per Token — the advisor strategy is this rule reached from the cost side…
- Ramp
Cost Per Task Over Cost Per Token — the blended token-price and model-tier series, read as a…
- Shared-Budget Compute Allocation
Cost Per Task Over Cost Per Token — the deployment version: batching several tasks under one budget…
- Shared Harness, Differentiated Surfaces
Cost Per Task Over Cost Per Token — the model-selection face of the same design tension: Anthropic…
- Typed Decision Verifiers
Cost Per Task Over Cost Per Token — cost per workflow as the denominator: the launch prices a whole…
- Unproductive Self-Verification
Cost Per Task Over Cost Per Token — the economic argument this failure mode inverts. "Stronger…
Related articles
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Agent-Authored Harness Optimization
An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Nine instances (Cli…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
