Sources#
Summary#
Hold a benchmark score fixed and ask what the cheapest available model costs to reach it, per question, on each date. That series is the price of fixed capability. Epoch AI's The plunging price of thought (Luke Emberson & David Roodman, 2026-09-22, empirical) is the corpus's first direct measurement of it. The headline:
- ~47% per quarter (~13× per year) on the five primary benchmarks since 2023: FrontierMath tiers 1–3, OTIS Mock AIME, GPQA Diamond, Chess Puzzles and Mystery Game Puzzles. The model-free estimate and the preferred model's estimate agree exactly, at 47.0%.
- By domain: math falls fastest at 50–52% per quarter (16–19×/yr); game puzzles fall slowest at 39–43% (7–10×/yr).
- Over time since a level first debuts as SOTA: 66% per quarter (75×/yr) at debut, and 32% per quarter (4.7×/yr) two years later.
- The worked example: OpenAI's o3 (2025-01-31) reached a guessing-adjusted 75% on GPQA Diamond at about $0.30 per question. GPT-5.6 Luna, released just under 18 months later, reached the same score for $0.0004. Epoch calls that a 725-fold drop. The "75%" is 75% of the way from chance to perfect, which is 81.25% raw.
Compared in log points, the authors say this is about 4× the rate of DNA sequencing (1.84×/yr, 2001–25), 6× compute (1.51×/yr, 1940–2001), 18× lithium batteries (1.16×/yr, 1991–2024) and 54× US residential electricity (1.05×/yr, 1892–1973).
Why this is a different number from "tokens got cheaper"#
Every earlier estimate the report cites measured the price per token of a model that clears a threshold, not the cost of clearing it:
- Appenzeller/a16z (2024): 10×/yr.
- Epoch's own March 2025 analysis: 9–900×/yr across six benchmarks.
- Demirer, Fradkin & Tadelis (2026), on the Artificial Analysis index.
Reasoning models broke that shortcut. They spend far more tokens to get good performance out of a smaller, cheaper-per-token model, so a per-token series no longer tracks what a given score costs. This is Cost-per-Task Over Cost-per-Token's distinction applied to a time series. The report uses the cost-per-task form, as do the two estimates closest to its method:
- Ihle (LessWrong, 2025-09), on coding tasks: costs halving every 1.4–2 months (64–380×/yr).
- Gundlach et al. (March 2026): 5–10×/yr, with prices falling faster at the high end.
How one run becomes a cost curve. Epoch borrows a procedure from CAISI's DeepSeek evaluation. Take a transcript from a run with a high budget or no budget. For any per-question token cap X, count the questions the model got right in fewer than X output tokens. Score the rest at the chance rate. Sweeping X gives a whole accuracy-versus-spend curve from one run, and Epoch does this at every reasoning-effort setting.
The obvious objection was tested. A model told its budget might do better than one silently cut off. For GPT-5.2, o4-mini, Claude Sonnet 4.6 and Gemini 3.5 Flash on GPQA (Figure 3), high-effort runs with an announced budget did improve at low budgets. They still did not beat the lower-effort settings under silent truncation. So truncation curves taken across effort levels cover most of a model's cost-performance range.
Two more data steps:
- Epoch filled the gap its own leaderboard history had left in low-effort and small-model runs.
- It ran open-weight models with no API on rented hardware. Validated against five that did have an API, the hardware cost came within 30% of API pricing.
What the number rests on#
- The frontier, not the average. The rate describes the cheapest model at each score, which assumes a user who switches to it immediately. The authors say plainly that no real user does. In Table 1, a regression over all runs instead of the frontier gives just 5.3–10.2% per quarter, with several negative signs. The authors dismiss this as "implausible." Read the other way, it says the average benchmarked model at a fixed score gets cheaper several times more slowly than the cheapest one, so the headline is a frontier rate and should be read as one.
- No standard errors, by choice. The frontier is an extremum, so one result can move it for months. The sampling is ad hoc, since Epoch mostly tests models that look near-frontier. Grid points are autocorrelated. The authors also argue that a resampling bootstrap cannot put a frontier outside the full-sample one. So they "abandon the formal paradigm of statistical inference." Treat the figures as measurements whose spread is shown by specification choices, not by intervals.
- The specification spread is wide.
- Changing only how the accuracy grid is averaged (Figure 7, GPQA) moves the bottom line from 42.9% to 58.0% per quarter (9.4× to 32.1× per year).
- The two dispreferred frontier fits require the curve to sit wholly outside the data. They give 27.2% and 38.1% for the primary average.
- The appendix's accuracy-side models give 53–68%.
- The authors pick the 47.0% combination and give their reasons, but a reader can't treat 13× as tighter than roughly 9–30×.
- "Since 2021" is an extrapolation. The data start in 2023. The claim that prices have fallen this fast since GPT-3's API opened (2021-11-18) rests on one MMLU pair: GPT-3 scored 43.9 at $60 per million tokens, and Llama 2-7B scored 45.3 at $0.20, about 31× per year. Both are non-reasoning models, so per-token and per-task are closer there. Figure 1's five-year AI line is the 2023–26 rate drawn backwards.
- Benchmaxxing. Mystery Game Puzzles keeps its game secret so nobody can train for it. It falls a little slower than the average: 44.0% versus 47.0% model-free, and 38.7% versus 47.0% in the preferred fit. The authors read this as the benchmaxxing critique being "valid but not fatal." That is the corpus's only measured estimate of how much benchmark-specific training inflates a cost trend: a few points per quarter. See Compute-Controlled Benchmarking.
- It chains a changing basket. Averaging across benchmarks, times and score levels is "effectively chaining price drops for various services." It does not track one price the way a kilowatt-hour does. The authors grant that comparing it to electricity is "apples and oranges" and still defend it as meaningful.
The shape: fastest right after SOTA#
Table 2's Box-Tidwell cost-frontier model gives the rate by quarters since a score first became SOTA. Averaged over the primary benchmarks, it runs 66.3% at debut, 51.8% two quarters on, 42.4% after four quarters and 32.4% after eight. The effect is concentrated in three benchmarks:
| Benchmark | At SOTA | After 8 quarters |
|---|---|---|
| AIME | 83.3% | 18.3% |
| GPQA | 80.7% | 24.6% |
| FrontierMath T1–3 | 63.6% | 29.8% |
The two game benchmarks go slightly the other way (Chess 41.8% → 45.6%, Mystery Game 35.9% → 40.0%). All six secondary benchmarks decay.
The authors' explanation is a short-lived premium: the lab that first reaches a score can charge for it until competitors catch up. They conclude that "supra-normal profits in LLM service provision may be fleeting for any given model," which bears on AI Product Economics Maturation's margin projections. They also point out that the competition near SOTA is between closed-weight models. In the report's second figure, only one of eight traces has an open-weight model as the second point (QwQ-32B at 25% on AIME). They allow that open models may still be pushing from behind. That is a different place to look for open-weight pressure from The Open-Weight Frontier Gap's Elo margin.
Where it is slow: agentic coding#
Among the secondary benchmarks, SWE-bench Verified falls slowest: 27.5% per quarter model-free, or about 3.6× per year. Its rate at SOTA is 32.9%, lower than any primary benchmark's. FrontierMath tier 4 (2025-07) is the other slow one at 26.0%. The report doesn't comment on this. The benchmark closest to what this corpus mostly cares about, agents fixing real repositories, has the price that falls slowest.
Two readings fit what's in the corpus, and nothing here separates them:
- Agentic tasks spend more tokens per task as models improve, which is the premise of Cost-per-Task Over Cost-per-Token.
- SWE-bench Verified's score ceiling is distorted by the defects and leakage recorded on Benchmark Task Defects (Spec–Test Mismatch) and Evaluation-Time Answer Leakage.
Supersession: "10–100× per release cycle"#
Latent Capability Overhang carried Noam Brown's statement (practitioner-opinion) that "the cost of any given capability drops 10–100× with each model release cycle" of two to three months. At Epoch's measured rates, a 2–3-month cycle buys about 1.5–1.9× on average, and about 2–3× for a level that has just become SOTA. The steepest rate anywhere in the report, FrontierMath tier 4 v2 at SOTA (85.3% per quarter), is about 7× per quarter. The single best example, 725× over 18 months, works out to about 3× per quarter.
So the per-cycle figure is too high by roughly an order of magnitude. The direction and the conclusion Brown draws from it survive: waiting is cheaper than extracting a capability now. The size of the saving does not survive. The superseded text is kept on that page.
What it does to budget-denominated scores#
A score stated in dollars loses comparability over time. FrontierMath Erdős Benchmark's $300 cap is the corpus's clearest case. The worry that "3% at $300" in September 2026 won't mean the same thing a year later, raised on Compute-Controlled Benchmarking, now has a rate: at FrontierMath tiers 1–3's 53% per quarter, a fixed score costs about 20× less a year on (0.469⁴ ≈ 1/20), so the same $300 buys a very different level of capability. Expenditure Horizon's dollar axis has the same issue for the agent curve, though its human curve deflates at wage rates. A dollar-axis score should therefore carry its price date, just as Economic Benchmark Construct Validity argues a leaderboard row should carry its release date.
Frontier rate versus realized price#
The frontier rate is much faster than what buyers see. Across Ramp's customers, the realized blended price per token fell about 41% over six months to September 2026, roughly 3× a year (empirical; see Cost-per-Task Over Cost-per-Token). ICONIQ's blended token index ended August 2026 roughly where it started in December 2025 (AI Product Economics Maturation). These don't contradict Epoch. They measure a spend-weighted mix that has been shifting toward and away from frontier tiers, and Epoch's third caveat, that no user stays on the cost frontier, predicts a gap in exactly this direction. What nobody has measured is how much of the 13× reaches a buyer who doesn't switch models every quarter. Market-Priced AI Exposure (the AI Premium)'s premium on frontier, paid-for usage is priced at the part of the curve where this report finds prices fall fastest.
Connections#
- Cost-per-Task Over Cost-per-Token — the unit distinction this measurement depends on. Every earlier price-trend estimate was per token, and reasoning models are why that stopped being enough
- Compute-Controlled Benchmarking — the budget-disclosure argument this rate turns into a dated requirement. A dollar budget loses about half its meaning each quarter at fixed performance, and the secret-game benchmark gives the corpus's only measured size for the benchmaxxing effect on a cost trend
- Latent Capability Overhang — the "10–100× per release" premise, now measured at about 1.5–3× per cycle and marked superseded there
- Open-Weight Elicitation Irreversibility — repeats that per-generation figure in its audit question, and carries the same correction
- Expenditure Horizon — a dollar-denominated horizon whose agent curve moves left over time at about this rate
- FrontierMath Erdős Benchmark — the $300-per-attempt protocol on the same family as the fastest-falling primary benchmark (FrontierMath T1–3, 53.1% per quarter)
- Inference Efficiency as Capability — the levers (KV-cache cuts, quantization, MoE, drafters) that make up part of this aggregate rate. This report measures the total and doesn't break it down
- The Open-Weight Frontier Gap — a different reading of open-weight pressure: near SOTA, price competition is closed against closed
- AI Product Economics Maturation — survey-projected margin growth against a measurement that says a SOTA premium lasts only a few quarters
- Market-Priced AI Exposure (the AI Premium) — a market premium on frontier usage, where prices fall fastest
- Economic Benchmark Construct Validity — release date as the forgotten control on a capability comparison. The price date is its twin for a cost comparison
- Benchmark Task Defects (Spec–Test Mismatch) — one possible reason SWE-bench Verified's cost trend is the slowest measured
- Epoch AI — the authors, with data and code at github.com/droodman/inference-cost
- US Center for AI Standards and Innovation (CAISI) — source of the transcript-truncation method that turns one run into a cost curve
- Artificial Analysis — the index behind an earlier per-token estimate. Epoch can't use it because it publishes no transcripts
Open Questions#
- How much of the frontier rate reaches a buyer who doesn't switch? Epoch measures ~13×/yr for the cheapest model at each score. Ramp's realized blended per-token price fell about 3×/yr, and that mixes tier shifts with list-price cuts. Falsifiable: keep a fixed agentic workload and a buyer panel's actual model choices, record cost per task every quarter, and compare it with the frontier rate for the same score.
- Is the "fastest at SOTA" decay a durable pattern or an artifact of the fit? It shows up on three of five primary benchmarks. The game benchmarks run the other way. The time curvature often failed to estimate and was then forced to linear. There are no intervals. Falsifiable when Epoch updates the dataset: score levels that first become SOTA in 2026 should fall about 66% per quarter in their first quarter and about half that rate two years on. Trigger: the next data release from github.com/droodman/inference-cost.
- Why is SWE-bench Verified the slowest-falling benchmark (27.5% per quarter)? Is it rising tokens per task on agentic work, a distorted ceiling from defective and leaky tasks, or slower competition in coding? Falsifiable: rerun the same truncation analysis on SWE-Bench Pro Verified's repaired, isolated task set and see whether the rate moves toward the ~47% average.
Sources#
- The plunging price of thought — Luke Emberson & David Roodman (Epoch AI), "The plunging price of thought", epoch.ai, 2026-09-22,
empirical. A long web report with a technical appendix. Data and code are at github.com/droodman/inference-cost, with an interactive overlay at droodman.github.io/inference-cost. Sources used here: the key takeaways, the o3 → GPT-5.6 Luna example (footnote 1 for guessing adjustment), Previous work, Data (the CAISI method, the announced-budget test, the rented-hardware validation within 30%), the modelling preferences, Table 1 (all columns), Table 2's by-quarter decline columns (the parameter columns aren't cited), the Limitations section including Figure 7's four grid variants (47.0 / 42.9 / 58.0 / 47.7%, from alt text and prose), and the Discussion (MMLU extrapolation, cross-technology rates, closed-vs-open competition near SOTA). Figure warning: the values in Figures 1–2 were read approximately off chart images at ingest and aren't cited here. Figures 3–7 survive only as alt text. Every number above comes from prose, captions or the two HTML tables, whose rows match the prose totals (the primary averages of 47.0% and 66.3% → 32.4%). The report deliberately gives no standard errors.
Cited by 17
- Compute-Controlled Benchmarking×5
Compute has several units (tokens, dollars, wall-clock). They diverge (a more efficient model wins…
- Latent Capability Overhang×4
The overhang persists because of a rational disincentive. ~~The cost of any given capability drops…
- Cost-per-Task Over Cost-per-Token×3
Price Of Fixed Capability — this page's unit distinction applied to a price trend: once reasoning…
- Epoch AI×3
A price trend for fixed capability (2026-09). The plunging price of thought (Luke Emberson & David…
- Expenditure Horizon×3
Price Of Fixed Capability — why a dollar horizon needs a price date: at fixed performance the agent…
- Market-Priced AI Exposure (the AI Premium)×3
Price Of Fixed Capability — the price side of the frontier margin this premium concentrates on: the…
- US Center for AI Standards and Innovation (CAISI)×2
A budget-truncation method others reuse. CAISI's evaluation of DeepSeek models (2025-09-30, NIST,…
- FrontierMath Erdős Benchmark×2
Price Of Fixed Capability — how fast the $300 cap loses its meaning: Epoch's own price-trend report…
- Open Questions Backlog×2
Price Of Fixed Capability: Is the "fastest at SOTA" decay a durable pattern or an artifact of the…
- Open-Weight Elicitation Irreversibility×2
The Brown compile left an open question in the log: who audits released models for latent dangerous…
- The Open-Weight Frontier Gap×2
the plunging price of thought — Emberson & Roodman (Epoch AI), "The plunging price of thought",…
- AI Product Economics Maturation
Price Of Fixed Capability — a measured bound on how long a capability premium lasts: the cost of a…
- Benchmark Task Defects (Spec–Test Mismatch)
Price Of Fixed Capability — SWE-bench Verified has the slowest fixed-performance cost decline in…
- Economic Benchmark Construct Validity
Price Of Fixed Capability — the cost-side twin of this page's release-date control: a…
- Evaluation-Time Answer Leakage
Price Of Fixed Capability — leakage on SWE-bench-family tasks is one candidate reason SWE-bench…
- Inference Efficiency as Capability
Price Of Fixed Capability — the aggregate this page's levers feed into: the cheapest cost of a…
- Evals & Benchmarks
Price Of Fixed Capability — How fast the cheapest way to reach a fixed benchmark score gets…
Related articles
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- Benchmark Contamination and Decontamination
Sun, Zhan & Gales (Cambridge): per-sample distribution distances expose that aggregate-accuracy decontamination can wor…
- Google DeepMind
Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…
