Sources#
- Announcing FrontierMath Erdős
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
- Gemini 3.5 Flash-Lite Model Card
- Gemma 4 Technical Report
- HarnessTax: How Much Does the Harness Matter for Coding Agents?
- How UK AISI and EvalEval Are Making Benchmark Results Reproducible
- Inkling: Our Open-Weights Model
- Kimi K3 Model Card
- More compute, more capability: Why AI agent evaluations need to account for test-time compute
- Noam Brown – Agent swarms, alignment, & recursive self-improvement
- OEIS Open: How many conjectures can language models turn into theorems?
- On the Navier–Stokes Millennium Prize Problem
- One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
- Recursive Self Improvement for Coding Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- Shortcutting the Fix: Identifying and Categorizing Agentic Exploits in Software Engineering Benchmarks
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- The plunging price of thought
- UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
Summary#
The evaluation consequence of Large-Scale Test-Time Compute: if capability is a function of inference budget, then a benchmark number reported without its budget is meaningless. Noam Brown's essay is aimed squarely at the "benchmark grid" — the standard release artifact with benchmarks on the x-axis, models on the y-axis, and a single score per cell. His fix: put compute on the x-axis. Report performance as a function of tokens, cost, or time; or fix a budget and compare within it (practitioner-opinion).
The motivating incident: when GPT-5.5 launched, the grid showed only a few-percentage-point gain over GPT-5.4, and the initial reaction was skepticism that it was meaningfully better. It was — 5.5 was simply far more compute-efficient, reaching the same or better answers with much less thinking. 5.4 at max settings thinks longer per response. Control for thinking time and 5.5 is "a substantial jump," which matched users' day-to-day experience once they played with it. The grid hid the improvement because it didn't equalize compute.
Benchmark-maxxing: gaming the score with scaffolds#
The sharper worry is that the grid is trivially inflatable. Benchmark-maxxing is Brown's term for scaffolding tricks that raise a benchmark number without a real capability gain once compute is held equal: run the model five times and take the best answer; add an LLM judge to pick the strongest of N candidates. These "look a lot better on paper but are actually not better once you control for the amount of test-time compute." It is Goodhart's law moved out of the training loop and into the eval-reporting layer: the measure (grid score) becomes the target, so it stops measuring capability and starts measuring willingness to spend inference on best-of-N.
The standing defense against optimizing to a benchmark (as opposed to for the underlying skill) is a held-out private set that isn't publicly available — Brown says OpenAI tries not to optimize for specific benchmarks, but "once you put out a benchmark it's always at risk of just being optimized for."
The bad equilibrium#
Brown frames the grid's persistence as a coordination failure, not a disagreement. Privately, researchers agree the x-axis should be cost/tokens/time — "yeah, that makes sense, we should do that." But the published answer is "people expect us to publish the grid," and people expect it because everybody publishes the grid. Everyone knows it's a bad equilibrium and nobody wants to move first. Writing the essay is an explicit attempt to give the field permission to defect: so that "next time there's a model release, a company can feel comfortable not publishing the grid, at least not at the top line." This is Goodhart named as a collective-action problem rather than an individual temptation.
Independent adoption: UK AISI reports curves, not scores (July 2026)#
Brown wrote the critique; the UK AI Security Institute operationalized it, and as a government evaluator rather than a competing lab it is the cleanest party to defect from the grid. Its July 2026 study is empirical and its conclusion is Brown's prescription almost verbatim: "Evaluations should report capability curves, especially when performance may still be rising" — and "agent capability cannot be interpreted without the compute budget used to estimate it." A capped score, AISI warns, can "make model comparisons unequal" and "quietly mislead" a decision-maker into treating an under-resourced evaluation as evidence of a low-capability model.
The prescription is now AISI's practice, not just its recommendation. It evaluates frontier models across multiple budgets (including very large ones for the hardest tasks), reports reliability and reach against budget, and is defining "minimum informative budgets" — a budget declared sufficient only once a model's reach stops rising with more compute. That last is the operational answer to "how much is enough?" that a single grid number never had.
And the same evaluator's disclosure floor, three weeks later (2026-07-23)#
The section above is AISI at its best on the budget axis. Its joint Kimi K3 assessment with
CAISI (UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities, empirical) is the same institution at its
worst on two others, and both failures are ones no vendor grid in this page commits:
- The comparator is anonymous. Every US figure — a 76.2% ExploitBench ladder score, 20 of 41 arbitrary-code-execution solves, step 28.5 of a 32-step cyber range — is attributed to "the most cyber-capable U.S. models," a group never enumerated anywhere in the document. Figure 2 makes the asymmetry graphic: ten PRC-lab models individually named and plotted against an unlabeled US aggregate trendline. Vendors cherry-pick which rival to show; this shows none, which is worse for the same reason a grid without budgets is worse than one with them — nothing in it is checkable. Nobody can reconcile 76.2% against any published card.
- The budget is a bare number. "Within the 100M-token limit" and "the standard 100M token limit" are the only statements of spend, with no per-task or per-attempt qualifier, from the organization that established that a capped score "cannot be interpreted without the compute budget used to estimate it." The corpus's other AISI-sourced 100M figure (Opus 5 solving the same range 8 of 10) is per-attempt, which is the likely reading, but the document does not say.
Worth recording precisely because the counter-argument is available and does not fully answer: a safety publication has a disclosure trade-off a benchmark table does not, and naming which US models were de-safeguarded is closer to publishing a capability roadmap. That explains the anonymity of the comparator; it does not explain the anonymity of the budget unit.
Standardized run records, and where "verified" stops helping (2026-09-22)#
The two AISI sections above are single-paper disclosure — a good curve, then an anonymous
comparator. Its next release changes the unit: jointly with the EvalEval Coalition
(How UK AISI and EvalEval Are Making Benchmark Results Reproducible, empirical, HF blog), AISI is publishing
verified Evaluation Cards — standardized benchmark+run+model-config records under EvalEval's
Every Eval Ever (EEE) schema — for five benchmarks (HealthBench, FrontierMath, Humanity's Last
Exam, SWE-Bench Pro, Terminal-Bench 2.0) across six models (Claude Opus 4, 4.5, 4.6; GPT-5, 5.2,
5.4), plus two cyber evals on a partially overlapping model set. The cards accompany a new AISI
paper, How Inference Compute Shapes Frontier LLM Evaluation — the same test-time-compute-curve
research line as the July 2026 study above, now packaged as machine-readable run records rather
than only a blog post, and explicitly built on AISI's own prior efficiency and rigor tooling
(OptStop, HiBayES).
One dataset from the release is a direct instance of a curve this page already tracks, with an explicit non-comparability warning attached. The HLE trajectory chart plots cumulative tasks solved (of a 100-task sample) against lowest observed tokens to success, under an expanded-budget, oracle-feedback protocol — the model gets correctness feedback after each attempt and keeps retrying as the budget rises: Claude Opus 4 and Opus 4.6 both reach 100/100, GPT-5.4 reaches 95, GPT-5 88, GPT-5.2 71, Claude Opus 4.5 70. The source is explicit that these are not standard fixed-budget HLE accuracy scores and should not be compared with published HLE numbers — the protocol's oracle feedback and expanded budget are themselves the variable under study, not a replacement scoring convention.
The Terminal-Bench 2.0 release is the sharper test of what "standardized record" buys. Its chart plots the study's own scores (diamonds, split by no-feedback vs. with-oracle setup, SE whiskers) against circles pulled from EEE for the same five models (Claude Opus 4.5/4.6, GPT-5/5.2/5.4 — "no comparable entries elsewhere for Claude Opus 4") — exactly the cross-source check this page's open question below asks for. But the chart itself is rendered client-side and exposes only its axis range (30.0%–80.0% binary success) and legend to a reader who does not script the SVG/canvas; no per-model numeric value is recoverable from the published page. "Verified" here means the record's provenance is checkable — a named benchmark version, run setup, and model config — not that its values are legible without tooling, which is a narrower claim than the run-record audits below.
Worked example: Gemma 4's headline table (July 2026)#
Brown published the critique in June 2026. Three weeks later Gemma 4 shipped the artifact, and it is worth naming because the paper is otherwise careful (empirical, sixteen tables).
Table 5 — the report's most-cited table, and the basis of its "leap in performance" claim — compares Gemma 4 in thinking mode against Gemma 3 27B non-thinking. AIME 2026 goes 20.8 → 89.2; Codeforces Elo goes 110 → 2150. No thinking-versus-non-thinking ablation for the same Gemma 4 model appears anywhere in the document. The generational improvement and the inference-budget increase are therefore confounded in the headline number, and the report never states the thinking budget it allowed.
What makes the case instructive rather than merely damning is that the same report does control, elsewhere, without remarking on it:
- Table 9 (long context) is run "without thinking" on both sides. It is the one clean generational comparison in the paper — and its gains are real and large (LOFT Text Retrieval @128k: 8.6 → 79.5).
- Table 6 (vision) repeats the confound: Gemma 4 thinking versus Gemma 3 27B non-thinking with Pan & Scan.
So the fix is not beyond the authors; it is applied inconsistently and never flagged. This is what Brown means by a bad equilibrium rather than a disagreement — the compute-controlled table exists inside the same PDF as the uncontrolled one, because nobody expects the top-line table to control and everybody expects the long-context table to.
There is a second-order cost specific to this paper. Gemma 4's actual contribution is inference efficiency — a 37.5% smaller KV cache, sub-gigabyte quantization, a drafter head. Those are precisely the gains a compute-uncontrolled grid cannot display, exactly as the grid hid GPT-5.5's efficiency advantage over 5.4. By publishing the standard grid, the report obscures its own best result.
A vendor half-defects: Inkling publishes its own curve (July 2026)#
Inkling's release is the first vendor artifact in this corpus to do what Brown asked for — partially. TML sweeps its controllable effort dial from 0.2 to 0.99 and publishes performance against mean generated tokens on Terminal Bench, HLE, and IFBench: capability curves, exactly AISI's prescription, and the framing is Brown's ("looking at the full cost curve allows developers to choose the best model for each use case"). The curve shows what a grid cell can't: Inkling matches Nemotron 3 Ultra on Terminal Bench at roughly a third of the tokens.
The defection is half because it is asymmetric: competing models appear at their default operating points — single dots against Inkling's curve — and the main benchmark table is still a standard grid at effort=0.99. A curve for me, points for thee is progress on the bad equilibrium, but it also flatters the model with the dial: the comparison Brown's critique actually demands (every model swept over its own budget axis) remains unpublished, now with the additional wrinkle that the effort dial is a trained capability (Inkling) rather than a serving parameter, so the sweep is cheap only for the vendor who trained it in (vendor-claim).
Harness disclosure is not compute control: Kimi K3's footnotes (July 2026)#
Kimi K3's model card (vendor-claim) is the most granular evaluation disclosure in this corpus and the sharpest illustration that granularity and control are different axes. Its 45-row grid is a standard grid. What sits underneath it is not standard at all.
What it discloses, in a class of its own:
- Every score is pinned to a named harness. Not "DeepSWE: 67.5" but "67.5 with the Kimi Code harness, 67.3 with mini-SWE-agent, on the DeepSWE v1.1 task set." Per-model harness pairings are listed benchmark by benchmark — Claude Code for the Claude models, Codex for the OpenAI models, Kimi Code for K3 — because the harness is now a first-class variable in the score, exactly as model-versus-scaffold contribution argues.
- Reasoning effort is named per model. All at max, except GPT-5.5 at "xhigh"; K3 at temperature 1.0 with top-p 0.95 for single-step and 1.0 for agentic tasks.
- Grader-side confounds are disclosed against the vendor's own interest. Fable 5 "hit fallbacks on 35% of the tasks" in the SWE-Marathon evaluation, "which may have negatively impacted its measured performance." On Kimi Code Bench 2.0, Fable 5 hit 13 fallbacks and 1 refusal of 80 tasks; 10 of 80 entered GPT-5.6 Sol's cyber guard; GPT-5.5 refused 3. On Agents' Last Exam, the leaderboard's Fable 5 entry has 40% of tasks "annotated as downgraded." A vendor volunteering that its rivals' numbers were depressed by the rivals' safety machinery is the opposite of benchmark-maxxing.
- Benchmark modifications are stated. SWE-Marathon was run on an H20-recalibrated branch of the pre-v1.1 tasks (Docker images, performance gates and reference oracles recalibrated for H20; correctness and anti-cheat validators untouched). PostTrainBench was re-run on H20 rather than the official H100.
- A self-deflating number is reported. On its own in-house Kimi Code Bench 2.0, K3 scores 73.7 under the rival Claude Code harness and 72.9 under Kimi Code — and the card prints 72.9 in the table. Likewise BrowseComp: 91.2 with a 300K-token compaction strategy, disclosed alongside 90.4 with the full 1M window and no context management.
What it does not do — and this is the point. Nowhere is there a token count, a dollar cost, a wall-clock figure, or a curve. "Reasoning effort = max" is a setting, not a budget: it says all models were told to think hard, not that they spent comparable compute doing so, and it is precisely the reading of "max" that Brown's GPT-5.5-versus-5.4 anecdote shows to be uninformative. The comparison is also structurally asymmetric — K3 is evaluated on Moonshot's own Kimi Code harness while competitors are quoted at "the best score across harnesses" or lifted from leaderboards run by third parties at unknown budgets. Where Inkling published a curve for itself and points for everyone else, Moonshot publishes provenance for everyone and a home-field harness for itself.
There is a second finding here that belongs to safety rather than evaluation: those fallback fractions are the first third-party measurements of how often Fable 5's classifier fallback fires on ordinary technical work.
The disclosure gap, measured: harness as an undisclosed budget surface (Arena.ai, 2026-09)#
Kimi K3's card names the harness per score but publishes no budget for it. HarnessTax (HarnessTax: How Much Does the Harness Matter for Coding Agents?, empirical) measures what that omission is worth: holding the model fixed and varying only the harness (Claude Code, Codex CLI, Pi) across 21 combinations on SWE-bench Lite, the same model at the same success rate costs up to 2× more depending on which harness ran it — Claude Fable 5 at 97.8% (Claude Code, $1.329) vs. 96.7% (Pi, $0.666). A leaderboard that names the harness but not its cost — exactly K3's disclosure shape — is silent on a variable that moves the bill 2–5× with success held within ±2–5%. This is the harness-choice instance of the page's thesis: the grid controls for the model and the task, rarely for the scaffold's own budget, and the scaffold's budget turns out to be large.
A price axis inside the grid: Gemini 3.5 Flash-Lite (July 2026)#
DeepMind's Gemini 3.5 Flash-Lite card (2026-07-21, vendor-claim) does something none of the others do: the first two rows of the benchmark table are prices — input and output $/1M tokens — printed for every model in the comparison, the vendor's own and its rivals' alike. Not a footnote, not an appendix, not a separate pricing page: a cost row sitting above SWE-Bench Pro in the same grid a buyer reads.
That is the first symmetric cost disclosure in this corpus. Inkling published a curve for itself and points for everyone else; Moonshot published provenance for everyone and a home-field harness for itself; DeepMind publishes the same cost figure for all four columns and gives itself no special treatment on that axis.
And it is still not compute control. Price per token is a rate; the number a buyer needs is rate × tokens-per-task. The card reports no token counts, no wall-clock, no thinking budget, and no effort setting for any of the four models — so the grid now displays a cost column that cannot be turned into a cost. The gap is not academic: it is exactly the quantity Brown's GPT-5.5-versus-5.4 anecdote turns on. A model priced at half the rate that emits three times the tokens is more expensive, and this table would show it as cheaper. Adding price without tokens can therefore make the grid more confidently misleading than a grid with no cost information at all, because it invites an arithmetic the data does not support (Cost-per-Task Over Cost-per-Token).
A second defect is orthogonal to compute and worth separating out: comparator selection. DeepMind's July 2026 model is set against GPT-5.4 mini and Claude Haiku 4.5 while its own comparator is its immediate predecessor. This corpus has been tracking GPT-5.5, GPT-5.6 Sol and the Claude 4.8/5 generation for weeks; the card does not date its rivals, state whether newer efficiency-tier models from either lab exist, or say when the rivals' scores were collected. Compute control and comparator currency are independent axes of grid honesty, and a vendor can defect on one while quietly failing the other.
So the corpus now has four vendor half-defections with four shapes: Inkling gave the curve without the competitors, Kimi gave the competitors' conditions without the curve, DeepMind's Gemini line gave everyone's price without anyone's token count, and Gemma 4 gave none of it at the top line while controlling quietly further down. None of the four did what Brown asked. The bad equilibrium is not holding because vendors refuse to disclose — Moonshot discloses more than AISI would need, and DeepMind puts dollars in the headline table — it is holding because the axis nobody puts on the page is the one that costs money to sweep. Prices are free to publish; tokens-per-task must be measured.
Picking the axis: METR says dollars, and measures why (July 2026)#
Every artifact above answers "should there be an axis?"; METR's expenditure-horizon note (2026-07-21, empirical) is the first in this corpus to argue which axis and back the choice with a measurement. Its answer is dollars, and the reason is not accounting taste:
- On an agentic AI R&D task, tokens are the minority of the bill. Across six agent runs on the NanoGPT speedrun at up to $10,000 each, experiment compute was ~70–90% of the cost of most trajectories — GPU time spent running and validating candidate training recipes, not model inference. A tokens-only x-axis on this task would be plotting against the smaller 10–30% of what was actually spent, and would rank a token-frugal agent that runs many experiments above a token-heavy one that runs few. Dollars are the only unit that absorbs both.
- Dollars are the unit the decision is actually made in. "Returns to expenditure directly measure the economically-relevant variables for an AI R&D lab that is choosing between spending on human vs agentic labor." Tokens and wall-clock cannot be compared to a salary; dollars can, which is what lets METR put a human curve on the same axis.
That last point is the substantive extension, and it is what none of the four half-defections above attempt: METR plots a human returns-to-expenditure line on the same axis as the agent curve and reports where they cross. Brown asked for the budget to be named; AISI reports reach and reliability against budget; METR adds a second curve to compare against, so the x-axis carries a unit of human labour rather than only a unit of spend. On NanoGPT that calibration is ~$2,500 per 1% speedup, and the six agent runs land at $0–$3,300.
Two caveats this page should carry, because they are the cost of the choice:
- The dollar axis absorbs harness inefficiency. METR's agents had continuous access to 4 H100 nodes and ran experiments freely, which is where the 70–90% came from; METR calls its own harness "likely inefficient" and expects an optimized one to lower cost for the same optimization. So a dollar-denominated curve measures agent-plus-harness, and comparing two labs' curves compares their harnesses too. METR's defence — that shifting the curves horizontally "would not dramatically change the expenditure horizon" — is read off the curves' flatness at the right end, not tested against a cheaper harness.
- Denominating in money imports a price. The $2,500/1% human rate assumes a $150/hour wage, and METR's own sensitivity table shows the resulting horizon moving from $120 to $14,400 for one model across a 10× range of the human-cost assumption. The dollar axis is honest about what was spent and inherits an assumption about what labour is worth — an exposure the token axis does not have.
Routing and consensus are subject to the same budget question#
The critique generalizes to the vendors whose value proposition is a scaffold. A routing / consensus layer (send each sub-task to the right model, or aggregate several models' answers) can indeed beat any individual model on a benchmark. But Brown's principle collapses the apparent win into the same question: does the routed ensemble beat the same model simply thinking longer at equal test-time compute? Consensus-among-models is just another way to spend inference; unless it's compared at a fixed budget — and shown to hold on real use cases, not just the benchmarks it was tuned on — the improvement may be an artifact of spending more, or of over-fitting the eval.
Benchmark-maxxing, measured: harness evolution at a matched budget (July 2026)#
Brown's argument that scaffolding tricks inflate scores was an argument. Wang, Zhu, Hu et al. (arXiv 2607.12227, Ai2 / UW, 2026-07-14, empirical) run it as an experiment against a named method class, and it is the first time this critique appears in the corpus with data attached.
The target is automatic harness evolution — an agent that iteratively rewrites its own scaffold using benchmark feedback and reports the final score on that same benchmark (Agent-Authored Harness Optimization). The observation that makes it a compute-control problem: harness evolution is itself a search procedure that repeatedly evaluates and revises candidates against task feedback. So it belongs on the same axis as parallel sampling and sequential refinement, and the honest comparison holds feedback and inference budget matched across all three. Nobody had done that.
Fix K = 5 for every method on Terminal-Bench 2.1 across three frontier models, and the ordering inverts the published one: parallel sampling averages 72.3 pass@1 without unit-test feedback and 86.0 with, against harness evolution's 67.4 and 75.8 — the evolution arm finishing below the do-nothing baseline (68.2) in the first setting. Evolving a harness on held-out tasks transfers +0.6pp.
Two things generalize beyond harness evolution:
- The pass@1-versus-pass@k split is a portable diagnostic for benchmark-maxxing. Harness evolution's pass@5 (86.2) sits level with parallel sampling's flat 86.0 while its pass@1 (75.8) trails by ten points. The authors' inference: "if harness revision genuinely produced better harnesses, we would expect the improvement to be reflected in pass@1. Instead, the benefit only materializes when we can select among multiple trajectories." A gain that appears only under best-of-N selection is best-of-N, whatever the paper calls it — which is exactly Brown's charge, now with a test that separates the two.
- Search-set and evaluation-set overlap is a second, independent confound. When the tasks a method optimizes against are the tasks it reports on, gains "may reflect adaptation to task-specific patterns." That is Benchmark Contamination and Decontamination's leakage argument arriving through the scaffold rather than through training data, and the remedy is the same held-out private set Brown prescribes — applied to the optimization loop, not just to the model.
The result cuts both ways for this page's thesis, and the paper says so. Its §5.2 concedes Terminal-Bench may be the wrong instrument — agents already score highly, and "a minimal setup consisting of a shell tool and a basic prompt already suffices for most solvable tasks," so there is little for scaffolding to buy. A budget-matched comparison run on a benchmark with no headroom can be uninformative in the other direction, which is a caveat the compute-control literature has not had to state before: controlling the budget fixes one confound and does not manufacture sensitivity where the benchmark has none.
The purest specimen: an undefined effort tier, called a control (DarwinX, July 2026)#
Seventeen days after the paper above, DarwinX (arXiv 2608.07545, Salesforce AI Research, 2026-07-31, empirical) reports the opposite result on the same benchmark and does what the Kimi section above warns about — with one aggravating difference. Moonshot published effort settings and never claimed they were budgets. DarwinX publishes tier names, titles a section "the gain is the harness, not compute," and calls the resulting comparison "an effort-controlled comparison." It is the first artifact in this corpus where a vendor-tier label is explicitly asserted to be the compute control, which makes it the cleanest available specimen of the confusion this page exists to name.
What full-text verification finds:
- The tiers are never defined. "medium," "high" and "xhigh" appear throughout the results table and are given no token count, no turn count, and no wall-clock bound anywhere in 33 pages. Appendix B's per-benchmark protocol table — base model, evolution data, report data, selection signal, report metric — has no compute column at all. No dollar figure appears in the paper.
- So the "control" compares tier names across vendors. "A neutral harness at higher effort (Terminus-2, xhigh) reaches only 78.0%, below DarwinX's 83.2% at high effort, so raw effort in another harness does not reproduce the gain." That is a real statement about escalating a tier inside a different harness. It is not a normalized budget, because nothing establishes that one harness's
xhighand another'shighbuy comparable compute — which is exactly Brown's GPT-5.5-versus-5.4 point. - The paper's own headline pair is not tier-matched. The matched-model gain it calls load-bearing runs
GPT-5.5 / default(75.5%) againstGPT-5.5 / high(83.2%). The single genuinely tier-matched row pair in the table is DarwinX's 84.7% against OpenAI's own reference 81.8%, bothGPT-5.6 Sol / medium— +2.9 points, its smallest headline. - It concedes the gap in §9 ("the archive, parent selector, recombination operator, and inference effort are not independently randomized") and, to its credit, states in §3.2 that "public leaderboard rows use different models and effort settings, so they provide context rather than controlled comparison." The paper knows; the abstract and §4.1 title do not carry the caveat.
- The only real compute numbers in it are measured medians, not declared budgets — 380K vs 89K tokens and 22 vs 11 turns on the six newly-solved tasks (and, from the figure rather than the prose, 172K vs 125K on the 69 already solved). Those are the right quantities and they are reported as a diagnostic of where the extra spend went, not as a cap either arm was held to.
The generalizable point, and it is a third axis beyond the four half-defections above: a tier name is a request, a budget is a bound, and the difference becomes invisible the moment a paper prints the tier in the same column as the score. Moonshot's card at least leaves "reasoning effort = max" looking like the setting it is. Presenting the same object as a control converts a disclosure into a claim the disclosure cannot support — and it is what turns an uncontrolled comparison into a paper that reads as though it answered the budget question. The published-tier convention now needs the counterpart Brown asked for from the start: whoever names a tier should also publish the tokens it spent.
There is a second, subtler compute-control failure in the same paper worth separating out, because it is about the search rather than the inference. Harness evolution is a search procedure, so its own budget is part of what a comparison must hold fixed — and DarwinX states none: the TB2.1 loop runs "over many generations" with no generation count, no rollout total, and no cost, while the paper it contradicts caps its search at K = 5 rounds. Two empirical results on one benchmark, four orders of magnitude apart in search spend, neither denominated (Agent-Authored Harness Optimization).
The author says the control is unaffordable at the interesting scale (September 2026)#
This page's prescription — fix a budget and compare within it, or plot the whole curve — assumes the comparison can be bought. Brown (Dwarkesh Podcast, 2026-09-17, practitioner-opinion) reports, about his own flagship result, that at frontier scale it cannot be:
"We do measure it up to 16 or so agents in our published blog posts. The problem is that it's very hard to push that science to 10,000 agents because it's just so expensive."
And the specific control the result invites, refused outright: "We don't know how long it would take a single agent to solve Navier-Stokes, because we haven't done that experiment yet." His stated plan is to ablate at 64 / 128 / 256 and infer, with the concession that "it's going to be very hard to push that all the way to 10,000 and know for sure what the benefit was that we actually got from using 10,000 agents versus 1,000."
This is the page's argument arriving at its own boundary, and it is worth stating precisely because it is not the usual failure. The vendor here is not hiding the budget — the budget is the headline (10,000 agents, 130 billion tokens, 88 hours). What is missing is the counterfactual at a different budget, which is the other half of what a compute-controlled comparison requires. The measured region is 1/4/16 agents in a blog post; the reported region is 10,000; the party with the compute says the middle will stay empty. So the result is simultaneously the best-disclosed and least-controlled headline in the corpus.
Three consequences:
- Disclosure and control come apart. Every disclosure exemplar on this page (Kimi K3's footnotes, Gemini's price rows, Inkling's published curve) treats publishing the budget as the fix. This case publishes the budget fully and remains uninterpretable, because no adjacent budget was run. "Report the compute" is necessary and, at the extreme, not sufficient — what a reader needs is two points on a curve, and one of them can cost more than the result.
- It gives the "plot the whole curve" prescription a price ceiling. METR's dollar-axis argument and optstop's sample-truncation both assume the curve's points are individually affordable. At 88 GPU-hours × 10,000 agents they are not, and the honest form of the prescription becomes "plot the curve where you can, and label the extrapolated region."
- And it is a hard case for the forecast-from-cheap-runs hope this page and Large-Scale Test-Time Compute both file as an open question. Brown's own plan is that forecast — measure to 256, infer to 10,000 — stated by someone who does not believe it will settle the question.
The primary document lands the same week and makes the case sharper, not softer (2026-09-21).
OpenAI's own announcement (On the Navier–Stokes Millennium Prize Problem,
2026-09-08, vendor-claim) is the source Brown was describing, and it discloses more than he did:
~10,000 concurrent agents, 88 hours, 2.7 million inter-agent messages, ~130 billion output
tokens on Navier–Stokes, and 4.9 million messages / ~300 billion output tokens across every
problem attempted. Three observations this page should carry.
First, the disclosure is in four units and none of them is dollars — agents, hours, messages, tokens. That is the richest budget disclosure in the corpus and it still cannot be compared with FrontierMath Erdős Benchmark's $300, or with METR's dollar axis, without an undisclosed price per token for an undisclosed internal model. A budget reported in units nobody else reports is a disclosure that does not travel.
Second, the post does contain one adjacent configuration, which is one more than the interview admitted: the same campaign's Euler result at "nearly 100 agents… approximately 50 hours." It is not a control — different problem, model retrained mid-campaign, and the large run was seeded with the small run's output — but it is worth recording exactly because it looks like one. A reader who wanted two points on a curve could take them from this post and be wrong, which is the failure mode this page should name as loudly as the absent control itself.
Third, the campaign changed configuration mid-run by design: agents were moved off other problems, re-prompted with the Euler resolution, upgraded to a further-trained checkpoint when one became available, and cross-pollinated through Codex. So even the single published point is not a fixed configuration — there is no budget at which "10,000 agents solved Navier–Stokes" is a reproducible statement, and the disclosure's completeness makes that visible rather than hiding it.
The budget written into the score, and an author who refuses the better headline (September 2026)#
Every exemplar above treats the compute budget as something to disclose alongside a score.
FrontierMath Erdős (Epoch AI,
Announcing FrontierMath Erdős, empirical) is the corpus's first benchmark that makes
the budget constitutive of the score: an attempt is $300 of inference and 72 hours of working
time on one problem, one attempt per problem, harness and item list open-sourced. A run at a different
budget is not a worse-disclosed score; it is not a score.
The demonstration of what that buys is the part worth copying. Epoch ran the same pre-release model off-protocol at larger budgets with varied agent setups, and it did better — 5 of 68 problems solved for over $220,000, against 2 of 68 for roughly $20,000 under the protocol. It published those runs in full, with per-solution costs, and then refused them the label in bold: "These attempts are not a FrontierMath Erdős score." The headline stayed at 3%. This is the exact move the bad equilibrium above predicts nobody makes — a benchmark author holding a larger number it produced itself, and declining to report it as the result — and the reason it was available is structural rather than moral: the budget is in the definition, so the better run is disqualified by construction instead of being a judgment call the author has to win against their own incentives.
Two further properties, one strong and one a limit:
- The off-protocol data is published as data rather than suppressed, which gives the reader the ladder the previous section says is usually missing: per-problem attempt counts (7/7, 5/5, 2/5, 1/4, 1/4) and per-solve costs from $47 to $1,384, on the same problems, same model. So a benchmark can both enforce a budget and supply the adjacent-budget counterfactual, provided the two are labelled differently. That is a cheaper answer to the Navier-Stokes problem above than it looks — the unaffordable thing there was a second point on the curve at frontier scale, and here the second point cost 11× the first.
- The axis is dollars only. No token counts, no per-model API prices, no wall-clock beyond the 72-hour cap. Under the x-axis question below that is the buyer-appropriate choice for a benchmark whose users are deciding whether to spend on a research problem — but it means a model that gets cheaper per token shows up here as strictly more capable, and the 3% is not separable into "better reasoning" and "more tokens per dollar."
The same author publishes the curve, not just the cap (OEIS Open, 2026-08)#
The protocol above fixes the budget at one point and scores there, which answers "how capable at
$300?" and nothing about the shape. Six weeks earlier the same author published
OEIS OPEN (OEIS Open: How many conjectures can language models turn into theorems?, arXiv 2608.11941,
empirical), which does both: caps are constitutive ($50 per conjecture on the 492-item set, $200 on
a 100-item subset) and the full spend curve is published — for every run, the fraction of
conjectures resolved as a function of the spend at the moment each one was resolved, so reading the
curve at $x$ estimates a run capped at $x$. Across eight runs the rise is roughly linear in
log-spend at about ten percentage points per tenfold increase, with no plateau at the cap.
That is the strongest available answer to this page's standing complaint that a single number hides its budget: publish the whole curve from one capped run, at no extra compute, by recording when each solve happened. Two honest limits, one of which the paper states and one it does not. Stated: "the estimate is imperfect because agents are told their budget, which may affect their behaviour" — an agent told it has $50 does not behave like the first $50 of an agent told it has $200, so the reconstructed curve is not the same object as a ladder of independently capped runs. Unstated: the curve is right-censored at the cap, so it can show the absence of a plateau below the cap and can never show where one begins.
It also supplies the price contrast this page's other Epoch exemplar lacks, from the same evaluator in the same quarter: $6–$10 average per resolved conjecture with a $47 maximum here, against $172 to $1,384 per solve and ~$67,000 marginal on FrontierMath Erdős. Same instrument, same units, two denominators — which is what a dollar axis is for.
The variable that beats compute on a frontier cross-section: release date (2026-09-22)#
Everything on this page treats compute as the missing control. Zhu (Oxford Internet Institute, arXiv 2608.29420, empirical, single-author unrefereed preprint) measures a different one on a hash-pinned twelve-benchmark Artificial Analysis snapshot and finds it larger — and finds that adding compute to it makes things worse, not better.
The leading capability factor across those twelve benchmarks tracks model release date at a logistic R² of 0.505 (OLS 0.477), and residualising every benchmark on date before re-fitting drops that factor's share of common variance from 74.5% to 59.6%. On the compute-known subsample of 58 models the same date-only adjustment removes 16.5 points; adding log training compute jointly removes only 9.3 — plausibly because later models are also larger, so the two predictors split their shared variance. Prior psychometric work on score matrices (Kearns; Ilić & Gignac; Ruan et al.) all controls for scale; the paper's line is that on a frontier snapshot "the calendar carries what scale carries across generations."
Two consequences for this page, and they pull in opposite directions.
Compute control is still the right answer to benchmark-maxxing, which is a within-run fairness property: a score produced by a hidden best-of-N scaffold is not comparable to one produced by a single pass, and only a stated budget makes those two numbers mean the same thing. Nothing here touches that.
Compute control is not the right answer to cross-model comparability, which is what a leaderboard row actually invites. On a frontier cross-section, two models released nine months apart differ mostly by calendar, and a study that controls for parameters or FLOPs while leaving release date uncontrolled has removed the smaller of the two confounds and can report a shrunken correction for its trouble. The disclosure item this evidence demands is one no board treats as load-bearing and every board already has: the release date, with the prescription that a small gap between contemporaneous models be date-adjusted, or the comparison restricted to a single release window, before it is read as capability.
The caveats are the ones that apply to the whole paper: n = 96 complete cases, one operator, one date, a 14.9-point primary estimate whose bootstrap interval ([−5.3, +32.7]) is indistinguishable from zero, and an author who reports that hypothesis as failed rather than promoting the specification that passes.
The disclosure item beneath the budget: per question, or shared (2026-09-22)#
This page's whole argument is that a score without its budget is undefined. Fan et al. (arXiv 2608.07968, UMD, empirical) expose a prior choice nobody states because nobody sees it: every benchmark in the corpus gives each item its own budget. That is a scaffold decision, and it is not neutral.
The direct measurement (Appendix B, aligned scoring, seven reasoning models): give a model one budget for N questions instead of B/N for each, and the score moves — +2.6 points at N=5, where front-loading lets it finish a few questions with confidence, −0.4 at N=10, and −5.0 points at N=20, for every one of the seven models. Same model, same total tokens, same questions; the only change is whether the budget is partitioned by the harness or by the model. The authors are careful that the uniform arm is not an oracle — independent prompts also remove cross-question interference — but the direction and its monotonicity in N are unambiguous.
Three consequences for a page about disclosure:
- A per-question budget is a silent uplift on any benchmark whose items are administered separately, and it is an uplift that grows with the size of the batch a real deployment would hand the model. A leaderboard row is measured in the regime that flatters the model most.
- The grid cannot see the capability at all. Conventional one-question-at-a-time evaluation creates no tradeoff across problems, so a model that cannot ration across them scores identically to one that can. The paper's framing — global budget allocation is "not captured by conventional per-question evaluation" — is this page's complaint about compute restated one level up: not "at what budget?" but "budget over what unit?"
- The paper practises what it asks of others only partly. Its own mathematics budget B is never given a numeric value anywhere in the paper; only the code-domain B = 3,000 is stated, calibrated as ≈3× a single question's median reference cost. A study whose subject is budget pressure reports its primary-domain results without the constant that sets it — a clean instance of the disclosure gap this page tracks, in a paper arguing for budget-aware evaluation.
The budget's price date: fixed-capability cost falls ~47% a quarter (2026-10-01)#
A dollar budget is only a control on the date it was set. Epoch's The plunging price of thought (Emberson & Roodman, empirical) measures how fast it goes stale. The cheapest cost of reaching a fixed score falls ~47% per quarter on five benchmarks since 2023, fastest just after a level debuts as SOTA (66%/q) and slowest on SWE-bench Verified (27.5%/q). So a dollar-denominated score needs its price date as much as it needs its budget. The same report's secret-game benchmark gives the corpus's only measured size for benchmark-maxxing in a cost trend (44.0% against 47.0%). Full treatment, with the specification spread (42.9–58.0%) and the absence of standard errors, on The Price of Fixed Capability.
Connections#
-
The Price of Fixed Capability — the rate at which a dollar budget loses its meaning: ~47% per quarter at fixed performance (13×/yr), measured per task rather than per token, with the benchmaxxing gap on a secret benchmark sized at a few points per quarter
-
Shared-Budget Compute Allocation — the unit question beneath the budget question: every benchmark here silently budgets per item, and moving the same total tokens to one shared budget costs all seven tested models 5.0 points at N=20 while buying 2.6 at N=5
-
Interactivity Benchmarks — the voice-side instance of this page's problem: a cross-vendor chart that labels a reasoning-effort setting on every bar, runs the home model at High against a competitor at Medium, and pairs it with a measured cost on a named workload rather than a price row
-
Economic Benchmark Construct Validity — the control this page never named, measured against the one it is built on. Release date explains the leading capability factor at R² = 0.505, and adding log-compute to a date adjustment shrinks the correction (16.5 → 9.3 points on the same 58 models). Compute disclosure remains the answer to benchmark-maxxing; release-date adjustment is the answer to cross-model comparability, and they are different problems wearing the same clothes. See the section above
-
OEIS Open Benchmark — the same prescription with the curve published rather than only the cap: caps constitutive ($50 / $200 per conjecture), and a solve-rate-versus-spend curve reconstructed from when each solve landed inside the capped run, rising ~10 points per 10× spend with no plateau — free from one run, right-censored at the cap, and biased by the agent being told its budget
-
FrontierMath Erdős Benchmark — this page's prescription taken to its conclusion: the $300 / 72-hour budget is part of the score's definition rather than a footnote, and the author publishes its own larger-budget run (5/68 for >$220k against 2/68 for ~$20k) while explicitly refusing it the label of a score
-
Epoch AI — the third-party evaluator that wrote that protocol, and the corpus's clearest instance of a benchmark author acting against its own headline incentive
-
The Navier–Stokes AI Claim — the first-party budget disclosure behind that headline, in four units and no dollars, with one adjacent configuration (~100 agents / ~50 hours on Euler) that looks like a control and is not
-
Autonomous Scientific Discovery — the corpus's best-disclosed and least-controlled headline at once: the Navier-Stokes run publishes its full budget (10,000 agents, 130B tokens, 88 hours) and no adjacent budget, so "report the compute" is satisfied and the number is still uninterpretable
-
Evaluation Horizon Versus Release Cadence — the calendar version of this page's budget problem, and the one no disclosure discipline reaches: a cost axis prices the curve and says nothing about whether there is time to run it before the next model ships
-
Multi-Agent Collective Intelligence — where the unaffordable ablation bites: the pathway's central quantitative hope is a multi-agent scaling law, and the only party able to measure one at the rhetorically interesting scale says it will not be measured there
-
Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves — measures what a named-but-unbudgeted harness (Kimi K3's disclosure shape above) actually costs: the same model at the same success rate runs up to 2× apart in dollars depending only on which of three harnesses ran it
-
RSI Autonomy Levels (B0–L5) — this page's charge extended twice by a September 2026 RSI survey. Its §3.3.4 supplies the sharpest statement of the validity half — "once a benchmark is queried adaptively, it effectively becomes part of the optimization surface rather than a passive measurement instrument" — and its L5 evaluation protocol extends the budget half in a direction nothing here covers: original and revised improvement mechanisms must be compared under matched total budgets including the cost of evaluating the mechanisms, not just the cost of running them. The same survey's L3 section asks for the matched-budget control in the self-training literature (Rationale Bootstrapping (STaR)), where it has never been run — the proposed arm is to freeze the curriculum's learner-state input, or replace adaptive selection with a learner-independent schedule, at matched total acquisition-plus-learning budget
-
Headroom-Closed Index (HCI) — the only instrument in the corpus that puts a number on the vendor-provenance discount this page argues for: benchmark-owner tables weighted 3, independent common-harness evaluations 2.5, combined reports 2, model-author tables 1, and first-party values multiplied by a further 0.75. It prices disclosure where this page demands it — and then reproduces this page's central omission, since every HCI value pools scores taken at unknown and unequal test-time budgets with no compute axis anywhere
-
Evaluation-Time Answer Leakage — the environment-side answer to this page's certification question, and the source of the disclosure item above: a benchmark score is undefined not only without its compute budget but without the prompt it was produced under and the exploitation rate that prompt permitted. One paragraph appended to a stock agent prompt moves Pass@1 by up to 13.3 points in one direction on one benchmark and up to 3.5 in the other on another, while cutting judged exploitation by 41–73 points — an uncontrolled variable of scaffold magnitude that no grid records
-
Evaluation-Time Answer Leakage — a second condition a score is undefined without, and the first one anybody has certified. This page's rule is that a number without its compute budget means nothing; Zheng et al. show a number without its sandbox boundary means nothing either, and on SWE-Bench Pro the boundary was worth more than most budget differences — 14–26 points for six of seven models. The methodological transfer runs the other way too: their certification is not a claim about the score but a published procedure plus a trajectory audit (named operation classes, per-class before/after counts, confirmed-access counts), which is the shape a "no benchmark-maxxing" attestation would have to take
-
Skill Lift — a vendor benchmark that meets part of this page's disclosure bar and misses the rest. It pins the snapshot commit, states attempt counts (85% single-attempt), and says outright that it reports no confidence intervals; it names no cost or effort budget per arm, which matters more here than usual because the with-skill and without-skill arms demonstrably spend different token counts — one skill cut tokens 76.9%, another raised them 120.3%. An ablation whose two arms differ in spend is a compute-uncontrolled comparison by this page's own definition
-
Inference-Time Architecture Search — a specimen that half-obeys this page and half-embodies its critique: Archon emits exactly one final response (pass@1, not pass@k) and optimizes an accuracy-versus-inference-calls frontier internally, then reports +14.1% over GPT-4o and Claude 3.5 Sonnet without putting those baselines on the same axis. The gain is real; it is a multi-layer multi-model stack measured against a single call, which is precisely the benchmark-maxxing shape
-
Continuous Self-Modification Under Review — SOTA claims whose comparison arm is a leaderboard citation. Ouroboros's Terminal-Bench, OSWorld and CL-Bench margins are measured against published baselines on different models and different harnesses, not re-run, with the audited Terminal-Bench score sitting roughly two binomial standard errors above the strongest of them. The one arm where the baseline was re-run under a matched protocol (SWE-bench Pro, 655 paired tasks after symmetric decontamination) comes back statistically indistinguishable
-
Open-Ended Discovery Harnesses — a comparison that is controlled against one baseline and not the other, in the same table. SwarmResearch and the multi-agent baseline CORAL both run to a $50/task cap on the same runtime and model; the evolutionary baseline EvoX runs 100 iterations at ~$23.50/task average — roughly half the spend, and the arm where the reported margins are largest. Its second experiment repeats the pattern one level down: the winning orchestrator-guided arm uses a stronger model for the orchestrator than the fixed-scaling baseline has anywhere
-
Large-Scale Test-Time Compute — the root cause: capability scales with inference budget, so a score without a budget is undefined
-
Reward Hacking — benchmark-maxxing is Goodhart at eval-report time, the sibling of reward hacking in the training loop
-
Evaluation Awareness & Grader Gaming — the in-model version of gaming a measure; benchmark-maxxing is the same pressure applied by the evaluator rather than the model
-
Latent Capability Overhang — the flip side of the same axis: if grids under-report because they under-spend, released models hold capability nobody has paid to reveal
-
Benchmark Score Redundancy — the sibling cost reduction on the benchmark-count axis: where this page compresses eval cost by naming the compute budget per benchmark, that page predicts a model's held-out benchmark scores from ~5 probes (the matrix is rank-2); the catch it flags is that BenchPress runs on exactly the uncontrolled public grid this critique targets, so it inherits the grid's heterogeneity and vendor-optimism. Its second source, DeepMind's CollabEval, opens a third cost axis — the annotation budget within one benchmark, where a skipped prompt costs neither inference nor scoring — and is the arm that answers the uncontrolled-grid objection rather than inheriting it: three of its five score matrices are single-lab runs under a uniform harness, and the correlation it exploits is strongest there. Worth noting the direction of travel: it reduces eval spend without touching the compute-per-item this page asks to be reported, so the two compose cleanly (state the budget, then label fewer items at it)
-
Task Time-Horizon Scaling — the compute-controlled successor metric: reliable task length at a budget, rather than accuracy at an unnamed one; both face benchmark saturation
-
Expenditure Horizon — the axis question answered and then extended: dollars, because experiment compute is 70–90% of an agentic R&D trajectory's bill, and because a human returns curve can be plotted on a dollar axis and not on a token one — the crossing point is the score
-
Responsible Scaling Policy Evaluations — the safety-eval instance of the same demand: a threat-model determination reported without its compute budget is as under-specified as a capability score
-
Gemma 4 — the worked example: a thinking-mode model benchmarked against a non-thinking predecessor in its own headline table
-
Jagged Intelligence (Ghosts, Not Animals) — the confound cuts across model sizes too: Gemma 4's reasoning wins over a 10×-larger predecessor are partly bought with inference, while its knowledge losses are not recoverable that way
-
Inference Efficiency as Capability — the gain an uncontrolled grid structurally cannot show, which is why efficient models are systematically under-credited
-
The Open-Weight Frontier Gap — Arena Elo inherits the same defect: human preference scored at an unnamed inference budget
-
Open-Weight Elicitation Irreversibility — the stakes when the un-budgeted determination is a safety one and the weights are public
-
UK AI Security Institute / US Center for AI Standards and Innovation (CAISI) — the independent evaluators who adopted "report capability curves" in practice, and who published the corpus's most anonymous comparator three weeks later; empirical backing for the whole critique and a counter-exemplar on disclosure
-
Noam Brown — the source and author of the essay
-
Measuring Beyond Accuracy Saturation — this critique applied inside a saturated benchmark: Nadgir et al.'s efficiency axis plots accuracy against tokens and dollar cost (GPT-5.3-Codex ~60% cheaper than an equal-accuracy peer; token-cost and dollar-cost rank agents differently), which is "put compute on the x-axis" restated — and generalizes it, arguing accuracy should also be reported alongside reliability and model-vs-scaffold contribution, not only a compute budget
-
Benchmark Contamination and Decontamination — a sibling benchmark-trust critique on a different confound: this page says an unnamed compute budget makes a score meaningless; that page says training-data leakage makes it meaningless (memorization inflates it). Both attack taking the headline number at face value, and both propose recovering what the single number hides (the capability curve here; the clean per-sample distribution there)
-
Kimi (Moonshot AI) — the far end of the disclosure axis: per-benchmark harness pinning, named effort settings, self-deflating numbers, and rivals' fallback/refusal fractions volunteered — with no compute budget on any of it
-
Capability-Gated Model Fallback — the safety mechanism whose firing rate K3's footnotes accidentally measure from outside Anthropic
-
Cost-per-Task Over Cost-per-Token — why a price row is not a cost: the bill is rate × tokens-per-task, and only the rate is ever published
-
How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them? — the cluster synthesis: the unnamed compute budget is the first of five corruption channels, and capability curves at stated budgets are the put-compute-on-the-x-axis move of the five-part replacement portfolio
-
Agent-Authored Harness Optimization — both halves of this page's critique, on one method class, and now the method class where the critique's result is contested. Cline's headline is a harness-uncontrolled comparison (88.8% at $49.8 on a harness tuned for 17 hours against that benchmark, set beside Fable 5's $552 and GPT-5.6 Terra's $400 from unstated configurations); the class as a whole was compute-uncontrolled until Wang et al. ran the matched-budget arms above, which is where benchmark-maxxing stopped being an argument and became a measurement; and DarwinX then reported the opposite result on the same benchmark without running the arm — arguing its extra compute is well-allocated (4.3× tokens on newly-solved tasks) rather than equal, which is a different claim and the one this page's newest section takes apart
-
Process vs Outcome Reward Models — the same demand arriving from the reward-modelling side, stated by a lecturer who then cannot meet it: "usually the ORM and PRM should come from the same data for a good apples-to-apples comparison," and separately, the unanswered question of how to split a fixed budget between generator samples and verifier compute. The ORM-vs-PRM label-efficiency comparisons in that literature are explicitly not controlled — an ORM needs k labels per problem where a PRM needs k × steps
-
Weak-Verifier Ensembling — a rare half-observance: Weaver plots success rate against total inference FLOPs rather than reporting a single number, which is more than most, but its headline comparison against a proprietary model is not on that axis
-
Tree Search over Agent Trajectories (LATS) — a headline result published without the axis this page insists on: LATS converts test-time compute into quality across expansion, rollout, judging and backup, and its own lecturer concedes "the cost-benefit was not really analysed in the paper"
-
Adaptive Stopping in Evaluation Sampling — the other cost lever the same institution is building, and the one this page is most likely to conflate with the forecasting research noted above. AISI's [[raw/optstop-bayesian-optimal-stopping-llm-evaluations|
optstop]] (Pilditch, arXiv 2608.14425,empirical) does not predict high-budget performance from cheap runs: it fixes the budget and cuts the sampling needed to place each point precisely, stopping each item and each model-task grouping once its Bayesian credible interval is narrower than a declared threshold. Extrapolation across the compute axis and truncation within a point are different operations with different guarantees, and its own equivalence test certifies only the second. They compose — adaptive stopping is part of what makes running the full curve affordable — but nothing about a validated truncation transfers to a validated forecast -
Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — this page's demand restated as a legal-drafting requirement. Second-party reproducibility under a declared budget is one of seven properties an obligation-bearing benchmark would need, and it is one of only two the wiki shows are achievable today: AISI's minimum informative budgets and Kimi K3's per-benchmark harness pinning are the two halves of a perimeter protocol nobody has assembled. The tier-is-not-a-budget finding is the one that transfers hardest — a regulator reading a vendor-declared effort tier as a compute control is the DarwinX error with a market-access decision attached
-
AI-Assisted Error Analysis — a live claim awaiting this page's control. Preliminary unpublished benchmarking (Husain & Dasgupta) reports general-purpose coding agents finding failure modes more exhaustively than dedicated eval-discovery platforms, with no budget or effort tier reported for either arm — the undefined-effort-tier pattern this page names, in a comparison whose authors also sell the alternative
Open Questions#
- Can you certify "no benchmark-maxxing" — verify a reported score used a stated, reproducible compute budget rather than a hidden best-of-N scaffold? Sharpened (2026-09-10) — the tractable certification is of the environment, not the score: Zheng et al. (
empirical) certify a different score property on the same kind of artifact and show what the attestation looks like when it works. They publish the procedure (repository reconstruction to a single commit, test-artifact deletion with Git hooks disabled, instance-ID hashing, a named host blocklist), then audit the trajectories the runs already produced: eleven named local operation classes and six online ones with per-class before/after operation and task counts, plus a high-precision path check yielding confirmed answer-file access on 103 tasks locally and 49 over the network before, zero after. Nothing there depends on trusting the reporter's description of its own scaffold — the trajectories are the evidence, and a third party holding them could re-derive every number. The transferable form for this question: certify the run record, not the score, and require the log rather than the claim. The gap that remains is who holds the trajectories — this was self-audited, and the paper concedes a determined adversary routes around a blocklist. Sharpened again, two days later, with a disclosure item this page did not have: Ludwig et al. (NVIDIA,empirical) change one paragraph of the agent's prompt — nothing else, same harness, same models, same budget — and Pass@1 moves by up to 13.3 points on SWE-bench Multilingual (and by −3.3 to +3.5 on DeepSWE, so the sign is not even fixed). That is a swing of the same order as the scaffold effects this page already treats as uncontrolled, produced by a variable no leaderboard reports and no compute axis captures: the harness prompt is part of the budget disclosure surface, and nobody discloses it. It also gives the certification question a second, cheaper artifact alongside the run record — publish the exploitation rate next to the pass rate, since a score with no exploitation rate is as underspecified as a score with no compute budget, and both are recoverable from trajectories the evaluator already holds. The caveat that keeps it open: their exploitation rate is itself an LLM-judge verdict with no human validation, so it certifies less than a token count does. Partially answered from a third direction, 2026-09-21, and this one is definitional rather than evidential: Announcing FrontierMath Erdős (empirical) makes the budget part of what the score means — $300 and 72 hours, one attempt per problem, harness and item list open-sourced — so a hidden best-of-N scaffold does not produce an inflated score, it produces no score. The author then demonstrated the property on itself: its own off-protocol runs at larger budgets reached 5 of 68 problems for over $220,000 against 2 of 68 for ~$20,000, were published in full with per-solve costs, and were refused the label in bold. Certification here costs nothing to verify, because there is nothing to attest — the claim "this was produced under the stated budget" is either true or the number is not a FrontierMath Erdős score at all. What this does not do is close the question as posed, and the gap is specific: nothing stops a third party reporting an off-protocol number as if it were the score, and the only party that could check is Epoch, which cannot re-run a competitor's private harness. The transferable lesson sits alongside "certify the run record, not the score" rather than replacing it — define the budget into the metric, then publish the disqualified runs as data — and it is cheap only for a benchmark whose author also runs every evaluation, which is not how leaderboards with vendor submissions work. Extended 2026-09-22, from the standardization-not-audit direction: How UK AISI and EvalEval Are Making Benchmark Results Reproducible (empirical) pairs UK AISI with the EvalEval Coalition to publish verified Evaluation Cards — standardized benchmark+run+model-config records under the Every Eval Ever (EEE) schema — for five benchmarks across six models, accompanying AISI's How Inference Compute Shapes Frontier LLM Evaluation. This is a third register, distinct from the prior two: not a self-audit of trajectories (SWE-Bench Pro Verified) and not a definitional budget cap (FrontierMath Erdős), but a shared schema meant to make a benchmark+run+model-config triple checkable across publishers rather than within one paper. Its own release shows the ceiling on that promise: the Terminal-Bench 2.0 comparison plots this study's own scores against circles pulled from EEE for the same models — the intended cross-source check — but the published chart exposes only axis range and legend, not per-model numbers, to a reader who isn't scripting its SVG/canvas. "Verified" here means the record's provenance is checkable, not that its values are legible without tooling — a standard that certifies less than the trajectory audits above, and cheaper to produce at scale because of it. Sized, not certified (2026-10-01): Epoch's price-of-thought report (empirical) gives the corpus its first magnitude for benchmark-maxxing, measured indirectly. Mystery Game Puzzles hides the identity of its game so nobody can train for it. Its fixed-performance cost falls 44.0% per quarter model-free (38.7% in the preferred fit), against 47.0% for the five-benchmark average. The authors read the gap as the critique being "valid but not fatal": a few points per quarter of a cost trend, on one secret benchmark, with no intervals. It also supports the run record direction above. With CAISI's truncation method, a single high-budget transcript produces a model's whole accuracy-versus-spend curve, and announcing the budget to the model did not beat the lower-effort settings. So a third party holding the transcript can rebuild the budget curve without re-running anything. It does not certify anything about a hidden scaffold. See The Price of Fixed Capability. - Compute has several units (tokens, dollars, wall-clock). They diverge (a more efficient model wins on cost but not always on tokens). Which x-axis is the honest one, and does it depend on the buyer? (AISI reports against tokens on a log axis, and notes that as cost-per-token falls, the high budgets that reveal capability become progressively cheaper to reach.) Partially answered (2026-08): METR argues dollars, and supplies the measurement that makes the argument bite rather than assert — on an agentic AI R&D task, experiment compute is ~70–90% of trajectory cost, so tokens and dollars are not proportional and a token axis omits most of the spend. It also names the buyer the axis is for (a lab choosing between human and agentic labour, which only dollars can price) and demonstrates the payoff: a second, human curve can be drawn on a dollar axis and cannot be drawn on a token one. Still open in two respects — the finding is one task class (long-horizon optimization with heavy experiment compute; a chat or single-shot benchmark inverts the ratio), and the dollar axis buys comparability at the cost of importing a wage assumption and the evaluator's own harness efficiency. Extended 2026-09-21 with a second task class and a second buyer, both choosing dollars: Announcing FrontierMath Erdős (
empirical) denominates an open-research-mathematics benchmark in dollars per problem — $300 per attempt, per-solve costs from $47 to $1,384, ~$20,000 for a 68-problem run against >$220,000 for the off-protocol ladder — and reports no token counts at all. The buyer named here is not a lab choosing between human and agent labour but someone deciding whether to point a model at an unsolved problem, and for that buyer the price of a solve is the whole question. Two things this adds rather than repeats. First, the task class is the opposite of METR's on the experiment-compute axis: there is no training run and no GPU experiment, so this is nearly all inference, and the dollar axis is still chosen — which weakens the reading that dollars won only because experiment compute dominated. Second, it makes the cost of the choice legible in the other direction: with no token counts, a model that becomes cheaper per token is indistinguishable here from a model that reasons better, and 3% at $300 in September 2026 will not be comparable to 3% at $300 a year later. That is the standing objection to a pure dollar axis, stated as a dated artifact rather than as a worry. Extended 2026-09-21 by the mirror case, which picks every axis except dollars (On the Navier–Stokes Millennium Prize Problem,vendor-claim): OpenAI reports its Navier–Stokes campaign in agents, hours, inter-agent messages and output tokens (~10,000 / 88 / 2.7M / ~130B on the problem; 4.9M / ~300B across the campaign) and names no dollar figure at all. The pairing is the useful part. Epoch's dollars are comparable across vendors and rot with the price list; OpenAI's tokens and messages do not rot, and are uncomparable by construction — the model is unreleased, so no price per token exists and no third party can convert. So the honest answer to "which axis" now visibly depends on who can convert between them: a dollar axis presumes a published price, a token axis presumes a known model, and a first-party report of an internal system satisfies neither. This also adds a unit the question did not have — inter-agent messages, which is a communication-volume axis with no analogue in any single-agent budget, and which nothing in the corpus knows how to price. Extended 2026-10-01 with the rate at which the dollar axis goes stale: Epoch (empirical) measures the cheapest cost of a fixed score falling ~47% per quarter (~13×/yr) across five benchmarks, and ~53% per quarter on FrontierMath T1–3. At that rate the "3% at $300 will not be comparable a year later" objection above is about a 20× effect at fixed performance. The same report settles one sub-question outright: per-token price series are the wrong unit for a trend once reasoning models exist. Every earlier estimate (a16z 10×/yr; Epoch 2025 9–900×/yr) measured price per token of a threshold-clearing model. The report uses realized cost per task, turning tokens into dollars by truncating transcripts at each budget. The honest axis is therefore dollars plus a price date. Tokens stay comparable over time but not across models, and dollars are comparable across models but only on the same date. This doesn't pick an axis per buyer, so the question stays open. See The Price of Fixed Capability. - Does a compute-controlled evaluation regime advantage frontier labs (who can afford the full curve) over academics and third-party evaluators who can't? Sharpened (2026-07): a government evaluator (AISI) does run the full curves — so it is affordable to a well-funded public body — but AISI itself flags that "the most informative evaluations may be expensive" and is researching how to forecast high-budget performance from cheap runs precisely to relieve that cost. So cost is the binding constraint even for a funded third party; it just isn't fatal to one. Partially answered on a sibling axis (2026-07): BenchPress shows the analogous cost problem on the benchmark-count axis is largely solvable — a model's full 133-benchmark scorecard is recoverable from ~5 probes to within ~3.93 points — but that reduces which benchmarks to run, not the compute-per-benchmark the full-curve question is about, so it relieves eval cost on a different axis than the one this question poses. Partially answered (2026-08-14) — the affordability research this question was waiting on shipped, and it is regressive. The same evaluator released [[raw/optstop-bayesian-optimal-stopping-llm-evaluations|
optstop]] (Pilditch, arXiv 2608.14425,empirical), an open-source adaptive stopping framework that removes 57.2–97.3% of planned trials (mean 81.1%) from a nine-cell validation matrix with a pooled truncation effect of+0.0003(97% HDI [−0.002, +0.003], ROPE ±0.02). So the cost of running a curve is falling for a third party, by roughly a factor of five in the demonstrated configuration, released publicly rather than held internally — which cuts against the advantage half of this question. The sting is in the conditional the paper states itself: the savings scale with how over-sampled the design already was. The headline is a 200-item × 10-epoch design; the same framework on a leaner 100-item × 5-epoch design at the same threshold saves 59.5%, and the paper concedes that leaner configurations will necessarily afford less scope for early termination. An evaluator who could already afford 10 epochs per item gets 81%; one already running lean — precisely the party this question worries about — gets substantially less, and one running 1–3 epochs on agentic tasks is untested. The relief is real, and it is largest where it was least needed. Still open on the compute-budget axis proper, where the forecasting work this question originally referenced remains unpublished. - Gemma 4 controls for compute in its long-context table and not in its headline table, without comment. Is partial control worse than none — does it lend the uncontrolled tables borrowed credibility?
- Is a grid with a price row but no token counts more misleading than a grid with no cost information? Falsifiable directly: run the four models in DeepMind's table on one agentic benchmark, record tokens-per-task, and check whether the price-implied cost ranking survives. Partially answered from the voice domain (2026-09-21), and it swaps one missing half for the other. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (Google,
vendor-claim) reproduces an Artificial Analysis chart titled Cost per Hour of Input Audio, computed on a Big Bench Audio subset — a realised cost on a stated workload, not a rate: Gemini 3.8 Live $0.84, Gemini 3.1 Flash Live minimal $1.50 / high $1.75, Gemini 3.8 Live Extended Thinking $3.50, Grok Voice Think Fast 2.0 $4.80, GPT-Live-1 Astra $5.83. That is the artifact this question asks for, produced by a third party rather than the vendor, and the same lab that published the price-row-without-tokens grid is the one reproducing it. Two reasons it does not close anything. The denominator is an hour of input audio, fixed by the user, not a task and not a token — so the price-implied ranking is tested against a duration baseline, which is legible for a conversational product and meaningless for the agentic benchmark this bullet names. And the numerator is undisclosed: nothing states whether backend tokens, output audio or thinking time are inside it, and the methodology footnote printed on the chart points at the vendor's own page rather than the evaluator's. The general lesson is worth keeping even so: a cost column becomes checkable the moment the workload is named, and the axis a vendor picks for the denominator is itself a disclosure choice. Also worth noting for this page's own subject — every bar on all four of that post's charts carries an explicit reasoning-effort label (High / Medium / Minimal), more compute disclosure than most cross-vendor grids offer, and the home model is charted at High against a competitor charted at Medium with no explanation, which is the compute-control defect this page exists for appearing inside the one grid that bothered to label the setting. - Moonshot reports Fable 5 hitting fallbacks on 35% of SWE-Marathon tasks and 40% of Agents' Last Exam tasks "downgraded". Would re-running those with safeguards disabled change the ranking — and should a leaderboard publish the safeguarded score, the unsafeguarded one, or both? Partially answered (2026-07-23) on the should, by UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities. A government evaluator faced the choice and picked one: US closed-weight models evaluated "with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities," with a one-sentence disclosure that public versions have them enabled. So the answer in practice is publish the unsafeguarded score and label it — defensible for a safety evaluation whose subject is the model's ceiling, and it inherits a defect the question anticipated: the comparison arm (an open-weight model run as hosted) was not de-safeguarded, so a single table now mixes both conventions. The would it change the ranking half is still unmeasured — nobody has published paired safeguarded and unsafeguarded scores for the same model on the same suite.
Sources#
-
The plunging price of thought — Luke Emberson & David Roodman (Epoch AI), 2026-09-22 (
empirical, web report, data and code at github.com/droodman/inference-cost). Cited here for the ~47%/quarter fixed-performance cost decline and its 53.1% FrontierMath T1–3 row (Table 1, model-free), the per-token-versus-per-task critique of earlier estimates, the CAISI truncation method and the announced-budget test (Figure 3, from prose), and the Mystery Game Puzzles benchmaxxing comparison (44.0% / 38.7% against 47.0%). No standard errors, by design. Full treatment on The Price of Fixed Capability -
One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation — Louis Yiven Zhu (Oxford Internet Institute, single author, unrefereed preprint), arXiv 2608.29420, 2026-08-29, 25pp,
empirical. Cited here for the release-date section above: §5.2 (logistic date R² = 0.505 against OLS 0.477; the 74.5% → 59.6% adjustment and its [−5.3, +32.7] bootstrap interval; the 14.9 / 24.1 / 16.5-point specification comparison; the 9.3-point joint date-plus-log-compute drop), §2 (the scale-as-control lineage this inverts), and Appendix N (the pinned snapshot and the refused Data API request). Table-free citation except Table 4, which was reconciled againstpdftotext -layout. Full treatment on Economic Benchmark Construct Validity -
Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations — Toby D. Pilditch (UK AI Security Institute), Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations, arXiv 2608.14425, 2026-08-14 (
empirical, 32pp). Cited here only for the distinction between truncating sampling at a fixed budget and forecasting performance at a budget never paid for, and for Section 3.1's efficiency figures plus Table 4's leaner-design comparison (59.5% at 100 items × 5 epochs against 81.1% at 200 × 10) used in the open-question annotation above. Author's own tool, author's own institution, one experimental configuration. Full treatment on Adaptive Stopping in Evaluation Sampling -
Noam Brown – Agent swarms, alignment, & recursive self-improvement — Noam Brown (OpenAI) with Dwarkesh Patel, Dwarkesh Podcast, 2026-09-17 (
practitioner-opinion). §00:00:00 only: the 5.6 Ultra Mode 1/4/16 plots (four agents "twice as fast" at "2x more" cost), the statement that the science cannot be pushed to 10,000, the un-run single-agent Navier-Stokes counterfactual, and the 64/128/256 ablation plan with its own stated limit. The referenced blog-post plots are not inraw/, so the 1/4/16 figures here are quoted from a verbal description of a chart, with no axes, no benchmark names and no error bars — the weakest possible provenance for a scaling claim, and recorded as such. Full treatment on Multi-Agent Collective Intelligence -
Shortcutting the Fix: Identifying and Categorizing Agentic Exploits in Software Engineering Benchmarks — Ludwig, Ahmad, Majumdar & Ginsburg (NVIDIA), Shortcutting the Fix, arXiv 2609.06780, 2026-09-06 (
empirical, 16pp): §2.1's paired vanilla/principled prompt design (identical harness, models and budget; one paragraph changed) and Table 1's Pass@1 / Pass@3 / exploitation columns across five models × two benchmarks. Exploitation is a majority vote of three open-weight LLM judges with no human validation. Full treatment on Evaluation-Time Answer Leakage -
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents — Zheng, Shang, Jiang, Tian, Zhu, Ma, Yuan & Zhang (ECNU / Shanghai AI Lab / Fudan), SWE-Bench Pro Verified, arXiv 2609.08149, 2026-09-08 (
empirical, 37pp). Cited here for the certification form — a published isolation procedure plus Tables 5–7's operation-class trajectory audit — and for the size of the uncontrolled-condition effect it exposes. Self-audited, one model deep on the leakage tables. Full treatment on Evaluation-Time Answer Leakage -
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown — No Priors interview (2026-06-26); the benchmark-grid critique, benchmark-maxxing, the bad-equilibrium framing, and the routing/consensus argument (
practitioner-opinion) -
Rethinking the Evaluation of Harness Evolution for Agents — Yike Wang, Huaisheng Zhu, Zhengyu Hu et al. (Allen Institute for AI / University of Washington, arXiv 2607.12227, 2026-07-14,
empirical): the abstract's two charges (matched feedback and inference budgets; search and evaluation sharing one benchmark), §3 the unified budget formalism, §4.2–4.3 Tables 1–2 and the pass@1-versus-pass@5 diagnostic, §4.4 Table 3 the held-out split, §5.2 the concession that Terminal-Bench may lack both the headroom and the harness sensitivity a fair test needs. Tables 1–3 verified exact against the PDF at ingest (no collapse, no shift); Figure 1 reproduces Table 1's averages. Full treatment on Agent-Authored Harness Optimization -
DarwinX: Evolving Agent Harnesses Through Natural Selection — Zhang, Dai, Tan, Yang et al. (Salesforce AI Research / Agentforce, arXiv 2608.07545, 2026-07-31,
empirical, 33pp): §4 + Table 2 the Terminal-Bench 2.1 leaderboard rows with theirModel / effortcolumn; §4.1 the self-described "effort-controlled comparison" and Figure 5's measured per-task compute medians; §3.2 the paper's own concession that leaderboard rows "provide context rather than controlled comparison"; §9 "inference effort [is] not independently randomized"; Appendix B Table 8, the per-benchmark protocol table with no compute column. No effort tier is defined and no dollar cost is stated anywhere in the document — verified across the full body and appendices at ingest. Tables 1–5 checked cell-for-cell againstpdftotext -layout(clean: no collapse, no shift, captions above);canary-recall20/20, recall 1.00. Figure 5 read under the image two-pass rule and supplies the already-solved token medians (125K → 172K) the prose omits. COI: Salesforce evaluating its own proprietary agent. Full treatment on Agent-Authored Harness Optimization -
Gemma 4 Technical Report — Table 5 (thinking vs non-thinking across generations), Table 9 (the controlled long-context comparison), Table 6 (the confound repeated on vision) (
empirical) -
More compute, more capability: Why AI agent evaluations need to account for test-time compute — UK AISI (2026-07-02,
empirical): "report capability curves"; multiple-budget evaluation, reliability/reach-against-budget, and "minimum informative budgets" as adopted practice -
UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities — UK AISI / US CAISI, 2026-07-23 (
empirical, joint government evaluation): the same evaluator's disclosure floor — an unnamed "Top U.S. Models" comparator carrying every US figure in the document (76.2%, 20/41, step 28.5), a Figure 2 that names ten PRC models against an unlabeled US aggregate trendline, a "100M-token limit" with no per-task or per-attempt unit, and the safeguards-disabled convention stated in one sentence in "Detailed Results." All three figures viewed per the image two-pass rule; Figure 2 prints no data labels, so nothing quantitative can be read off it beyond gridline estimates -
Inkling: Our Open-Weights Model — the effort-sweep chart: a vendor publishing its own capability curve, against competitors' default operating points (
vendor-claim) -
Gemini 3.5 Flash-Lite Model Card — Google DeepMind, 2026-07-21 (
vendor-claim): the July 2026 evaluation table with input/output $/1M as its first two rows across four models, no token counts or effort settings anywhere, and GPT-5.4 mini / Claude Haiku 4.5 as the cross-vendor comparators -
On the Navier–Stokes Millennium Prize Problem — OpenAI (no byline), "On the Navier–Stokes Millennium Prize Problem", openai.com, 2026-09-08 with a 2026-09-10 update, ~1,900 words,
vendor-claim. Cited here for the budget disclosure only: ~10,000 concurrent agents, 88 hours, 2.7M messages and ~130B output tokens on Navier–Stokes, 4.9M and ~300B across all problems, plus the ~100-agent / ~50-hour Euler configuration and the mid-run model swap. No dollar figure appears anywhere in the post, and the model is unreleased, so no third party can convert its units into anyone else's. First-party about its own system, disputed, unverified. Full treatment on The Navier–Stokes AI Claim -
Announcing FrontierMath Erdős — Adamczewski & Burnham (Epoch AI), 2026-09-01 (
empirical, web article, ~2,250 words): the $300 / 72-hour / one-attempt protocol as a definition rather than a disclosure, the five-model score table, and the off-protocol cost ladder published-but-disqualified (5/68 for >$220k against 2/68 for ~$20k, per-solve costs $47–$1,384). Bounds: the only non-zero score is a pre-release GPT-6 Astra nobody outside Epoch can re-run; no token counts anywhere, so the dollar axis here is not convertible; the scored run is one attempt per problem, making 2/68 versus 0/68 a two-event difference. Full treatment on FrontierMath Erdős Benchmark -
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT — Cunningham, Shetty, Cheng & Rush (METR, 2026-07-21,
empirical): the "Calibrating to money vs calibrating to time" section and the measured 70–90% experiment-compute share of agent trajectories — the corpus's first argued-and-measured answer to which x-axis, plus the human-curve overlay that turns a budget axis into a comparison. Full treatment on Expenditure Horizon -
Kimi K3 Model Card — §3 footnotes 1–4 (2026-07-26,
vendor-claim): per-benchmark harness pairings, named reasoning-effort settings, the H20-recalibrated SWE-Marathon branch, the 73.7-vs-72.9 cross-harness disclosure, BrowseComp with and without 300K compaction, and the fallback/refusal fractions for Fable 5, GPT-5.6 Sol and GPT-5.5 -
OEIS Open: How many conjectures can language models turn into theorems? — Tom Adamczewski (Epoch AI), arXiv 2608.11941, 2026-08-12, 27pp,
empirical. Cited here for the $50/$200 constitutive caps, the Figure 2 spend curve and its ~10-points-per-decade slope (read from the figure image; the paper states the slope in prose as well), the per-solve cost figures from footnote 9, and the paper's own caveat that agents are told their budget. No table row is cited. Full treatment on OEIS Open Benchmark -
HarnessTax: How Much Does the Harness Matter for Coding Agents? — Pan, Yang, Arabzadeh, Chiang, Stoica & Zaharia, Arena.ai blog, 2026-09-16/18 (
empirical): cited here for the cost-vs-success grid showing a named harness's undisclosed budget can move cost 2–5× at matched success. Full treatment on Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves -
How UK AISI and EvalEval Are Making Benchmark Results Reproducible — Ghosh, Chim, Joshi, Srishti, Kennedy, Solaiman (EvalEval Coalition) with McFadyen, Tan, Coz (UK AI Security Institute), HuggingFace blog, 2026-09-22 (
empirical): the joint Evaluation Cards / Every Eval Ever release for five benchmarks across six models, cited here for the HLE oracle-feedback trajectory numbers (with the source's own non-comparability warning against standard fixed-budget HLE) and the Terminal-Bench 2.0 own-study-vs-EEE comparison chart. Both data visualizations are client-rendered iframe embeds; transcribed via a headless-browser read ofbody.innerTextafter mount, per the raw's own provenance note. The Terminal-Bench chart's per-model values are not recoverable from that capture — only axis range and legend text render as text. Full treatment also on UK AI Security Institute
Cited by 56
- Inference Efficiency as Capability×7
The headline is "an approximate 2.5× improvement in overall scaling efficiency over Kimi K2" — and…
- Agent-Authored Harness Optimization×6
Cost fell as score rose, which is the mechanistically interesting part: most of the gain came from…
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them?×6
Incentive shaping — benchmarks steer what labs optimize; retiring them doesn't remove the pressure,…
- Large-Scale Test-Time Compute×6
This is the training-side statement of Compute Controlled Benchmarking's reporting rule ("a gain…
- Open Questions Backlog×6
Compute Controlled Benchmarking (88d) — Gemma 4 controls for compute in its long-context table and…
- UK AI Security Institute×5
uk aisi evaleval benchmark reproducibility — How UK AISI and EvalEval Are Making Benchmark Results…
- Interactivity Benchmarks×4
Cost per hour of input audio, on a Big Bench Audio subset (Artificial Analysis) — Gemini 3.8 Live…
- Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It×3
Concept articles: Domestic Frontier Pacing and Frontier Ai Standards Body (the two proposals and…
- Headroom-Closed Index (HCI)×3
2. A provenance-weighted consensus score where several sources report the same model under the same…
- Kimi (Moonshot AI)×3
Against that: K3 is evaluated on the Kimi Code harness throughout while competitors are reported at…
- The Navier–Stokes AI Claim×3
This page exists because the event was already load-bearing across a dozen wiki pages before the…
- The Open-Weight Frontier Gap×3
Arena Elo carries no compute budget. Gemma 4's entries are thinking-mode models; the table doesn't…
- The Price of Fixed Capability×3
A score stated in dollars loses comparability over time. Frontiermath Erdos Benchmark's $300 cap is…
- Reward Hacking×3
Compute Controlled Benchmarking — benchmark-maxxing is reward hacking moved to the eval-reporting…
- Aakanksha Chowdhery×2
That is the mechanism side of Compute Controlled Benchmarking's reporting rule ("a gain visible…
- Tree Search over Agent Trajectories (LATS)×2
Cost was never analysed. Mirhoseini states this outright — every expansion, every rollout, every…
- AI-Assisted Error Analysis×2
the study is not out. It is worth flagging alongside Compute Controlled Benchmarking's
- Benchmark Contamination and Decontamination×2
Compute Controlled Benchmarking — a sibling reason a headline benchmark number can't be trusted at…
- Benchmark Score Redundancy×2
Would a public probe set become a Goodhart target? If "run these 5 benchmarks and infer the rest"…
- US Center for AI Standards and Innovation (CAISI)×2
Compute Controlled Benchmarking — its comparator anonymity is a disclosure counter-exemplar
- Capability-Gated Model Fallback×2
Caveats: these are counts from one vendor's evaluation runs, unaudited, with no per-task detail and…
- Continuous Self-Modification Under Review×2
The Terminal-Bench result is honestly bounded by its own authors. 387/445 raw; a trajectory audit…
- Cost-per-Task Over Cost-per-Token×2
Compute Controlled Benchmarking — the evaluation-side statement: a published price is a rate, and…
- Economic Benchmark Construct Validity×2
Second, adding compute makes the correction smaller, not larger. On the same 58 models, date alone…
- Epoch AI×2
Epoch's FrontierMath Erdős announcement does something benchmark authors rarely do: it reports a…
- Evaluation Horizon Versus Release Cadence×2
The budget problem is that capability is a curve over spend, so a single number is under-specified.…
- Expenditure Horizon×2
Two further claimed advantages: testing against a problem humans have already extensively optimized…
- FrontierMath Erdős Benchmark×2
This is the part with the widest reach beyond mathematics. Epoch's stated target is the opacity of…
- Gemma 4×2
The report's headline comparison (Table 5) puts Gemma 4 in thinking mode against Gemma 3 27B…
- Google DeepMind×2
Compute Controlled Benchmarking — the lab is now on both sides of it: Gemma 4's headline table is…
- Inference-Time Architecture Search×2
Note the design choice that makes the comparison meaningful at all: Archon emits exactly one…
- Inkling×2
Compute Controlled Benchmarking — publishes its own effort/performance curve; competitors still…
- Jagged Intelligence (Ghosts, Not Animals)×2
Compute Controlled Benchmarking — the compression comparison is confounded by thinking mode, which…
- Measuring Beyond Accuracy Saturation×2
Compute Controlled Benchmarking — the efficiency axis here (accuracy vs tokens vs dollar cost;…
- Multi-Agent Collective Intelligence×2
That is an infeasibility claim about the ablation, from inside the only organization that could run…
- Noam Brown×2
The benchmark grid is broken. Single-number benchmark tables don't control for test-time compute,…
- Process vs Outcome Reward Models×2
Compute Controlled Benchmarking — the generator-vs-verifier budget split the lecture names and does…
- Responsible Scaling Policy Evaluations×2
This sharpens two things already latent on this page. The RSP's reliance on "we use it daily and it…
- RSI Autonomy Levels (B0–L5)×2
Two protocol requirements attached to it are worth carrying: the original and revised mechanisms…
- Shared-Budget Compute Allocation×2
Compute Controlled Benchmarking — the benchmarking consequence: per-question is itself an unstated…
- Skill Lift×2
85% of published skills ran one attempt per task; 15% ran two. No confidence intervals are reported…
- Adaptive Stopping in Evaluation Sampling
Compute Controlled Benchmarking — the same evaluator, the same cost problem, a different lever, and…
- Artificial Analysis
Compute Controlled Benchmarking — AA's charts label a reasoning-effort setting per bar, which is…
- Autonomous Scientific Discovery
The cost is disclosed and enormous, which nothing else here does. 10,000 agents × 88 hours is the…
- Cline
case-study. Cline benchmarks its own harness, publishes its own scores, and does so on a suite it…
- Evaluation-Time Answer Leakage
Compute Controlled Benchmarking — the environment-side answer to its certification question. That…
- Harness Tax: Coding-Agent Cost Multiplies Across Harnesses While Success Barely Moves
Compute Controlled Benchmarking — harness choice is an undisclosed budget surface: two agents…
- Latent Capability Overhang
Compute Controlled Benchmarking — the reporting twin: grids under-report capability because they…
- Evals & Benchmarks
Compute Controlled Benchmarking — Noam Brown's critique: the single-number benchmark grid is broken…
- OEIS Open Benchmark
Compute Controlled Benchmarking — the budget written into the score's definition again, and here…
- Open-Ended Discovery Harnesses
Compute Controlled Benchmarking — the EvoX comparison is the uncontrolled half: a ~2.1× dollar…
- Open-Weight Elicitation Irreversibility
Compute Controlled Benchmarking — the capability-side sibling: a determination without a budget is…
- OpenAI
Inference-time-scaling research and its evaluation critique. Noam Brown — one of the pioneers of…
- Rationale Bootstrapping (STaR)
The controls it asks for are cheap and nobody runs them: freeze the acquisition mechanism's…
- Task Time-Horizon Scaling
Compute Controlled Benchmarking — reliable task length at a stated budget is a compute-controlled…
- Weak-Verifier Ensembling
Compute Controlled Benchmarking — the discipline the FLOP-efficiency curve half-observes: it does…
Related articles
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Measuring Beyond Accuracy Saturation
Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturate…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
