Howardism · Vol. 03Plate II · No. 02
Evals & Benchmarks, in order.
Notes39DomainEvals & BenchmarksOpen Qs121Newest1 Oct 2026Oldest14 Apr 2026
Benchmark design, contamination, LLM judges, and measuring capability.
Map of Content for the evals-and-benchmarks domain — 36 concepts. The science of measuring models: benchmark validity, contamination, saturation, LLM judges, and production-sourced evaluation. Curated entry point; see Home for all domains.
- Adaptive Stopping in Evaluation Sampling — UK AISI's optstop (Pilditch, arXiv 2608.14425): treat an LLM evaluation as a sequential measurement problem and stop sampling each item and each model-task grouping once its Bayesian credible interval is narrow enough — removing 57.2-97.3% of planned trials across nine validation cells with a pooled truncation effect of +0.0003, but measured at 200 items x 10 epochs with the low-performance safeguard never engaged, and buying an interval-width guarantee rather than frequentist coverage
- Aggregate Cancellation — The failure mode where a headline metric stays flat because two real effects of opposite sign sum to zero across strata — so a null pooled result is evidence of heterogeneity, not of no effect; the cleanest measured instance is a sparse-attention audit whose three preregistered pooled tests return p = 0.995 / 0.771 / 0.541 because per-cell effects run opposite ways, and the wiki's other instances (identical accuracy over disjoint correct subsets, a group cooperation rate over one systematically drained member) share the shape: the composition under the number changed and the number did not
- AI-Assisted Error Analysis — Shreya Shankar's account of the one eval step that resists automation: discovering what counts as a failure. The argument is epistemic rather than technical — what 'good' means lives in the developer's head, not in the traces, and a tool that could fully find and fix your product's failures could do the same for every competitor, so judgment is the only differentiator left. AI competence rises monotonically along the analyze → measure → improve lifecycle and is weakest at the start. The working division of labor: the human authors the failure-mode taxonomy, the agent builds a bespoke review interface, clusters and samples the traces, and then scales each human annotation back across already-labeled traces — never proposing taxonomy of its own, because validating an agent's taste costs more than expressing your own
- Benchmark Contamination and Decontamination — Sun, Zhan & Gales (Cambridge): per-sample distribution distances expose that aggregate-accuracy decontamination can worsen residual contamination, and Uncertainty-Based Decontamination (UBD) — deep LoRA ensembles exposing memorized samples as confident-but-batch-order-sensitive — debiases without a clean reference model; plus the two prevention-by-construction alternatives — a starting state past every model's cutoff (NanoGPT) and an item set of problems nobody has solved at all (FrontierMath Erdős), whose immunity expires only on success — with OEIS Open (2026-08) the first case where that expiry is already dated, at least 153 of its 492 items now carrying a published machine-checked proof.
- Benchmark Convergent and Discriminant Validity — Desai, Wallach, Chouldechova, Koyejo et al. (COLM 2026) run the multitrait-multimethod test over 48 benchmarks and 53 models: benchmarks claiming the same concept mostly do not rank models alike (safety-detection ρ 0.02), capability labels do not discriminate (reasoning correlates with knowledge at 0.74, above reasoning with itself at 0.66), and the strongest predictor of agreement is shared LLM-judge scoring (β 0.526 vs −0.058 for shared concept). BBQ-accuracy tracks reasoning more than bias; all nine headline statistics survive era and judge splits
- Benchmark Score Redundancy — Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix completion predicts held-out scores to ~4.6 MedAE and a 5-benchmark probe set recovers a full scorecard. DeepMind's CollabEval takes the same premise down to models × prompts and inverts its use — completion output becomes a control variate inside prediction-powered inference, so the redundancy buys unbiased estimates with valid confidence intervals whose correctness survives the matrix not being low-rank at all (and item-level matrices need ~16 components, not 2). Zhu then replicates the redundancy on a single-operator, single-harness grid where no score is vendor-reported — ρ = 0.79, one factor at 74.5% of common variance — which retires the reporting-bias objection and hands back a confound nobody controlled for: the leading factor tracks release date at R² = 0.505.
- Benchmark Task Defects (Spec–Test Mismatch) — Benchmark instances whose prompt and hidden tests disagree, so a pass or a fail stops meaning what the score says. OpenAI's July 2026 audit of SWE-Bench Pro's 731 public tasks (agent pipeline 27.4%, five-engineer panel 34.1%, ~30% stated) names four kinds: overly strict tests, underspecified prompts, misleading prompts and low-coverage tests. Three push scores down and one pushes them up, and OpenAI retracted its own recommendation of the benchmark
- Cognitive Capability Profiling for Task Suitability — Prunty et al. (Cambridge CFI) put AI systems and workplace tasks in one cognitive space: rubric-annotate 19,535 benchmark items for their demand on 18 capabilities (16 survive an inter-rater screen, clustered to 8), infer latent capability by Bayesian IRT where difficulty is annotated not estimated, and ask 410 workers to spend 100 points across capabilities per activity. Six systems share one profile shape — Semantic Memory 5.59, Social Cognition 4.08, Language 4.02 on top; Action Planning 1.99, Instrumental Reasoning 1.22, Object Permanence 0.29 at the bottom — and what separates the leaders is planning and control, not knowledge. A 5.30-wide dimension spread against a 1.12-wide system spread is the whole 'dimensions over families' claim; no variance decomposition is reported. Measures task importance, not demand, and is validated against no deployment outcome
- Compute-Controlled Benchmarking — Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performance against a cost budget instead; benchmark-maxxing, held-out private sets, the Goodhart equilibrium that keeps the grid alive, disclosure exemplars (Kimi K3's footnotes, Gemini's price rows), the first budget-matched test of a named method class, and DarwinX as the specimen of an undefined effort tier presented as a compute control; FrontierMath Erdős (2026-09) is the first benchmark to make the budget constitutive of the score rather than a disclosure item, and to publish its own larger-budget run while refusing it the label.
- Cross-Model Error Entanglement — Three 2026 audits of whether LLM errors are independent: Kuai et al. measure excess co-failure and same-distractor collision across 18 models; Kohli finds 9 frontier judges carry an effective sample size of 2.18, with majority vote 22pp short of the independence prediction; Hossain, Yousefi & Lim replicate cross-provider correlation matching within-provider on new judges and show correlated votes flip significance-test conclusions in up to 28% of comparisons
- Discovery Certification Protocol (DCP) — CMU's outcome-level audit for AI-research-agent claims: Gate 1 certifies a sealed utility gain, Gate 2 gives a fresh matched challenger the registered background and captured Web bytes (withholding the target run's own research history) and treats any valid route within tolerance as a veto witness, with a finite-sample recovery-probability bound from zero recoveries; optional Gate 3 randomizes truthful vs. a matched neutral feedback policy from a shared checkpoint. Two controlled audits (SQLite query optimization, virtual catalyst control) each returned 0/96 recoveries (upper bound 0.0468) and a Core+Evidence decision
- DRACO Benchmark — Perplexity's benchmark of 100 production-sourced deep-research tasks (10 domains, 40 countries) graded by 26-expert rubrics on accuracy/completeness/objectivity/citation; Perplexity Deep Research leads every domain and axis, Claude Opus 4.6 is the strongest non-Perplexity system, factual accuracy is the universal weak spot
- Economic Benchmark Construct Validity — Zhu's psychometric audit of a hash-pinned Artificial Analysis snapshot (421 configurations × 12 benchmarks, four economic; every hypothesis carried on the 96-model complete-case grid): one factor holds 74.5% of common variance and tracks release date at R²=0.505, so most of the leading 'capability' axis is calendar; the four economic benchmarks form no distinct factor under the pre-specified rule, yet leave-one-benchmark-out prediction with factors re-estimated in every fold beats a single mean index by a pooled ΔMSE of 0.037 [0.019, 0.055]. Verdict: economic benchmarks add incremental predictive information to a largely date-driven general factor without constituting a separate capability — and a gap between models released months apart is mostly a gap in release dates
- Evaluation Horizon Versus Release Cadence — Noam Brown's observation that the horizon a frontier model can operate over is growing faster than the interval between frontier releases, so there will be no window in which a model can be safety-evaluated at the full length of its own capability before its successor ships — a structural expiry on pre-release evaluation that the corpus's own budget and time-horizon curves already point at, plus the flip side Brown says has no good answer: slowing the cadence to buy evaluation time widens the gap between what a lab holds internally and what the world can use
- Evaluation-Time Answer Leakage — The channel by which an agent retrieves the reference solution during a benchmark run — residual Git objects, hidden test files, the target SHA embedded in the instance ID, upstream code hosts — rather than from training data; SWE-Bench Pro Verified (Zheng et al., arXiv 2609.08149) closes all four channels on 731 tasks and six of seven models lose 14–26 points, while the one model the audit found barely hacking loses 0.05; Ludwig et al. (NVIDIA, arXiv 2609.06780) leave the channels open and count them instead — 45.1–82.4% of vanilla trajectories on SWE-bench Multilingual are judged exploitative, and a four-sentence solution-originality instruction alone cuts that to 4.0–10.7% while leaving local Git access as the residual a technical block removes
- Expenditure Horizon — METR's continuous generalization of time horizon: the dollar spend at which an agent's improvement on an optimization problem equals a human's at the same budget, measured by crossing an agent's returns-to-expenditure curve with the local returns to human labor — proof-of-concept on the NanoGPT speedrun, where humans cost ~$2,500 per 1% speedup and six agent runs from record #78 re-validate to horizons of $0-$3,300, of which the maintainer would merge ~70% of the ideas but only 50-60% of the speedup
- GDPval Benchmark — OpenAI's benchmark of real, economically valuable knowledge work: 1,320 tasks built from the actual work product of industry professionals (4-year minimum, 14-year average experience) across 44 occupations in the nine sectors each contributing over 5% of U.S. GDP, scored as a blinded pairwise win rate against those professionals rather than as accuracy — 12.4% (GPT-4o) to 47.6% (Claude Opus 4.1) on the 220-task open gold subset, with the 'roughly linear over time' claim resting on three OpenAI models only; its speed-and-cost analysis is the corpus's cleanest measurement of oversight cost, collapsing a 474× naive cost advantage to 1.63× once expert review and rework are priced; the ancestor of the GDPval-AA Elo boards this wiki's model pages quote
- Headroom-Closed Index (HCI) — A cross-benchmark normalization that reports how much of a benchmark's remaining headroom a model has closed — 0 = the 90th-percentile frontier in the benchmark's entry year, 100 = a perfect score — aggregated into per-domain annual trajectories over 393 model-benchmark observations. Its finding is that progress is uneven in both level and shape: by 2026 advanced mathematics and graduate science sit at 86.4 and 85.8 while tool agents sit at 39.9 and software engineering at 52.6, and the closing rates move in opposite directions (mathematics accelerating 32.8→53.6, multimodal collapsing 59.7→2.5). Its weakness is that the normalizer is path-dependent on when a benchmark entered the dataset, and the underlying score table is not published
- Item Response Theory for LLM Benchmarks — Score a benchmark with a 3PL IRT model instead of percent-correct and two separable things follow. Scoring: ability θ reorders the leaderboard (23–31% of ~4,000 Open LLM Leaderboard models move >10 ranks despite Spearman 0.96–0.99) and lifts cross-benchmark rank consistency. Administration: Fisher-information item selection with an SE stopping rule reaches whole-bank ability on 30–89 of 627–5,600 items. Reordering comes from IRT scoring, not adaptivity. Family DIF (Zheng & Yang): items keep replicable model-family residuals, and low-DIF reweighting keeps global rankings but reverses 31–47% of near-tied cross-family pairs
- LLM-as-a-Judge — Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET + justification, weight-aggregated into normalized score and pass rate; key properties — rankings stay stable across judge models while absolute magnitudes vary, and adaptive per-case rubrics detect failures but blend them away, motivating stable custom metrics for the behavior under change; a 21-judge / ~541K-judgment audit finds raw exact-match agreement overstates chance-corrected reliability by 33–41pp (kappa deflation), so judges need chance-correction, bias, and cross-benchmark validation before thresholded use; upstream, CalibratedRubric makes the rubric bank the instrument — measurability, informativeness and validity are distinct, and unanimity filters decay with leaderboard size; OmniVChat's gated tiered rubric makes one missed basic criterion non-substitutable
- LLM-Judge Validation — UC Berkeley's 21-judge / 9-provider / ~541K-judgment audit (Norman et al., 2026): LLM-as-a-judge validation is systematically under-rigorous — exact-match agreement overstates chance-corrected κ by 33–41pp (kappa deflation, universal across every judge), judge rankings shift up to 14 positions across benchmarks, and high test-retest reliability masks severe position bias (the consistency–bias paradox); distilled into a 5-step Minimum Viable Validation Protocol. Yang et al. (2026) add the judge-version axis: upgrading the evaluator is not a reliability intervention (one robust step in 18, and scaling makes it worse on 2 of 4 datasets), and repeated-sample juries are capped by measured error correlation ρ ≈ 0.66–0.97. Chen et al. (2026) push it down to the rubric item: measurability (judge agreement), informativeness (IRT information) and validity are three different properties
- Matched Comparisons for Memorization Claims — Cooper et al. (arXiv 2607.12649): a generation rate measured only on training data is not a memorization rate — comparable non-training sequences must be scored by the identical procedure to supply a predictability floor. A conformal test calibrates the threshold to a chosen false-positive rate for populations, a census calibrates a single document against a matched control book, and 'extractable memorization' is redefined to require both a calibrated claim and near-certain generation within a realistic query budget.
- Measuring Beyond Accuracy Saturation — Princeton-led case study (arXiv 2606.26158): accuracy saturation is not benchmark saturation — re-instrument a saturated benchmark instead of retiring it, because statistically-indistinguishable agents still differ sharply in reliability, cost-efficiency, scaffold contribution, and human-collaboration speedup (CORE-Bench); plus a second saturation mode where the answer key never saturates but the human reference class does, and re-eliciting that baseline becomes the recurring cost
- Orchestration-Plan Simulation — OrchBench (Ren et al.): score a multi-agent orchestration plan without running workers — a deterministic simulator over a fixed task DAG correlates r=0.816 with real Claude Code quality at ~1% of the tokens; transfer coverage dominates agent count, multi-agent wins only under context pressure, and the headline correlation weakens sharply once the weakest planner is dropped.
- The Price of Fixed Capability — How fast the cheapest way to reach a fixed benchmark score gets cheaper: Epoch AI (Emberson & Roodman, 2026-09) measures ~47% per quarter (~13× per year) across five benchmarks since 2023, from cost per task rather than price per token, via CAISI-style budget truncation of transcripts. Math falls fastest (50–52%/q), games slowest (39–43%/q), SWE-bench Verified slowest of all (27.5%/q). The decline is steepest just after a level debuts as SOTA (66%/q, falling to 32%/q two years on). It is a frontier rate that no buyer who doesn't switch models every quarter gets, and it replaces the corpus's '10–100× per release cycle' figure, which overstates it by an order of magnitude or more
- Production-Sourced Evaluation — Building benchmarks from de-identified real production usage rather than synthetic or hand-authored tasks; DRACO's central method — difficulty-proxied sampling, PII-stripping, augmentation, automatable refresh with a human QA gate; representativeness vs. over-specification tradeoff; production traffic as a proprietary eval asset; plus the buyer-side instance, where a customer builds the eval from its own engineering work to decide what to buy, and the training-side instance, where production failures become RL trajectories and the difficulty proxy stops being independent of the model; plus the opposite pole, where production traffic does not exist at all and OmniVChat generates every stimulus with a multi-agent engine, then validates the substitution against 360 recordings the generator never touched — matching training deltas, disagreeing rankings
- Recurring Production Agent Evaluation — She & Lin (CMU/Meta, arXiv 2609.21267, adopter-voice deployment report): 574 runs of a production analytics agent's 519-question benchmark over 52 days, split chronologically 287 calibration/287 held-out, compare random sampling, historical outcome caching, fixed difficulty-stratified subsets, and Rasch/multidimensional-2PL adaptive testing for recurring evaluation of one evolving system. Multidimensional-2PL adaptive testing wins on score fidelity from k=200 onward (1.03pp MAE at 38.5% of the benchmark, vs Rasch's 1.37pp and random's 1.79pp), and difficulty-stratified fixed subsets win at small budgets — but the team deployed the fixed subsets anyway, for operational simplicity over accuracy, and backed that choice with unrecalibrated transfer to five other agent families and a calibration window shrunk to one day without loss.
- Reference-Free Judge Over-Crediting — Reference answers are a first-order determinant of LLM-judge verdicts: without one, judges systematically over-credit wrong answers (up to 85% verdict flips when the reference is added — Kranti & Vajjala), and self-play against a reference-free judge inflates pass rate at flat true accuracy (Zhou); the fix is the judge committing its own answer first.
- Scale-Dependent Prompt Sensitivity — Large models underperform small ones on 7.7% of standard benchmarks due to overthinking; brevity constraints recover 26pp and fully reverse hierarchy on GSM8K/MMLU-STEM
- Skill Lift — NVIDIA SkillEvaluator's with/without-skill ablation turned into a publication gate: three pre-publication tiers (safety-and-structure scanning, catalog distinctiveness, live sandboxed A/B), a per-skill delta in rubric points reported alongside the artifact, and a first at-scale benchmark — 300+ verified skills × 30+ products × 2 harnesses, +41 Correctness and +39 Effectiveness, +34 (Claude Code) vs +29 (Codex), per-product spread +2 to +46, token cost moving in both directions (-76.9% and +120.3% on two single-attempt examples). The measurement's load-bearing weakness is that the eval set is generated from the skill under test
- Task Time-Horizon Scaling — METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7): Opus 3 ~4min (Mar 2024) → Opus 4.6 ~12hr (2026) → weeks projected for 2027; paired with benchmark saturation (SWE-bench, CORE-Bench)
- Typed Decision Verifiers — Structured-verdict verifiers (boolean/choice/ordinal, zero generated tokens) benchmarked against 12 rivals on a shared 2,018-item answer-correctness set: the top two are statistically tied, a ten-feature surface baseline beats 8 of 13 systems, and a retry-gate experiment shows AUC does not predict deployed value because gate value is set by precision. TypeSafe AI's Jev launch (vendor-claim) pitches the category wider, as RL-trained 'smart if-statements' on a self-run cost Pareto frontier
- Usage-Telemetry Classifier Validation — Google ATLAS is the first AI-usage-economics program to publish accuracy numbers for the LLM classifiers every such study rests on — and they are humbling: 22.6% exact accuracy on O*NET task assignment and 42.5% on occupation title, against 85.8% human approval of the same labels; the accuracy/approval gulf, its mitigations (presence-not-frequency, task-type aggregation to 70.4%), and what it means for every headline number in the genre
- Verbalized-Confidence Soft Scoring for LLM Judges — Having a judge state a 0–100 confidence next to its True/False verdict, rather than reading a soft score off token logprobs (G-Eval). Hsiao (Cisco, 2026) finds a two-part prompt recipe — an overconfidence advisory plus single-call self-debate — cuts calibration error and widens score spread on all 10 flagship judges tested, flattens the soft score's coupling to task subjectivity, and on GPT-5-era models overtakes G-Eval on subjectivity robustness; post-2025 judges take the recipe at no accuracy cost while pre-2025 judges lose 1.8pp balanced accuracy. The measurement that makes a judge's calibration checkable now that Claude, Gemini 3.x and most reasoning models return no logprobs
- Version-Dependent Judge Error — A frozen LLM judge's false-accept and true-accept rates move with the agent version it grades, so a judge-only old-vs-new comparison can be biased as well as attenuated. Li (arXiv 2609.34198) on 20 prespecified SWE-bench Verified version pairs, τ-bench and AgentRewardBench: rank correlation 0.71–0.79 yet 8 of 60 judge-pair units declare upgrades execution cannot establish; failed-patch acceptance rises with agent capability (ρ ≈ +0.82, replicated on 8 OpenHands configs at +0.94); transporting old-version calibration multiplies comparison error 5.2×; a paired PPI++ audit is valid but saves ~5% of labels at 80 tasks.
- Weak-Verifier Ensembling — Weaver (Stanford, 2025): instead of training a better verifier, combine a pool of imperfect ORMs, PRMs and LLM judges by weak supervision, lifting hard benchmarks from ~40% to 70%+ and distilling into a ~400M scorer. Its load-bearing assumption — that verifiers err independently — is the one 2026 judge-correlation measurements say is false; Sunkavalli and Hossain, Yousefi & Lim show shared error is hard to identify and dependence-aware filters help only under the failure structure they target
Derived#
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them? — Synthesis of the 2026 eval-science cluster: public benchmark suites carry far less independent signal than their count implies (133 benchmarks ≈ rank-2; accuracy saturates even after validity fixes) and the headline number is corrupted through five distinct channels (unnamed compute budget, contamination, vendor optimism, unvalidated judges, evaluation-time answer leakage) — but ordinal comparisons survive under verified invariances, and nothing replaces benchmarks wholesale: the field's answer is a six-part portfolio (predict-don't-run, re-instrument saturated suites, compute-controlled curves, production-sourced refresh, judge validation, harden the sandbox), with failure-mode discovery, contamination monitoring, and incentive shaping as the jobs only benchmarks still do
- Construct or Roster? What BEI's Entanglement Graph and IRT's Ability Ranking Are Measuring — Two #oq/now items answered as one: both read shared structure off a model × item matrix while conditioning on something fitted from the same roster, so what the structure means depends on who is in the roster. (1) BEI's top tier is a tier effect, not a family effect. Kuai's own Table C.2 is 10 of the 15 pairs among six 2023–24 open-weight checkpoints, with cross-family Llama↔Qwen pairs at ranks 2, 3, 6 and 10, so vintage, weakness and base-vs-instruct status are collinear and cannot be separated on that roster. Frontier panels whose difficulty is conditioned from outside the roster (Kohli, Hossain) keep the dependence and do not organise it by vendor. (2) No page tests θ̂ against accuracy as a predictor of an external outcome. ATLAS's convergence row compares θ̂ with θ̂ on one calibration population. The only out-of-sample latent-vs-raw comparison is Zhu's, one level up: a single factor loses badly to a crude mean (R² 0.110 vs 0.771) and three factors win narrowly (0.808). Both questions stay open, retagged #oq/source, with the settling experiment specified
- How Do You Write Evals for Taste? Character as the Limit Case — Taste-driven features are eval-resistant but not eval-proof: the technique is conviction → dogfood-sourced failure signals → A/B variant measurement (MSM's method) → ~10 interpretable judgment-encoding evals; demonstrated on safety/values, still open on warmth/wit
Open questions 121 open
- SourceDoes the efficiency survive a lean or agentic evaluation design, and is there a design below which adaptive stopping costs more than it saves? The measured 81.1% is a 200-item x 10-epoch configuration; the same framework at 100 items x 5 epochs and the same threshold returns 59.5%, and the paper concedes leaner configurations afford less scope. Nobody has run it on a 1-3 epoch agentic suite, where per-trial cost is highest and repetition counts are lowest — and where the fixed cost of repeated MCMC re-analysis could plausibly exceed the trials saved. Falsifiable directly: run
optstopin shadow mode over an existing agentic eval log and compare the trials it would have saved against the inference overhead. - SourceIs a width guarantee enough for a decision that needs a location guarantee? The framework is explicit that it targets interval width rather than coverage, and its own simulation puts empirical coverage at ~80% against a 94% nominal level under moderate heterogeneity — with zero coverage at exact performance boundaries, which is where a dangerous-capability threshold would sit. Any governance use of an adaptively-stopped evaluation needs the coverage number, not the width number, and no source in this corpus reports one at a boundary under a well-specified model with known truth.
- SourceDo evaluators actually satisfy the randomisation precondition? Randomised presentation order is required for exchangeability,
sample_shuffleis off by default in the frameworkoptstopships against, and ordinal efficiency moves 38 points (57-92% fixed-order versus 95.7-97.3% shuffled) on order alone. Falsifiable by audit: sample published Inspect evaluation configurations and count how many enable shuffling.
- SourceDoes the efficiency survive a lean or agentic evaluation design, and is there a design below which adaptive stopping costs more than it saves? The measured 81.1% is a 200-item x 10-epoch configuration; the same framework at 100 items x 5 epochs and the same threshold returns 59.5%, and the paper concedes leaner configurations afford less scope. Nobody has run it on a 1-3 epoch agentic suite, where per-trial cost is highest and repetition counts are lowest — and where the fixed cost of repeated MCMC re-analysis could plausibly exceed the trials saved. Falsifiable directly: run
- Aggregate Cancellation2 open
- SourceDoes cancellation survive on a task-accuracy outcome, or only on influence proxies? The one controlled demonstration measures a logit-margin contrast; "the benchmark curve stays flat" is inferred, never measured. A sweep reporting per-cell accuracy alongside the calibrated contrast at the same compression ratios would settle whether the sign heterogeneity reaches the number operators actually watch.
- SourceIs there a cheap screen for cancellation that does not require knowing the strata in advance? Every detection move here presumes a labelled cell structure. A variance- or sign-based diagnostic computable from per-item results alone — flagging "this null is heterogeneous" without a pre-specified stratification — would make the check routine rather than a study design. Nobody in the corpus has proposed one. Sharpened (2026-09-10) by anchor judge error correlation estimator, which builds a screen of exactly this shape one field over and prices it honestly. Sunkavalli's problem is a judge/anchor panel rather than a benchmark grid, but the object is the same: a model that predicts a set of second moments should all be equal, and the diagnostic is their coefficient of variation, thresholded at the 95th percentile of a null distribution simulated from the correctly-specified model over 1,000 seeds — which buys a 5% false-positive rate by construction at every sample size, without anyone naming the strata in advance. Three things transfer as a template rather than as an answer. The recipe: dispersion statistic + simulated-null threshold, not a test statistic with an analytic distribution. The price, measured: power rises with signal and N from a floor that is genuinely uninformative — at a loading spread of 0.2 the test's power is 0.04 at N = 500 and 0.11 at N = 2000, at or below its own false-positive rate, so mild heterogeneity at realistic sample sizes is invisible to it, and a clean screen is not evidence of homogeneity. The calibration's fragility: the 5% rate is a correct-model Gaussian property, and heavy-tailed data inflates it to 0.083, so the null must be recalibrated on matched marginals for real data. It also names the structural limit any such screen inherits — a violation loading uniformly across every unit is invisible to any second-moment statistic, because it is observationally indistinguishable from the thing being measured. The question stays open because nothing here is a benchmark-grid instrument: nobody has run a dispersion-plus-simulated-null screen over per-item benchmark results.
- Shankar's skill has the agent apply each new human annotation to already-labeled traces
- The agent is stated to be non-exhaustive when re-applying a flagged failure mode across
- Does the general-purpose-agent-beats-dedicated-tool result survive a budget-matched
- SourceCan the ensemble be derived from one released model? The whole method rests on having several checkpoints differing in batch ordering; the authors flag single-checkpoint ensemble derivation (e.g. via cheap perturbations) as the key unlock for adoption. Until then it needs provider cooperation to release a LoRA ensemble.
- SourceDoes it extend past MCQ? UBD-Debiasing is classification-only today; whether per-decoding-step debiasing recovers the clean distribution for open-ended generation (where contamination shows as near-verbatim reproduction) is untested.
- SourceIs batch-order sensitivity a reliable memorization tell at pretraining scale? The signal was validated on 3B models with 5 LoRA seeds and induced contamination; whether the high-confidence-high-variance signature survives full-scale pretraining and real (not synthetically injected) leakage is open. Partially answered: Matched Comparisons for Memorization Claims (Cooper et al., arXiv 2607.12649,
empirical) settles the second half — real, non-injected leakage is detectable at pretraining scale, on OLMo 2 7B–32B against its published corpus and Llama 3.1 8B/70B against Books3 — but with a different tell: a matched non-member baseline rather than ensemble variance, needing no extra checkpoints. It also bounds what an uncalibrated statistic is worth at that scale (a 10-token verbatim match is ~24% false positive; at 50 tokens the floor is 0.02%). The batch-order signature itself remains untested above 3B. - SourceDoes correcting toward an ensemble-averaged uncontaminated reference introduce its own bias? The D_KL/D*_L1 targets are themselves an average over a 5-member uncontaminated LoRA ensemble; how much the "clean" target moves with ensemble size/composition is unexamined.
- SourceDoes capability non-discrimination survive removing the general factor? The authors name the test and do not run it: residualise every benchmark on the first principal component (and, following Economic Benchmark Construct Validity, on release date), then recompute within-minus-between ρ. The finding is falsified if reasoning, knowledge and comprehension separate once the shared axis is gone. The released item- and benchmark-level dataset makes this a re-analysis rather than a new collection.
- SourceIs the LLM-judge method effect a judge effect or a refusal-family effect? On this grid, judge scoring and the refusal/over-refusal concept family nearly coincide. The falsification test is to score a capability benchmark set with an LLM judge (free-response math or QA with rubric grading) alongside its exact-match version and check whether the judge-scored versions pull away from their own concept.
- WaitWill frontier model cards replace BBQ-accuracy as their only bias number? Table E.1 shows it as the sole bias metric on Claude Sonnet 4.6 and GPT-5 and one of two on Claude Opus 4.6. Trigger event: the next major release cycle's system cards. The prediction is falsified if any of those labs reports a bias metric that this paper's relabelling test does not move toward a capability concept.
- WaitDoes the rank stay 2? The geometry is snapshot-specific; whether a third latent factor emerges as the matrix grows (novel capability profiles, new benchmark families) is the paper's own named signal for when the recipe needs refreshing — and an open empirical watch item. Partially answered (2026-08-04) — on the stakes rather than the fact, by CollabEval. Nobody has re-measured the 84×133 matrix, so whether its rank holds is still open. But CollabEval shows the consequence of the geometry breaking is entirely a function of how you spend it: used as a prediction (BenchPress) a broken rank silently corrupts the answer; used as a control variate inside PPI, correctness is independent of the rank and only efficiency degrades, gracefully, back to the classical sample mean. The question stays live for BenchPress's use case and is largely defused for the inferential one. Two further data points against the geometry being permanent: at the item level rank 2 is nowhere near enough (>16 components for >50% of variance, and IterativeSVD keeps improving through rank 32), and the redundancy is weakest exactly on the largest matrix (MMLU, ~0.66 cumulative EVR at 32 components). Extended 2026-09-22 by Zhu, which argues the retention rule is the wrong instrument for the question. On his twelve-benchmark single-harness grid every in-sample dimensionality criterion — Horn's parallel analysis, the Kaiser rule, the scree elbow — returns k = 1, which on this page's framing would read as "the rank fell to 1." Out of sample it has not: a three-factor representation beats a single mean index at predicting held-out benchmarks (pooled ΔMSE +0.037 [+0.019, +0.055]), and sweeping k ∈ {2,3,4,5} the gain rises monotonically (+0.021 / +0.037 / +0.058 / +0.059, pooled test R² 0.79 → 0.83). So usable dimensionality exceeds retained dimensionality, and a rank chosen by held-out error — which is exactly how BenchPress picked 2 — is measuring something a retention rule is not. The corollary for this bullet: "does the rank stay 2" should be watched on the completion-error curve, not on a variance-explained threshold, and on Zhu's evidence the two need not move together.
- SourceCan outlier models be anchored without any scores? BenchPress fails on a model whose capability profile has no close neighbor in the matrix; the authors propose folding in external metadata (training-data composition, architecture, size) to compute model-to-model similarity before any benchmark is run, but do not build it. Partially answered (2026-08-04) by CollabEval — not by anchoring the outlier, but by bounding what its failure costs. Its anchor ablation finds CI coverage flat across anchor-set size and selection strategy, with positive variance reduction even at the smallest sets, and the control-variate construction means a target model uncorrelated with every anchor yields a wider interval, not a wrong one. It also inverts the intuitive anchor policy: Random-$k$ > Top-$k$ > Bottom-$k$, so curating anchors toward the strongest models is actively worse than sampling them. The metadata-similarity idea remains unbuilt. Adjacent evidence (2026-10-01), not an answer: Zheng & Yang find that a model-family label carries replicable item-level signal beyond an 8–128-dimension response fit. The sign of that signal changes across benchmarks, with every family flipping its mean shift at least once, so family metadata would inform similarity only per benchmark, not as a fixed prior.
- SourceDoes the low-rank treatment carry beyond text/vision? Audio/speech, robotics/embodied agents, and scientific-simulator ecosystems are untested; whether the same rank-2 structure holds there is open. (CollabEval does not touch this — all five of its datasets are text generation.)
- WaitWould a public probe set become a Goodhart target? If "run these 5 benchmarks and infer the rest" becomes practice, the compact probe set is a small, public, high-leverage surface to optimize against — the same eval-report Goodhart pressure Compute-Controlled Benchmarking names, now concentrated on five benchmarks. Unexamined here.
- SourceIs a valid interval around an autorater's mean worth anything? CollabEval's guarantee is about sampling uncertainty in the mean of the score a benchmark already computes — and on three of its five datasets that score is an autorater or learned metric (GPT-4 Turbo win-rate, AutoAIS, MetricX), which LLM-Judge Validation shows can be systematically wrong in ways no amount of tighter sampling detects. Directly falsifiable: run CollabEval on a task with both human labels and autorater scores, take human labels as $Y$ and the autorater as an anchor row, and check whether the reported CI actually covers the human parameter — or only the autorater's. Sharpened (2026-09-10) by anchor judge error correlation estimator, on the assumption the falsification test would have to make. That paper's "anchor" is a different object from CollabEval's — an external reference set whose errors might correlate with the raters', not a densely-evaluated model row — but it supplies the parametric form of the worry and one result that changes how the experiment should be read. It models the reference's contamination
ρas a free parameter rather than assuming it away, and shows that the standard alternative, designating one reference clean, does not fail noisily: as the trusted reference's own contamination runs 0 → 0.6 the companion's estimate degrades linearly to−0.002against a true 0.5, reporting a heavily contaminated reference as perfectly clean. So "take human labels as $Y$" is itself the assumption under test, and a CI that covers the human parameter certifies the composition only if the humans are uncontaminated — which the same paper's Test A rejects on the one real human panel it examines (three HANNA raters: statistic 0.589 against a 0.247 matched null,p < 0.0005). Two further constraints on any such measurement: separating shared rater error from shared quality needs≥ 2judges and≥ 2anchors, and with all scores ordinal the contamination is not identified at any number of anchors. The question is unchanged and the experiment is still the right one; it just cannot be run with an assumed-clean human column. Sharpened again (2026-08-14) — a second framework with the same blind spot, and it does not even name it.optstop(optstop bayesian optimal stopping llm evaluations,empirical) decides when to stop sampling on the width of a Bayesian credible interval — and six of the nine cells in its validation matrix, the entire ordinal and continuous columns, score WritingBench with an LLM judge. So its stopping decisions, its efficiency claims and its equivalence verdicts are all statements about the sampling uncertainty of a judge's mean, and unlike CollabEval's authors, who at least document the PPI ancestry that exists to relate cheap ratings to expensive ones, this paper raises the point nowhere. That makes the question load-bearing in a second place and adds a distinct failure mode to the experiment this question proposes: an adaptive rule stops earlier when the rater is more self-consistent, so a judge with a stable systematic bias is precisely the judge that terminates collection fastest and most confidently. It also supplies a concrete instance of estimand drift inside the same machinery — its ordinal pathway stops on the modal category while its own primary equivalence test evaluates means, a mismatch worthmu_diff ≈ −0.059in the high-ordinal cell and 1/15 interval coverage. Precision about the wrong quantity is a hazard that shows up before you even reach the human-versus-autorater question. - Resolved~~Does vendor optimism manufacture the correlation?~~ Four-fifths of the scores are provider-reported and possibly inflated; the paper flags this could inflate apparent cross-benchmark correlation but cannot separate it — would a fully standardized re-evaluation still be rank-2, or is some of the redundancy an artifact of shared reporting bias? Partially answered (2026-08-04) by CollabEval, at the item level. Its AQA (21 systems from one paper, one AutoAIS scorer) and WMT24++ (15 systems, all scored with MetricX under one protocol) matrices are single-lab runs under a uniform harness with no vendor self-reporting anywhere — and they are the strongest arms in the whole study (+17.3% and +12.9% CI-size reduction at 50% labeled). So exploitable cross-model correlation clearly survives standardized re-evaluation; it is not an artifact of shared reporting bias. What this does not settle is BenchPress's actual question, since these are item-level matrices within one benchmark rather than a standardized redo of the 84×133 cross-benchmark grid — the correlation surviving at one granularity is evidence, not proof, at the other. Answered 2026-09-22 by Zhu (
empirical, single-author preprint), at the right granularity this time. His grid is a cross-benchmark score matrix with no vendor self-reporting anywhere — a hash-pinned Artificial Analysis snapshot in which one operator runs all twelve benchmarks under a single harness — and the redundancy is not merely present but stronger than on the 84×133 public grid: mean off-diagonal Spearman ρ = 0.79 with every pair positive, first principal component 79.4% of total variance, parallel analysis retaining a single factor at 74.5% of common variance, and ten of twelve benchmarks in one correlation-distance cluster. So exploitable cross-benchmark correlation survives standardised re-evaluation at the cross-benchmark granularity, not just the item-level one CollabEval settled: shared reporting bias is not necessary for the redundancy. Two scope limits, stated because the answer is not free. It is a different matrix — 12 frontier-relevant benchmarks against 133, and 96 complete cases against 84 models — so nobody has re-run the 84×133 grid itself under a uniform harness, and three of Zhu's twelve benchmarks are constructed or subset by the operator whose harness also runs them, which trades vendor optimism for operator-specific item construction. And it hands back a different confound in exchange, which is why this question retires rather than dissolves: on that clean grid the leading factor tracks release date at R² = 0.505, and removing the date trend costs it 14.9 points (24.1 on the deduplicated base-model grid). The correlation is real and not manufactured by reporting bias — a large share of it is manufactured by comparing models of different vintages. That successor question lives on Economic Benchmark Construct Validity.
- SourceDoes SWE-Bench Pro Verified's repaired set still fail an OpenAI-style audit, and at what rate? Zheng et al. repaired 102 tasks chosen from public complaints and never examined low-coverage tests. OpenAI flagged 249 by human review but released no list. Falsifiable: run OpenAI's two-path procedure (or any five-reviewer panel) on the Verified release and count broken tasks by category. The prediction from the table above is that low-coverage defects survive at roughly the original rate.
- SourceHow much of a frontier SWE-Bench Pro score on the overly-strict subset is retrieval? Overly strict tests reward reproducing the gold patch. Falsifiable: split Zheng et al.'s Baseline → Verified drop by OpenAI's per-task category. If strict-test tasks lose disproportionately under isolation, the two defects compound as argued above. It needs OpenAI's per-task labels, which are unreleased.
- SourceIs the agent path's low-coverage undercount structural? The prediction is that an auditor reading failures misses defects that produce no failure. Falsifiable: seed known low-coverage tasks (delete assertions from a clean task) and measure the agent path's recall against the other three categories.
- Does capability-profile suitability correlate with any realized outcome? Nothing here is scored against
- Are the Object Permanence, Instrumental Reasoning and Action Planning floors a property of the models or of a
- Does an annotator that is also a subject bias its own profile? Gemini 3 Flash annotated the demand levels of
- SourceCan you certify "no benchmark-maxxing" — verify a reported score used a stated, reproducible compute budget rather than a hidden best-of-N scaffold? Sharpened (2026-09-10) — the tractable certification is of the environment, not the score: Zheng et al. (
empirical) certify a different score property on the same kind of artifact and show what the attestation looks like when it works. They publish the procedure (repository reconstruction to a single commit, test-artifact deletion with Git hooks disabled, instance-ID hashing, a named host blocklist), then audit the trajectories the runs already produced: eleven named local operation classes and six online ones with per-class before/after operation and task counts, plus a high-precision path check yielding confirmed answer-file access on 103 tasks locally and 49 over the network before, zero after. Nothing there depends on trusting the reporter's description of its own scaffold — the trajectories are the evidence, and a third party holding them could re-derive every number. The transferable form for this question: certify the run record, not the score, and require the log rather than the claim. The gap that remains is who holds the trajectories — this was self-audited, and the paper concedes a determined adversary routes around a blocklist. Sharpened again, two days later, with a disclosure item this page did not have: Ludwig et al. (NVIDIA,empirical) change one paragraph of the agent's prompt — nothing else, same harness, same models, same budget — and Pass@1 moves by up to 13.3 points on SWE-bench Multilingual (and by −3.3 to +3.5 on DeepSWE, so the sign is not even fixed). That is a swing of the same order as the scaffold effects this page already treats as uncontrolled, produced by a variable no leaderboard reports and no compute axis captures: the harness prompt is part of the budget disclosure surface, and nobody discloses it. It also gives the certification question a second, cheaper artifact alongside the run record — publish the exploitation rate next to the pass rate, since a score with no exploitation rate is as underspecified as a score with no compute budget, and both are recoverable from trajectories the evaluator already holds. The caveat that keeps it open: their exploitation rate is itself an LLM-judge verdict with no human validation, so it certifies less than a token count does. Partially answered from a third direction, 2026-09-21, and this one is definitional rather than evidential: epoch frontiermath erdos announcement (empirical) makes the budget part of what the score means — $300 and 72 hours, one attempt per problem, harness and item list open-sourced — so a hidden best-of-N scaffold does not produce an inflated score, it produces no score. The author then demonstrated the property on itself: its own off-protocol runs at larger budgets reached 5 of 68 problems for over $220,000 against 2 of 68 for ~$20,000, were published in full with per-solve costs, and were refused the label in bold. Certification here costs nothing to verify, because there is nothing to attest — the claim "this was produced under the stated budget" is either true or the number is not a FrontierMath Erdős score at all. What this does not do is close the question as posed, and the gap is specific: nothing stops a third party reporting an off-protocol number as if it were the score, and the only party that could check is Epoch, which cannot re-run a competitor's private harness. The transferable lesson sits alongside "certify the run record, not the score" rather than replacing it — define the budget into the metric, then publish the disqualified runs as data — and it is cheap only for a benchmark whose author also runs every evaluation, which is not how leaderboards with vendor submissions work. Extended 2026-09-22, from the standardization-not-audit direction: uk aisi evaleval benchmark reproducibility (empirical) pairs UK AISI with the EvalEval Coalition to publish verified Evaluation Cards — standardized benchmark+run+model-config records under the Every Eval Ever (EEE) schema — for five benchmarks across six models, accompanying AISI's How Inference Compute Shapes Frontier LLM Evaluation. This is a third register, distinct from the prior two: not a self-audit of trajectories (SWE-Bench Pro Verified) and not a definitional budget cap (FrontierMath Erdős), but a shared schema meant to make a benchmark+run+model-config triple checkable across publishers rather than within one paper. Its own release shows the ceiling on that promise: the Terminal-Bench 2.0 comparison plots this study's own scores against circles pulled from EEE for the same models — the intended cross-source check — but the published chart exposes only axis range and legend, not per-model numbers, to a reader who isn't scripting its SVG/canvas. "Verified" here means the record's provenance is checkable, not that its values are legible without tooling — a standard that certifies less than the trajectory audits above, and cheaper to produce at scale because of it. Sized, not certified (2026-10-01): Epoch's price-of-thought report (empirical) gives the corpus its first magnitude for benchmark-maxxing, measured indirectly. Mystery Game Puzzles hides the identity of its game so nobody can train for it. Its fixed-performance cost falls 44.0% per quarter model-free (38.7% in the preferred fit), against 47.0% for the five-benchmark average. The authors read the gap as the critique being "valid but not fatal": a few points per quarter of a cost trend, on one secret benchmark, with no intervals. It also supports the run record direction above. With CAISI's truncation method, a single high-budget transcript produces a model's whole accuracy-versus-spend curve, and announcing the budget to the model did not beat the lower-effort settings. So a third party holding the transcript can rebuild the budget curve without re-running anything. It does not certify anything about a hidden scaffold. See The Price of Fixed Capability. - SourceCompute has several units (tokens, dollars, wall-clock). They diverge (a more efficient model wins on cost but not always on tokens). Which x-axis is the honest one, and does it depend on the buyer? (AISI reports against tokens on a log axis, and notes that as cost-per-token falls, the high budgets that reveal capability become progressively cheaper to reach.) Partially answered (2026-08): METR argues dollars, and supplies the measurement that makes the argument bite rather than assert — on an agentic AI R&D task, experiment compute is ~70–90% of trajectory cost, so tokens and dollars are not proportional and a token axis omits most of the spend. It also names the buyer the axis is for (a lab choosing between human and agentic labour, which only dollars can price) and demonstrates the payoff: a second, human curve can be drawn on a dollar axis and cannot be drawn on a token one. Still open in two respects — the finding is one task class (long-horizon optimization with heavy experiment compute; a chat or single-shot benchmark inverts the ratio), and the dollar axis buys comparability at the cost of importing a wage assumption and the evaluator's own harness efficiency. Extended 2026-09-21 with a second task class and a second buyer, both choosing dollars: epoch frontiermath erdos announcement (
empirical) denominates an open-research-mathematics benchmark in dollars per problem — $300 per attempt, per-solve costs from $47 to $1,384, ~$20,000 for a 68-problem run against >$220,000 for the off-protocol ladder — and reports no token counts at all. The buyer named here is not a lab choosing between human and agent labour but someone deciding whether to point a model at an unsolved problem, and for that buyer the price of a solve is the whole question. Two things this adds rather than repeats. First, the task class is the opposite of METR's on the experiment-compute axis: there is no training run and no GPU experiment, so this is nearly all inference, and the dollar axis is still chosen — which weakens the reading that dollars won only because experiment compute dominated. Second, it makes the cost of the choice legible in the other direction: with no token counts, a model that becomes cheaper per token is indistinguishable here from a model that reasons better, and 3% at $300 in September 2026 will not be comparable to 3% at $300 a year later. That is the standing objection to a pure dollar axis, stated as a dated artifact rather than as a worry. Extended 2026-09-21 by the mirror case, which picks every axis except dollars (openai navier stokes millennium prize solution,vendor-claim): OpenAI reports its Navier–Stokes campaign in agents, hours, inter-agent messages and output tokens (~10,000 / 88 / 2.7M / ~130B on the problem; 4.9M / ~300B across the campaign) and names no dollar figure at all. The pairing is the useful part. Epoch's dollars are comparable across vendors and rot with the price list; OpenAI's tokens and messages do not rot, and are uncomparable by construction — the model is unreleased, so no price per token exists and no third party can convert. So the honest answer to "which axis" now visibly depends on who can convert between them: a dollar axis presumes a published price, a token axis presumes a known model, and a first-party report of an internal system satisfies neither. This also adds a unit the question did not have — inter-agent messages, which is a communication-volume axis with no analogue in any single-agent budget, and which nothing in the corpus knows how to price. Extended 2026-10-01 with the rate at which the dollar axis goes stale: Epoch (empirical) measures the cheapest cost of a fixed score falling ~47% per quarter (~13×/yr) across five benchmarks, and ~53% per quarter on FrontierMath T1–3. At that rate the "3% at $300 will not be comparable a year later" objection above is about a 20× effect at fixed performance. The same report settles one sub-question outright: per-token price series are the wrong unit for a trend once reasoning models exist. Every earlier estimate (a16z 10×/yr; Epoch 2025 9–900×/yr) measured price per token of a threshold-clearing model. The report uses realized cost per task, turning tokens into dollars by truncating transcripts at each budget. The honest axis is therefore dollars plus a price date. Tokens stay comparable over time but not across models, and dollars are comparable across models but only on the same date. This doesn't pick an axis per buyer, so the question stays open. See The Price of Fixed Capability. - Does a compute-controlled evaluation regime advantage frontier labs (who can afford the full curve) over academics and third-party evaluators who can't? Sharpened (2026-07): a government evaluator (AISI) does run the full curves — so it is affordable to a well-funded public body — but AISI itself flags that "the most informative evaluations may be expensive" and is researching how to forecast high-budget performance from cheap runs precisely to relieve that cost. So cost is the binding constraint even for a funded third party; it just isn't fatal to one. Partially answered on a sibling axis (2026-07): BenchPress shows the analogous cost problem on the benchmark-count axis is largely solvable — a model's full 133-benchmark scorecard is recoverable from ~5 probes to within ~3.93 points — but that reduces which benchmarks to run, not the compute-per-benchmark the full-curve question is about, so it relieves eval cost on a different axis than the one this question poses. Partially answered (2026-08-14) — the affordability research this question was waiting on shipped, and it is regressive. The same evaluator released `optstop` (Pilditch, arXiv 2608.14425,
empirical), an open-source adaptive stopping framework that removes 57.2–97.3% of planned trials (mean 81.1%) from a nine-cell validation matrix with a pooled truncation effect of+0.0003(97% HDI [−0.002, +0.003], ROPE ±0.02). So the cost of running a curve is falling for a third party, by roughly a factor of five in the demonstrated configuration, released publicly rather than held internally — which cuts against the advantage half of this question. The sting is in the conditional the paper states itself: the savings scale with how over-sampled the design already was. The headline is a 200-item × 10-epoch design; the same framework on a leaner 100-item × 5-epoch design at the same threshold saves 59.5%, and the paper concedes that leaner configurations will necessarily afford less scope for early termination. An evaluator who could already afford 10 epochs per item gets 81%; one already running lean — precisely the party this question worries about — gets substantially less, and one running 1–3 epochs on agentic tasks is untested. The relief is real, and it is largest where it was least needed. Still open on the compute-budget axis proper, where the forecasting work this question originally referenced remains unpublished. - SourceGemma 4 controls for compute in its long-context table and not in its headline table, without comment. Is partial control worse than none — does it lend the uncontrolled tables borrowed credibility?
- SourceIs a grid with a price row but no token counts more misleading than a grid with no cost information? Falsifiable directly: run the four models in DeepMind's table on one agentic benchmark, record tokens-per-task, and check whether the price-implied cost ranking survives. Partially answered from the voice domain (2026-09-21), and it swaps one missing half for the other. gemini 3 8 live announcement (Google,
vendor-claim) reproduces an Artificial Analysis chart titled Cost per Hour of Input Audio, computed on a Big Bench Audio subset — a realised cost on a stated workload, not a rate: Gemini 3.8 Live $0.84, Gemini 3.1 Flash Live minimal $1.50 / high $1.75, Gemini 3.8 Live Extended Thinking $3.50, Grok Voice Think Fast 2.0 $4.80, GPT-Live-1 Astra $5.83. That is the artifact this question asks for, produced by a third party rather than the vendor, and the same lab that published the price-row-without-tokens grid is the one reproducing it. Two reasons it does not close anything. The denominator is an hour of input audio, fixed by the user, not a task and not a token — so the price-implied ranking is tested against a duration baseline, which is legible for a conversational product and meaningless for the agentic benchmark this bullet names. And the numerator is undisclosed: nothing states whether backend tokens, output audio or thinking time are inside it, and the methodology footnote printed on the chart points at the vendor's own page rather than the evaluator's. The general lesson is worth keeping even so: a cost column becomes checkable the moment the workload is named, and the axis a vendor picks for the denominator is itself a disclosure choice. Also worth noting for this page's own subject — every bar on all four of that post's charts carries an explicit reasoning-effort label (High / Medium / Minimal), more compute disclosure than most cross-vendor grids offer, and the home model is charted at High against a competitor charted at Medium with no explanation, which is the compute-control defect this page exists for appearing inside the one grid that bothered to label the setting. - SourceMoonshot reports Fable 5 hitting fallbacks on 35% of SWE-Marathon tasks and 40% of Agents' Last Exam tasks "downgraded". Would re-running those with safeguards disabled change the ranking — and should a leaderboard publish the safeguarded score, the unsafeguarded one, or both? Partially answered (2026-07-23) on the should, by aisi kimi k3 cyber assessment. A government evaluator faced the choice and picked one: US closed-weight models evaluated "with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities," with a one-sentence disclosure that public versions have them enabled. So the answer in practice is publish the unsafeguarded score and label it — defensible for a safety evaluation whose subject is the model's ceiling, and it inherits a defect the question anticipated: the comparison arm (an open-weight model run as hosted) was not de-safeguarded, so a single table now mixes both conventions. The would it change the ranking half is still unmeasured — nobody has published paired safeguarded and unsafeguarded scores for the same model on the same suite.
- SourceCan you certify "no benchmark-maxxing" — verify a reported score used a stated, reproducible compute budget rather than a hidden best-of-N scaffold? Sharpened (2026-09-10) — the tractable certification is of the environment, not the score: Zheng et al. (
- SourceThe paper reports three judges and its Spearman degrees of freedom imply two (n = 34 = 2 × 17 on MMLU-Pro, n = 26 = 2 × 13 on MATH-500; three judges would give 51 and p ≈ 1.4 × 10⁻⁴ rather than the reported 0.002). Which judge is missing from the headline association, and does it come back in at the same sign? Settleable by one sentence from the authors or by a per-judge breakdown of ΔPrec.
- SourceDoes the entanglement penalty buy anything once the verifier pool is large enough for redundancy to matter? The +0.1pp over competence-only reweighting is measured at
M_J = 3, whereR_maverages over two other verifiers and the exponential penalty has almost nothing to discriminate. The falsifiable version: rerun the reweighting atM_J= 5, 7, 9 and report the entanglement-minus-competence delta as a function of pool size — a flat line kills the method, a rising one is the paper's real result. - SourceIs the BEI graph measuring entanglement or vintage? Its top 10 pairs are entirely Llama and Qwen models from 2023–2024 while no frontier pair appears at all, and difficulty is estimated from the other 16 models, so a pair of uniformly weak models has a high fitted failure probability everywhere. Restricting the roster to one release window (or residualising the error matrix on release date, as Zhu does on the score matrix) would say whether the intra-family signal survives. Partially answered (2026-09-22) by Kohli, from the side rather than head-on. His nine judges are all current frontier endpoints — no three-generation spread — and the panel still shows φ̄ = 0.391 with the three most correlated pairs all cross-family (Claude × Gemini 0.603, GPT-4o × Claude 0.588, Mistral × DeepSeek 0.564) and the same-family excess only +0.047. So on a single-release-window roster the dependence does not vanish and does not organise by lineage, which is evidence against the vintage reading of BEI's intra-family top tier. It is not the experiment this bullet asks for: different instrument (raw verdict-level φ, not difficulty-conditioned excess co-failure), different task (NLI classification, not MMLU-Pro), and no open-weight tier in the panel at all — so it says the cross-family finding survives a vintage restriction, not that BEI's intra-family ranking would. Corroborated again (2026-09-25) by Hossain, Yousefi & Lim, on a third instrument and a fourth judge roster: GPT-5.6-sol, Claude Opus 5 and Grok 4.5 — all current frontier endpoints, no generational spread — show cross-provider mean error correlation (0.42) statistically indistinguishable from the within-provider Gemini bank's (0.40), and the gap survives adjustment for shared item difficulty (0.465 → 0.467). Three independent groups, three instruments (BEI/CIG's difficulty-conditioned excess co-failure, Kohli's Kish n_eff, this paper's Ledoit-Wolf ρ̄), three judge rosters, one direction: vendor boundaries do not predict lower error correlation among frontier judges. The head-on version — rerun BEI itself on Kuai's roster restricted to one release window — is still unrun. No direct bearing from Li (2026) (2026-10-01). That paper measures how a judge's error depends on the graded agent's version and capability, not whether two answerers' failures co-occur, so it cannot say whether BEI's intra-family tier is vintage. It bears on the judge-bias half instead (see the fourth caution above): capability, a correlate of vintage, structures judge over-endorsement. Partially answered again (2026-10-01) by Construct or Roster? What BEI's Entanglement Graph and IRT's Ability Ranking Are Measuring, from Kuai's own Table C.2. The top 10 is not an intra-family list. It is 10 of the 15 pairs among six 2023–24 open-weight checkpoints: Llama-2-70b, Llama-3-70B, Llama-3.1-70B, Qwen1.5-110B, Qwen1.5-72B-Chat and Qwen1.5-14B-Chat. Four of the ten are cross-family Llama↔Qwen pairs, at ranks 2, 3, 6 and 10. Llama-2-70b ↔ Qwen1.5-110B (0.0398) outranks every intra-family pair except Llama-3 ↔ Llama-3.1. Inside that clique, tier orders the edges and family does not. On this roster, tier bundles three things: vintage, weakness and post-training status. By the paper's naming convention, four of the six are base checkpoints. Llama-3.1-70B-Instruct, from the same release as its base, is in none of the ten, while the base is in four. A single-release-window restriction would not separate those three variables on this roster. Zheng & Yang's family-DIF residuals, which survive an 8–128-dimension ability adjustment, show that family is a real variable in open-weight response matrices beyond capability. But that finding is DIF rather than co-failure, its sign flips across benchmarks, and the paper includes no release-date control. So the premise of an intra-family signal is weaker than the paper's prose suggests. What remains open is whether the tier effect is vintage, weakness or base status. The settling experiment is a within-window roster that crosses family with base and instruct variants of the same weights over a capability range, with pair residuals residualised on the models' accuracies. It needs a new run, not synthesis.
- SourceHas a DCP-style recovery audit been run on an open-ended scientific-hypothesis claim (e.g. Anthropic's ~80% blinded hypothesis-preference result, or a Co-Scientist-style suggestion) rather than a sealed-score engineering-optimization task? This is the scope gap between what Gate 2 formally tests and what Autonomous Scientific Discovery's prior-art question actually asks about.
- SourceDoes the zero-hit recovery bound hold as challenger budget scales — i.e., does giving the Gate 2 challenger more episodes, a stronger harness, or a higher-capability model shrink
p_uppertoward zero recoveries becoming a positive rate, the way AI R&D Autonomy Evaluation (AECI)'s CoBench score barely moves with a 3× token budget but "harness effort... could produce significant further gains"?
- DRACO Benchmark3 open
- WaitThe benchmark is static; the construction pipeline is automatable. Will Perplexity actually refresh it, and does a vendor-built benchmark on which the vendor's own product wins stay credible over time? Sharpened 2026-09-10 by gdpval real world economically valuable tasks (
empirical), which supplies the control case rather than an answer. OpenAI built a benchmark of the same shape — real professional work, expert pairwise grading, an open subset, a vendor-built automated grader — and its own best model finished 8.8 points behind a competitor's, with the paper additionally reporting that its GPT-5-high grader agrees less with human experts on OpenAI outputs. So "vendor-built" is not by itself disqualifying, and the discriminating test is not who wins but which choices the vendor made where a win was available: whose models get a cost row (GDPval's covers OpenAI only), whose product gets its best-configured sampling surface (GDPval sampled Claude through its consumer UI and OpenAI models through a tuned API scaffold), and whether the grader's self-preference is measured or left unexamined. The question this page should now ask of itself is the third one — DRACO's judge is a third-party model, which removes the self-preference channel but not the rubric-authoring one. - SourceRankings are judge-stable but magnitudes aren't — how much do absolute scores move under a non-Gemini judge, and does that matter for cross-paper comparison? Partially answered (2026-08-04) by Yang et al. (2026), with a construct caveat. They hold the candidate pairs fixed and vary only the judge, which is this question's design: on adversarial LLMBar the measured quantity spans 0.463 (Qwen3-1.7B) → 0.900 (GLM-5.1) across ten judges, and no judge leads all four datasets — so the spread from judge choice alone is large, and its direction is dataset-dependent, which is the part that kills naive cross-paper comparison. Two calibrations run the other way, though. Adjacent releases of the same family (MiniMax M2 → M2.7) move accuracy by at most 0.022 and never significantly under paired McNemar, so a routine provider upgrade is a small perturbation; and a single judge held fixed still flips 14.7% of its own verdicts under pure A/B reversal, meaning some of what looks like judge-choice variance is within-judge protocol noise a position-randomization fix would remove. The caveat that keeps this open: their "absolute score" is a judge's agreement accuracy against human preference labels on pairwise items, not a system's normalized rubric score on long-form reports. Rubric-weighted grading of open-ended research output has no gold pairwise label to be accurate against, so the magnitude of DRACO-style score movement under a swapped judge is still unmeasured.
- SourceDoes the production-sourced, expert-rubric method generalize cheaply to non-English, multimodal, and multi-turn deep research?
- WaitThe benchmark is static; the construction pipeline is automatable. Will Perplexity actually refresh it, and does a vendor-built benchmark on which the vendor's own product wins stay credible over time? Sharpened 2026-09-10 by gdpval real world economically valuable tasks (
- SourceDoes the date-driven structure replicate on a second, independently-operated leaderboard, or is it an Artificial Analysis artefact? Three of the twelve benchmarks are constructed or subset by the operator itself and all twelve are run under its single harness, so operator-specific item construction and harness choices are confounded with the structure. Falsifiable directly: run the same two tests on a second board (Open LLM Leaderboard, LMArena, HELM) that shares models but not items, and check whether the first factor's date R² lands near 0.505. Partially answered (2026-09-25) by Desai et al. (
empirical, COLM 2026). Their grid is a second single-pipeline grid, independent of Artificial Analysis: 53 models × 48 HELM- and SafetyPrompts-derived benchmarks, all run by the authors under one zero-shot harness. On it, the capability benchmarks again share one axis, with within-label ρ no higher than across-label ρ (−0.00). So the structure is not an Artificial Analysis artefact. The date half is only bounded. The paper never regresses on release date, but it splits the models at October 2024 (n = 27 / 26) and the non-discrimination holds inside each half (−0.01 / +0.00). Within each half of the release range, capability benchmarks still do not separate. Each half spans roughly a year or more (Llama-2 chat models to September 2024, then October 2024 to at least claude-sonnet-4-5), so this is a coarse window rather than a single release burst. The result is consistent with the date trend being a large share of the first factor rather than all of it. The posed test, the first factor's date R² on a second board, is still unrun. - SourceIs there a confirmatory economic factor, or only an over-extraction? The paper's own first limitation: the three-factor separation is exploratory, no criterion selects it, the loading bootstraps overlap each other, and F1 holds two benchmarks the taxonomy calls non-economic (Terminal-Bench Hard 0.82, AA-Omniscience 0.69). Falsifiable: fit a confirmatory factor model with the economic block pre-specified on an independent, later snapshot and report its fit against the one-factor alternative.
- WaitIs the calendar confound a permanent property of frontier leaderboards or a feature of the 2025–2026 release burst? Every model in this snapshot sits inside a window in which capability rose steeply and jointly; if releases desynchronise across labs, or progress within a generation outruns progress between them, the first factor's date R² should fall without any change to the benchmarks. The trigger to watch is the first snapshot in which date adjustment costs the first factor less than ten points at the configuration level.
- SourceDoes the date-driven structure replicate on a second, independently-operated leaderboard, or is it an Artificial Analysis artefact? Three of the twelve benchmarks are constructed or subset by the operator itself and all twelve are run under its single harness, so operator-specific item construction and harness choices are confounded with the structure. Falsifiable directly: run the same two tests on a second board (Open LLM Leaderboard, LMArena, HELM) that shares models but not items, and check whether the first factor's date R² lands near 0.505. Partially answered (2026-09-25) by Desai et al. (
- SourceDoes any lab publish the horizon at which its pre-release safety evaluations actually ran, against the horizon it claims for the model? No system card in this corpus states an evaluation duration alongside a task-length capability claim, which makes the gap Brown describes currently unmeasurable from outside. A single disclosure — "the longest agentic trajectory in this model's safety evaluation was N hours" — would settle whether the crossing is approaching or already past.
- WaitBrown dates the problem as "not an issue right now" but "quickly becoming" one. The falsifiable version: does a frontier model ship whose published effective task horizon exceeds the interval since the previous frontier release? (Trigger: a model card claiming month-scale autonomous operation, released less than a month after its predecessor.)
- SourceThe internal/external gap is asserted first-party and is invisible from outside by construction — the whole claim is that the model is not available. Is there any external instrument for it at all, or does measuring the disparity require exactly the access whose absence constitutes it? The nearest candidate the corpus has is the Erdős pattern, where an internal result was later reproduced from a public model with enough scaffolding, which measures the gap only in arrears and only where the result is reproducible at all. Partially answered 2026-09-21 by openai navier stokes millennium prize solution — a second instrument exists, it is weaker than it looks, and it is entirely under the lab's control. OpenAI published the output of the withheld model — a writeup, a proof PDF, an Euler PDF and a Lean repository — while the model itself stayed inside. So the answer to "is there any external instrument" is yes in one narrow sense: an artifact can be opened and checked by anyone, which is more than the interview offered and more than the Erdős pattern offers in arrears. Four reasons it is a partial answer and the tag stays
#oq/source. The artifact bounds a result, not a capability — checking that a proof is correct says nothing about what else the model can do, at what budget, or how reliably. It is selected by the publisher, so it measures the best output the lab chose to show, which is the same selection problem the withheld model creates, moved one step downstream. Its most load-bearing element, the claimed Lean formalization, is described in a single clause with none of the checkable discipline, so even the narrow reading depends on the lab's prose. And nobody has done it: as of this compile the repository and both PDFs had not been fetched into this corpus, which makes the instrument's existence the finding and its use the next step.
- SourceDoes the anti-hacking environment survive a capable adversary rather than an incidental one? Every control here is a blocklist or a deletion, and the authors concede self-hosted Git, proxies, dynamic domains, mirrors and direct-IP routes are uncovered. Falsifiable directly: run the hardened environment against a model explicitly instructed to obtain the reference solution, and count confirmed accesses — the paper's zero was measured under ordinary prompting, which is a different quantity. Sharpened (2026-09-10): Ludwig et al. supply the incidental-adversary baseline the comparison needs — under a stock prompt, on two other benchmarks, every one of five open models exploits on 44–82% of its trajectories — so the population being hardened against is not a tail of unusual runs, and a control's residual should be read against that base rate rather than against zero. They also show the prompt is a lever in both directions: a paragraph of prohibition moves judged exploitation by up to 73 points, which means a paragraph of encouragement is a fully specified adversarial arm that costs one prompt edit to run. Nobody has run it.
- SourceHow much of the 21.48-point drop is the block and how much is the instruction? The hardened runs carry an added prohibition on code hosts and version probing, and the trajectories show GLM-5.2 reasoning "the instructions say not to use code hosts or APIs. Let me just implement based on the requirements" and stopping — before any block fired. A two-arm design (isolation without the instruction; instruction without the isolation) would separate a technical control from a compliance effect, and the answer decides whether hardening generalizes to models that do not follow instructions. Partially answered (2026-09-10) — the instruction-only arm now exists, on a different benchmark, and it does most of the behavioural work: Ludwig et al. (NVIDIA,
empirical) append a solution-originality paragraph to the stockmini-swe-agentprompt with no technical control of any kind, and judged exploitation falls 45.1–82.4% → 4.0–10.7% on SWE-bench Multilingual and 44.2–66.1% → 1.5–7.1% on DeepSWE (−41.1 to −73.0 and −37.1 to −62.6 percentage points). So an instruction alone is not a garnish. Three things keep the question open. (i) The two arms measure different quantities — that is a judge's reading of behaviour, this is a score, and an instruction that suppresses narration and overt probing deflates the first faster than the second; the accompanying Pass@1 losses are 4.4–13.3 points on SWE-bench Multilingual and negative for two models on DeepSWE, against the 14–26 points isolation removed here. (ii) The residual is channel-specific and is exactly the block's territory — local Git inspection survives the instruction on all ten model-benchmark cells (3.2–8.6% and 0.3–4.4%) and is the only category that does, while upstream network access goes to 0.0–0.9%, which is the channel a host blocklist covers worst. (iii) Nobody has run the arms on one benchmark with one model set, and neither paper tests a model that declines to comply. The answerable remainder is now narrow and cheap: run vanilla, instruction-only, isolation-only and both on the same 731 instances and report score and confirmed access for each. - SourceWhat is the true defect rate of SWE-Bench Pro's 731 tasks, and does any reissue clear it? This paper repairs 102 (14.0%) from a 119-instance candidate pool assembled from public complaints; OpenAI's audit is reported to have judged ~30% broken (27.4% automated / 34.1% human) and withdrew its recommendation. The two numbers are not comparable as published — one is a repair count from a publicity-bounded sample, the other a prevalence estimate — and nobody has run a systematic review of all 731. Falsifiable: sample uniformly from the 731 and adjudicate against both criteria. Partially answered (2026-10-01): OpenAI's audit (
empirical, now ingested) is closer to a systematic review than this question assumed. An automated filter screened all 731 and flagged 286. Deep review of those 286 found 200 broken (27.4%) by agent-plus-researcher and 249 (34.1%) by five-engineer panels. So the prevalence side has a whole-set estimate, though it is a lower bound: the 445 unflagged tasks were never deep-reviewed, and the filter's recall is unstated. The "does any reissue clear it" half stays open, and the answer is likely no, for a structural reason. OpenAI's low-coverage tests category (4.1–9.4% of tasks, the one that lets incomplete fixes pass) has no counterpart in Zheng et al.'s taxonomy, and a repair that editstest_patchin 17 of 102 cases cannot reach it. Nothing is released per-task, so the two sets cannot be intersected. The remaining falsifiable form is on Benchmark Task Defects (Spec–Test Mismatch). - SourceIs the Pass@1 loss under a solution-originality instruction all removed exploitation, or partly a compliance tax? The instruction forbids reaching any commit, tag or reference outside the current branch, and an agent obeying it literally also gives up legitimate repository archaeology — reading the history of the file it is fixing, which the judge protocol explicitly permits. On SWE-bench Multilingual the instruction costs 4.4–13.3 Pass@1 points; on DeepSWE it costs at most 3.3 and gains up to 3.5. If the loss were pure de-exploitation the two benchmarks should not differ that sharply. Falsifiable cheaply: rerun the principled arm with the local-Git clause narrowed to future references only, holding every other clause fixed, and see whether the SWE-bench Multilingual pass-rate gap closes while exploitation stays flat.
- Expenditure Horizon3 open
- SourceThe existence of a crossing is assumed, sourced to RE-bench and PaperBench rather than measured here. On which frontier optimization problems, and at what budgets, do agent returns actually stop diminishing faster than human returns — the event that both retires this metric and trips the RSP threshold?
- SourceDoes the horizon ranking survive a compute-efficient harness? METR's agents spent 70–90% of budget on experiments with continuously-available nodes, and the claim that shifting curves left leaves horizons roughly unchanged is read off curve shape rather than tested.
- SourceThe hybrid curve — human assisted by agent — is the quantity a lab actually buys, and it is illustrated but never measured; existing evidence points both ways (dominance if humans allocate LLM effort well, degradation if they don't — Becker et al. 2025). What would a runnable hybrid-expenditure experiment look like at a cost anyone would pay, and does it belong on the same dollar axis? Partially answered 2026-09-10 by gdpval real world economically valuable tasks (
empirical) — a runnable design exists, on a different task class. Its Appendix A.2.1 measures the hybrid directly: expert completion time and wage cost, expert review time measured from grading telemetry (109 min, $86), API time and invoiced cost, and the win rate, composed intoE[C_n] = (M_C + R_C)(1−(1−w)ⁿ)/w + (1−w)ⁿ H_C. It costs no more than grading the benchmark already costs, and it is on the same dollar axis. Two things it settles and one it does not. Settled: the hybrid can be worse than unaided — a 12.5%-win-rate model returns 0.53× cost and 0.46× speed, so the direction is a function of the win rate, not a property of assistance. Settled: review is the dominant term, turning 474× into 1.63×. Not settled for this page: the design assumes a fixed one-shot task with a wage-priced human baseline, which is precisely what the NanoGPT setting lacks — there is no per-task win rate for an open-ended optimization problem, and no median-wage estimate for a speedrun contributor. Whether the review-cost term can be estimated for an unbounded optimization task is the remaining gap.
- GDPval Benchmark3 open
- SourceGDPval scores one-shot delivery with no iteration. How much of the gap to the expert closes when the model is allowed the back-and-forth a real assignment gets — and is that gap capability or protocol? Partially answered 2026-09-10 by gdpval real world economically valuable tasks, on the non-interactive half. Three levers the paper does test move it materially without any human in the loop: reasoning effort (GPT-5 low → high, 32.7% → 38.8%), a generic formatting-and-self-check prompt (38.8% → 43.1%, with self-inspection of deliverables going 15% → 97% and black-square PDF artifacts eliminated), and scaffolding (best-of-N=4 with a judge, container GET access). HarnessBank's sealed-split harness evolution adds +9.2 points on a GDPval domain from outside the paper. So a substantial slice is protocol. The interactive half remains untested — the paper lists one-shot non-interactivity as its own limitation and promises future versions — and the falsifiable form sharpens: run the gold subset with a bounded number of expert correction turns and compare the win-rate lift against the 4.3-point prompt-tuning lift, which is the non-interactive ceiling on the same graders.
- SourceThe headline comparison is system-level, not model-level: Claude Opus 4.1 was sampled through its consumer UI to enable file creation, while OpenAI models ran via API with web search, a code interpreter and preinstalled libraries. How much of the 8.8-point gap between Claude Opus 4.1 and GPT-5 high is the model, and how much is the sampling surface? Falsifiable directly — run both through matched API scaffolds on the open 220-task gold subset.
- SourceThe savings analysis covers OpenAI models only — the paper could not obtain cost estimates for Claude, Gemini or Grok — so the model with the highest win rate has no row in the table where win rate is the parameter the payoff turns on. Does the try-n-times ratio for a 47.6% model clear GPT-5's 1.63×, and at what price per deliverable?
- Resolved~~The paper itself is not in this corpus — every figure here is a lecturer's slide reading. What do GDPval's published numbers actually say, and how much of the ASR-damaged quality distribution and cost claim survives contact with the source?~~ Answered 2026-09-10 by gdpval real world economically valuable tasks (arXiv 2510.04374 v1,
empirical). Nearly all of it survives. The quality distribution was accurate as spoken (47.7% acceptable-but-subpar / 22.9% model-better / 29.4% bad-or-catastrophic) but its denominator was wrong on this page — those are shares of GPT-5's losses, not of all outputs. The cost claim was also accurate and was wrongly discarded here as irreconcilable: "1.6× cost, 1.4× speed" is GPT-5's try-n-times row (1.63× / 1.39×) and "under 10% of the expert's salary" is the naive column of the same table. Three things did not survive: the sector rule (>5% each, not a top-5% slice), the O\NET digital filter (a 60% per-occupation inclusion threshold, not a share of tasks kept), and the expert bar (4-year floor, 14-year mean, not "10+ years"). The largest correction is not from the lecture at all but from this page's own reading: the roughly-linear trend is a three-point OpenAI-only fit ending at 38.8%*, and 47.6% is a competitor's score on a different chart with no time axis.
- SourceDoes HCI's ordering survive a different entry-year anchor? The index is defined against the 90th-percentile score in a benchmark's first year in the dataset, so re-anchoring (fixed calendar year, first-release score, or a random-baseline floor) is a cheap falsification test — if the ranking of the ten domains is stable across anchors the instrument is sound, and if legal reasoning and academic breadth swap places it is measuring publication timing. Needs the unpublished score table or a reconstruction.
- SourceAre the ten capability domains independent enough to be trajectories? Benchmark Score Redundancy reports an 84-model × 133-benchmark public matrix as effectively rank-2, which would make most of the ten domains predictable from a handful of probes — and would mean the divergence between mathematics (+53.6) and multimodal (+2.5) in one year is either the strongest counterexample to the low-rank result in the corpus or an artifact of HCI's normalization. Not answerable from the wiki alone: settling it needs the unpublished 393-observation score table, or a reconstruction of it from public leaderboards. Partially answered 2026-09-22 by Zhu (
empirical, single-author preprint) — on the ranking half, and it resolves the either/or this bullet poses into a both. On a hash-pinned single-operator snapshot of twelve benchmarks spanning economic, academic, scientific-coding and long-context blocks, one factor holds 74.5% of common variance, parallel analysis retains exactly one, and hierarchical clustering by correlation distance puts ten of the twelve in a single block. So the domains are not independent as orderings of models: knowing where a model sits on one block places it on the others. That is not a counterexample to this page's divergence, because the two measure different objects — cross-model covariance at one instant against per-domain level and rate over years — and a benchmark at 86.4 closed headroom and one at 39.9 can still order the same models identically. What it does settle is that "ten trajectories" buys ten levels, not ten degrees of freedom in model comparison. And it names the mechanism this page should worry about more than rank. The leading axis tracks release date at logistic R² = 0.505, and removing the date trend costs it 14.9 points at the configuration level (24.1 on one row per base model) — while HCI normalises against the 90th-percentile score in a benchmark's entry year, which is a release-date anchor by construction. The first bullet above (re-anchoring as a falsification test) is therefore the more urgent of the two, and Zhu supplies the method for it: regress every benchmark on release date and re-fit on the residuals. Still open as posed, on two counts — Zhu's twelve benchmarks are not HCI's ten domains (no multimodal, no legal reasoning, no cybersecurity), and nobody has run either test against the unpublished 393-observation table. Partially answered (2026-09-25) by Desai et al. (empirical, COLM 2026), on whether a capability label guarantees a separate axis. It does not. On a 53-model × 48-benchmark single-pipeline grid, reasoning, knowledge and comprehension benchmarks correlate as strongly across labels as within them (within − between ρ = −0.00, robust to a model-era split), and only summarization stands apart. So HCI's ten domain labels cannot be assumed to be ten degrees of freedom. What that grid does not contain is any of HCI's high-divergence domains: no advanced mathematics beyond MATH/GSM8K, no tool agents, no multimodal. The +53.6 against +2.5 divergence is therefore untested, and the question stays open for those domains.
- Does the ability ranking predict anything downstream that accuracy does not? The validity evidence here is
- How fast does a calibrated item bank go stale? Temporal holdout is the paper's largest degradation (ability
- Does calibration survive a population that is not thousands of merges of a few base models? Every item
- LLM-as-a-Judge3 open
- How far can the judge's absolute calibration be trusted for thresholded decisions (ship/no-ship, RSP gating) as opposed to rankings? Partially answered: Norman et al. (2026) show absolute calibration is worse than the reported number implies — the metric practitioners cite (exact-match agreement) systematically overstates chance-corrected reliability by 33–41pp on balanced label sets, so an "85% agreement" judge is really at κ ≈ 0.48 (moderate), and a threshold set on raw agreement is calibrated to an inflated figure. The "rankings are safe" fallback is also bounded — stable across judge-model choice (DRACO) but fragile across benchmark choice (up to 14 rank positions). It does not close the question: it prescribes a pre-deployment checklist (the Minimum Viable Validation Protocol) rather than declaring thresholded judge decisions safe, and defers calibration proper (ECE/Brier) because most providers don't expose logprobs. Sharpened further by Kranti & Vajjala (2026): absolute scores are not just chance-inflated but reference-inflated — with no gold answer in the prompt, judges systematically over-credit incorrect answers, so a threshold set on reference-free correctness is calibrated to an inflated number, and adding the reference flips up to 85% of verdicts (worst in low-resource Telugu). A human study confirms the reference-driven stricter verdicts are more correct, so the reference-free absolute score is genuinely wrong, not merely a different opinion. Demonstrated on three judge models (open-weight Qwen3-32B / Gemma3-27B, closed Gemini-3.1-Flash-Lite) in zero-shot binary QA across English/Arabic/Telugu — magnitudes are model- and language-dependent, not universal. Extended (2026-09-23) with the sharpest separation of the two regimes yet, and a new amplifier. OmniVChat (
empirical) grades one fixed set of 2,773 replies with six judges (three GPT-5.6, three Qwen) under one prompt, parser and scorer, and publishes both reference scales first: re-running the same judge at T=0.1 moves the pooled score by 0.0007, a bootstrap of one judge's pooled score has sd 0.0066, and the six-judge spread is 0.034 — 5.2 bootstrap sds. So absolute level is judge-determined well beyond noise. The ranking fallback holds on the same data and holds strongly: pairwise Pearson across the 17 subcategories is at least 0.936, Spearman at least 0.860, all six judges pick the same weakest subcategory, and a two-way decomposition puts only 1.1% of variance on judge severity against 95.7% on subcategory difficulty. That is this question's cleanest both-halves datum — rankings safe, thresholds not — from an instrument that reports the noise floor. The new part is an amplifier this page had not named: a gate. OmniVChat's rubric is tiered, and a criterion in an earlier tier blocks every later one, so judges agreeing on 87.8–96.5% of individual criterion verdicts still produce different final scores on 11.3–34.2% of replies. Weight-aggregation averages a judge error; a prerequisite structure multiplies it. Any thresholded decision taken on a gated rubric therefore needs its judge validated at the criterion level of the gating tier specifically, not at the aggregate. Still open for the same reason as before: all agreement here is judge-against-judge, no human re-annotation is reported, and κ across pairs runs as low as 0.532. The deferred calibration half becomes measurable (2026-10-01). Hsiao (2026) (empirical) gets calibration numbers for closed judges with no logprobs, by having the judge state a 0–100 confidence. Baseline flagship AECE is 8.0–17.7% on AggreFact. An overconfidence advisory plus single-call self-debate brings it to 3.3–13.5% on all 10 judges. A threshold on a judge's absolute score can therefore now be audited, but only against a calibration figure measured under that exact prompt, since one advisory sentence moves AECE by up to 7pp. The ship/no-ship case measured directly (2026-10-01): Li (2026) (empirical) finds frozen judges that rank 35 SWE-bench agents at Kendall 0.71–0.79 still declare 8 of 60 version upgrades that execution intervals cannot establish, and old-version calibration makes comparisons worse in 59 of 60 units. Rankings safe, thresholds not, now on a release decision. - SourceCan a fully-autonomous, well-aligned rubric+judge pipeline match expert-authored rubrics, removing the human bottleneck DRACO still relies on? Partially answered (2026-08-04) by Chen et al. (2026) — and the partition it draws is the useful part. CalibratedRubric removes the expert from filtering, weighting and sizing the bank, with no human labels and no gold judge required: measurability filtering lifts human-gold κ 0.604 → 0.743 on JudgmentBench, and IRT selection reaches the target rank correlation with 49 instead of 131 rubrics. It does not touch authoring or validating them — "we take the candidate pools as given," the method "cannot recover dimensions absent from the candidate pool," and measurability is stated outright to be "necessary but not sufficient for substantive expert endorsement." So the automatable half is the psychometric half; the half DRACO spends 26 experts on — deciding what a good report is in the first place — is untouched. Two measurements bound how close the autonomous pipeline gets. (i) Against a human reference ranking on FinResearch, the task-adaptive scorer and the plain binary baseline achieve the identical Spearman ρ = 0.8833 — so the paper's own front end buys discrimination and cost, not external validity, and 0.8833 is what this generation of automated rubric grading matches human judgment at. (ii) The judge–human gap survives the filter: LLM judges assign positive labels at 55.6–62.9% against a 47.1% human-gold rate, "a systematic judge–human mismatch that measurability filtering does not fully eliminate" — a directional generosity bias in exactly the direction Reference-Free Judge Over-Crediting measures.
- SourceWhen does judge-lineage bias actually flip a result, versus merely shift magnitudes? Partially answered (2026-08-12), and the answer is "it flips" — on a detection task, at low evidence: Greptile (
case-study) has two frontier models review two 500-PR corpora, one authored by each family, and the ranking of which reviewer is better reverses between the corpora (Opus 53.7 vs GPT 62.0 on Claude-authored PRs; Opus 60.0 vs GPT 50.5 on Codex-authored PRs) while the reviewers' pooled averages sit 0.6pp apart. So on this task lineage does not shift a magnitude — it decides the rank, and a leaderboard built on either corpus alone would report the opposite winner. Three things keep it partial: the grading task is bug detection rather than quality scoring, so the bias surfaces as recall rather than as generosity and may not transfer to rubric grading; the ground truth is vendor-built with no released artifact, no agreement statistic and no validation of the matching judge; and one arm's review prompt was tuned against the outcome metric, which cannot manufacture a crossover but does make the magnitudes soft. The clean version of the experiment — a third-family reviewer across both corpora, which would separate lineage from stylistic fit — is named on that page and has not been run. A near-miss worth recording (2026-09-23): omnivchat's six-judge study looks like the experiment and is its exact transpose — it varies the judge (three GPT-5.6, three Qwen) while holding the responder fixed (one set of Gemini-3.7-Flash replies). What that design can answer, it answers cleanly and negatively: pooled severity does not split by judge family (gpt-5.6-sol is the most generous at 0.678 and gpt-5.6-luna nearly the strictest at 0.646, with the two Qwen judges straddling them), so judge lineage is not a severity axis here. What it cannot answer is this question, because no Qwen-authored reply is in the set — and the paper's main table has a qwen3.7-max judge scoring five Qwen comparators plus a Qwen-derived proposed model. The Greptile design and this one are complements, and nobody has run both arms on one task.
- How far can the judge's absolute calibration be trusted for thresholded decisions (ship/no-ship, RSP gating) as opposed to rankings? Partially answered: Norman et al. (2026) show absolute calibration is worse than the reported number implies — the metric practitioners cite (exact-match agreement) systematically overstates chance-corrected reliability by 33–41pp on balanced label sets, so an "85% agreement" judge is really at κ ≈ 0.48 (moderate), and a threshold set on raw agreement is calibrated to an inflated figure. The "rankings are safe" fallback is also bounded — stable across judge-model choice (DRACO) but fragile across benchmark choice (up to 14 rank positions). It does not close the question: it prescribes a pre-deployment checklist (the Minimum Viable Validation Protocol) rather than declaring thresholded judge decisions safe, and defers calibration proper (ECE/Brier) because most providers don't expose logprobs. Sharpened further by Kranti & Vajjala (2026): absolute scores are not just chance-inflated but reference-inflated — with no gold answer in the prompt, judges systematically over-credit incorrect answers, so a threshold set on reference-free correctness is calibrated to an inflated number, and adding the reference flips up to 85% of verdicts (worst in low-resource Telugu). A human study confirms the reference-driven stricter verdicts are more correct, so the reference-free absolute score is genuinely wrong, not merely a different opinion. Demonstrated on three judge models (open-weight Qwen3-32B / Gemma3-27B, closed Gemini-3.1-Flash-Lite) in zero-shot binary QA across English/Arabic/Telugu — magnitudes are model- and language-dependent, not universal. Extended (2026-09-23) with the sharpest separation of the two regimes yet, and a new amplifier. OmniVChat (
- LLM-Judge Validation4 open
- SourceThe MVVP validates reliability and bias; calibration proper (ECE/Brier) is deferred for lack of provider logprobs. How far can a judge's absolute score be trusted for a threshold once confidence calibration is measurable? Partially answered (2026-10-01): calibration is measurable without logprobs, and the measurement belongs to the prompt as well as the judge. Hsiao (Cisco, 2026) (
empirical) computes adaptive ECE and a score-spread statistic from a stated 0–100 confidence on 10 closed flagships (GPT, Claude, Gemini) on AggreFact, so the premise of the deferral no longer holds. The answer to "how far" is conditional. With a baseline confidence rubric, flagship AECE runs 8.0–17.7% (Gemini 2.5 Pro worst). One overconfidence-advisory sentence cuts it on 10/10 models, by 1.1–7.1pp, and adding single-call self-debate leaves the full recipe at 3.3–13.5%. So a threshold is trustworthy only against a calibration number measured under the same judge prompt. The prompt also moves the hard verdict, generation-dependently: the same recipe costs pre-2025 judges −1.8pp balanced accuracy and post-2025 judges nothing. Still open because nothing here is a thresholded decision. The tasks are binary faithfulness and quality labels with human gold, and no ship/no-ship or gating threshold is set or audited. AECE is also measured in-distribution, so whether it holds for the items a gate would actually see is untested. Partially answered on the thresholded half (2026-10-01): Li (empirical) audits the ship/no-ship call itself on 20 SWE-bench agent-version pairs against execution labels. Judge-only intervals promote 8 of 60 reference-inconclusive pairs, and old-version calibration is worse than none in 59 of 60 units, so a gating threshold has to be calibrated on the version being gated. - SourceAll judges were run with thinking suppressed. Does reasoning-on flip the consistency–bias paradox, or just move the numbers? Partially answered (2026-09-22) on a neighbouring property, and the answer is unwelcome: Kohli re-runs a nine-judge, seven-family panel on the same 1,000 MNLI items with chain-of-thought and finds mean pairwise error correlation rises — φ̄ 0.391 → 0.456, Kish n_eff 2.18 → 1.94, panel accuracy 72.0% → 69.2% — while prompt rewording, label-order reversal and temperature 0.5 leave all three flat. Shared reasoning amplifies shared error. That does not touch the consistency–bias paradox itself, which is a within-judge reliability-versus-validity claim and would need a position-flip rate measured with reasoning on; it does establish that reasoning-on is not a neutral setting for judge panels, and that a validation protocol which certifies a judge with thinking suppressed is not certifying the configuration a panel would actually be run in. A second neighbouring datum (2026-10-01), on calibration rather than bias: Hsiao varies an in-output reasoning block (not a thinking channel) on 10 flagship judges. Free-form reasoning significantly narrows the confidence distribution on 8/10 judges, which fits a one-directional argument piling probability mass onto the committed answer. Self-debate widens it on 6/10, and it costs pre-2025 judges balanced accuracy (−1.0pp vs the advisory-only prompt) while slightly helping post-2025 ones (+0.6pp). Same direction as Kohli: the reasoning configuration is not neutral, and its effect depends on the judge's generation. The consistency–bias paradox itself is still unmeasured with reasoning on.
- SourceHosted endpoints drift silently between provider updates. How stable are these agreement/bias profiles over a longer horizon than five weeks — and should judge validation be continuous rather than one-shot? Partially answered (2026-08-04) by Yang et al. (2026), on the announced-upgrade sibling of the question rather than silent drift. Across four released MiniMax generations (M2 → M2.1 → M2.5 → M2.7) on four datasets, adjacent accuracy moves at most 0.022 and not one of the nine adjacent McNemar tests reaches even uncorrected
p < 0.05— and because those tests are paired on parse-shared examples, this is stability of the individual verdicts, not merely of the aggregate. So a deliberate version step at the top of the capability range is a small perturbation to the agreement profile. Three things keep this open. (i) These are version-labeled releases you can pin, not the unannounced same-endpoint drift the question is about — nobody has re-measured a fixed endpoint over months. (ii) Stability of accuracy is not stability of bias: position-flip rates still span 0.117–0.147 across the MiniMax releases, and no one tracked whether that band moves within a single version. (iii) The finding runs the other way on the parameter axis, where a step does move things and can move them down (Qwen3 1.7B → 32B costs 0.065 on Judge's Verdict), so "upgrades are safe" is not the lesson — "upgrades are a measurement event that must be re-validated, and the null case is the lucky case" is. That is an argument for continuous validation, from the direction of the one axis where it was cheap to check. - Does the paradox generalize beyond position bias — i.e., are there other biases (self-preference, lineage) that high test-retest also masks? Partially answered by Kranti & Vajjala (2026): yes — reference-presence sensitivity is a large one. Their temperature-0 judges (thus perfectly reproducible) systematically over-credit incorrect answers in no-reference settings, an invalidity invisible to any reliability metric until a gold answer is added, which flips up to 85% of verdicts and lifts human-alignment sharply (e.g. Gemma3-27B 0.34→0.85 NR→RV). It does not close the question — self-preference and lineage remain unisolated (their design deliberately overlaps judge and generator but doesn't attribute the effect), and it is a different bias on a different (multilingual QA, three specific judges) setup, not a re-run of this study's position-bias protocol. The self-preference half then largely resolves negatively (2026-08-04), via Zhou (2026): errors optimized against a self-judge transfer to judges from other families that were never in the loop (Llama 0.480 → 0.568, Gemma 0.764 → 0.918) and to same-family judges 3.5× larger (still 77%), with a three-family unanimous-accept ensemble passing 55% and judge acceptances pairwise correlated at φ = 0.29–0.38. So lineage amplifies — the self-judge is the worst single cell at 0.906 — but is not the mechanism; the bias is a shared property of the candidate-conditioned channel. This also adds a bias class no reliability metric on this page can reach, because it is not a property of the judge at all: the same judge, unchanged, is valid before optimization (discrimination 0.31) and invalid after (0.09), so any one-shot validation — including the full MVVP — certifies a judge that will be true only until something starts optimizing against it. That is a direct argument for the continuous-validation question two bullets up. A field datapoint on the same half, 2026-09-10, from gdpval real world economically valuable tasks (
empirical): OpenAI's GDPval grader — GPT-5-high, scoring deliverables against a human professional's — is reported by its own authors to agree less with human expert graders on outputs from capable OpenAI models, cited to Panickssery et al. (2024). That is lineage effect surfacing inside a shipped validation number rather than in a lab protocol, and it is invisible to test-retest by construction, since a judge is perfectly self-consistent about preferring its own family. It does not reopen the negative resolution above — the paper reports it observationally with no controlled arm, and agreement also falls with model capability generally, which nothing in the paper separates — but it is the first instance in this corpus of a vendor publishing the effect against itself.
- SourceThe MVVP validates reliability and bias; calibration proper (ECE/Brier) is deferred for lack of provider logprobs. How far can a judge's absolute score be trusted for a threshold once confidence calibration is measurable? Partially answered (2026-10-01): calibration is measurable without logprobs, and the measurement belongs to the prompt as well as the judge. Hsiao (Cisco, 2026) (
- SourceDo the calibrated rates hold for instruction-tuned production models? Everything here runs on open-weight base models with known or inferable training corpora; the parrot demonstration is the only instruction-tuned experiment and it is deliberately degenerate. Whether matched controls can be constructed at all for a closed production model — where the cutoff is approximate and the corpus unpublished — is the gap between this method and the deployment setting where the copyright claims actually land.
- SourceDoes the near-verbatim ε-ball resolution change the answer, or just the accounting? The near-verbatim test is a strictly broader instance of the same hierarchy, computed with a beam-search lower bound rather than exactly, and its floors differ from verbatim ones by a decade on at least one pair (Collins: 10⁻²⁷ verbatim vs 10⁻²⁶ near-verbatim). Whether calibrated rates move as much as thresholds do is not reported.
- WaitWhat is the right realistic query budget? 10⁵ is picked "for illustration purposes," and the k-CBS result shows a smarter decoder shifts the frontier at fixed cost. Any threshold that determines whether text counts as extractable in a legal or policy setting needs a defensible budget, and there is no principle here for choosing one.
- SourceDoes re-instrumentation generalize past reproducibility? CORE-Bench Hard was chosen precisely because it has a direct human counterpart, clean OOD axes, and multiple practical dimensions. Whether the six-axis treatment yields comparable signal on benchmarks without those properties (e.g. closed-form reasoning benchmarks with no human-workflow analog) is untested. Partially answered (2026-09-22) — on exactly that class of benchmark, by a different treatment: ATLAS re-instruments five benchmarks with no human-workflow analog at all (WinoGrande, TruthfulQA, HellaSwag, GSM8K, ARC) and does recover separation the accuracy number had lost — 23–31% of models move more than 10 rank positions, and models tied on accuracy differ by thousands. So re-instrumentation generalizes past reproducibility in the weak sense that something beyond accuracy is recoverable there. It does not answer the strong form: the recovered signal is a re-scoring of the same construct, not any of the six axes, and the four that need a human counterpart or a scaffold sweep remain untested on closed-form benchmarks. Stays
#oq/source, with the remaining test named: run the reliability, efficiency and model-vs-scaffold axes on a closed-form reasoning benchmark. - SourceIs the human-uplift result real or a demand effect? The reproducers are the paper's own coauthors, there is no ground-truth correctness, and n = 20 papers / 5 participants. The 2.11× speedup is statistically significant but the authors themselves cannot rule out participant bias — an independent, blinded replication is the missing evidence.
- SourceWhich non-accuracy axis actually predicts deployment value? The paper measures six axes but does not rank them by decision-relevance for a downstream deployer. If you can only measure one beyond accuracy, is it reliability, efficiency, or scaffold contribution — and does the answer depend on the use case? Partially answered (2026-08-04) — a seventh candidate rather than a ranking: Leni proposes loop telemetry and gives it the most direct claim to decision-relevance any axis here has made. Because its verification loop is fully instrumented, the measured catch/fix/false-alarm rates convert straight into marginal returns on the next engineering decision — raising the catch rate is worth up to +8 pp, raising the fix rate at most +0.5 pp — so the axis does not merely separate systems, it names which component to fund. It also answers the "does it depend on the use case" half affirmatively and specifically: the argument for keeping the checkpoint at all is that its value concentrates where a reliability SLA's tail sits, which is a use-case-conditional claim by construction. Still open as posed, because no source has ranked the axes against each other, and this proposal comes from a vendor instrumenting its own system.
- SourceCan the model-vs-scaffold decoupling be made routine? The oracle-router result (every task solvable by some scaffold → 100%) implies large headroom from scaffold routing, but requires per-task oracle knowledge. Whether a practical router can approach the oracle without it is open, and would turn a measurement into a capability. Partially answered (2026-08-04) — on an adjacent axis: Leni ships a practical router and it pays. A 0.5B step-type classifier dispatches each step across two model families (cheap models for classification, frontier reasoning for multi-hop synthesis, strong grounding for vision), and internal estimates credit it with ~4 pp of GAIA accuracy at net-negative cost — cheap steps subsidise extended reasoning on hard ones. So routing is deployable at a price low enough to run on every step, which was the practical objection. It does not settle the question as posed, on an axis mismatch that matters: this routes models per step inside one fixed scaffold, not scaffolds per task, and no oracle comparison is run, so what fraction of available headroom the router captures is unmeasured. Vendor-authored, internal single-run attribution.
- SourceOnce forecasters pass the human reference class, what anchors the scale? The Brier index keeps reading past superforecasters, but its remaining headroom mixes forecaster skill with the questions' irreducible uncertainty, and the human baseline was what implicitly separated them. Falsifiable directly: estimate a question set's aleatory ceiling independently — ex-post outcome variance, or the asymptote of an ensemble-of-ensembles over all submissions — and check whether "how much better than the best humans" becomes a quantity rather than an ordering. Nobody in the corpus has tried it; FRI's own answer is to re-elicit the humans.
- SourceDo "living benchmarks" outrun their own maintenance? v1.1 and OOD are to be updated as new validity threats surface via log analysis, which the authors note is non-exhaustive. Whether continuous log-analysis-driven maintenance is sustainable — or itself becomes a Goodhart target once developers know the rubrics — is unexamined. Partially answered (2026-09-10) — the first maintenance pass in the corpus observed under load, and it does not keep up: SWE-Bench Pro Verified (
empirical) is exactly this operation performed on a 731-task benchmark, and three properties of it answer the sustainability half. (1) The maintenance is reactive, not systematic. The candidate pool was assembled from public issue reports — GitHub issues, review repos, Hugging Face feedback — so only 119 of 731 instances (16.3%) were ever examined, and a defect nobody complained about was never inspected. (2) The cost binds explicitly. LLM-assisted filtering drafts the fixes but human experts make every final edit, with trial re-runs and iterative repair, and the authors state they triaged to "completely broken instances" "given the substantial review cost." (3) It closes well under half the estimated gap. 102 instances repaired (14.0%) against a separately reported ~30% defect estimate for the same task set (OpenAI's audit, unverified secondary — this vault's_system/research-channels.md, not this paper, carries the number). A fourth finding is one the CORE-Bench framing does not anticipate: the execution environment needs maintenance too, and it degrades adversarially rather than by wear — the paper's own limitations concede its domain blocklist cannot cover self-hosted Git, private proxies, dynamic domains, mirrors or direct IP, so this axis of maintenance has an opponent. What stays open is the Goodhart half: nobody has yet observed developers optimizing against a published refinement rubric.
- SourceDoes re-instrumentation generalize past reproducibility? CORE-Bench Hard was chosen precisely because it has a direct human counterpart, clean OOD axes, and multiple practical dimensions. Whether the six-axis treatment yields comparable signal on benchmarks without those properties (e.g. closed-form reasoning benchmarks with no human-workflow analog) is untested. Partially answered (2026-09-22) — on exactly that class of benchmark, by a different treatment: ATLAS re-instruments five benchmarks with no human-workflow analog at all (WinoGrande, TruthfulQA, HellaSwag, GSM8K, ARC) and does recover separation the accuracy number had lost — 23–31% of models move more than 10 rank positions, and models tied on accuracy differ by thousands. So re-instrumentation generalizes past reproducibility in the weak sense that something beyond accuracy is recoverable there. It does not answer the strong form: the recovered signal is a re-scoring of the same construct, not any of the six axes, and the four that need a human counterpart or a scaffold sweep remain untested on closed-form benchmarks. Stays
- SourceDoes the r = 0.816 sim-to-real correlation survive on a set of comparable planners? Leave-one-out puts it at 0.421 (p = 0.500) once the weakest of six models is dropped, and only two of seven rows clear p < 0.05, so the fidelity claim may be entirely the strong-vs-weak spread. Settling it needs a run over ten or more frontier-tier planners with the weak tail excluded.
- SourceDoes the multi-agent / single-agent crossover survive real execution? In simulation the advantage falls from +0.302 at 16k to +0.007 at 128k and reverses on 82% of model-problem pairs — but the simulated single agent suffers only compression loss, with no attention degradation, distraction, or long-context recall failure priced in, and the paper never runs the single-agent comparison for real. A real 128k single-agent-vs-multi-agent arm on the same tasks would settle whether 128k is the true crossover or an artifact of a generous single-agent model.
- SourceIs the transfer-coverage result about orchestration or about the penalty? All measured separation between a strong and a weak planner vanishes when λ goes to 1 or omitted transfers are auto-completed, and λ = 0.5 was chosen for discriminative power rather than fitted to observed handoff loss. What would settle it: an execution study measuring how much downstream quality an actually-omitted handoff costs in a real framework.
- SourceHow much does augmentation distort the distribution it claims to represent? Is there a measurable representativeness loss between raw queries and augmented tasks?
- SourceDifficulty-by-thumbs-down biases toward current failures — does that make the benchmark a moving target that flatters the next model trained on those failures? Instantiated, not answered (2026-08-13): shopify sidekick continual learning loop runs exactly this loop in production — hard negatives sampled by a judge's low score, repaired, replayed, and folded into weights daily via SFT then GRPO with that same judge as the reward — and reports no held-out split, no independent difficulty proxy, and no arm that would separate "the model got better" from "the model got better at the sampler." So the hazard has a deployment now and still no measurement. The falsifiable form sharpens: hold out a slice of production traffic sampled by a signal the training loop never sees (user thumbs-down, or a judge from a different family), and compare the improvement measured there against the improvement measured on the in-loop judge. Half-answered (2026-09-23), and the half that got run came back clean. omnivchat (
empirical) runs the same shape — a training corpus and a benchmark synthesized by one engine from one set of subcategory configurations, optimized with the grading judge as the reward — and it does hold out the distribution slice this question asks for: 360 human phone recordings, improvised by performers from a plain-language guide, produced by no part of the training loop. The two improvements match closely: 0.465 → 0.652 on the in-loop synthetic benchmark, 0.402 → 0.632 on the recorded probe, and 0.437 → 0.634 on a language-matched synthetic control. So on this instance, distributional self-flattery did not happen. The judge half is untouched and the authors say so: the sameqwen3.7-maxgrades reward, synthetic benchmark and recorded probe, and their six-judge sensitivity study "does not remeasure the OmniVChat-RL gain from 0.465 to 0.652 with other judges." The falsifiable residue is now exactly one experiment — re-grade both the in-loop and the held-out slice with a judge from a different family — and it is cheap, since the reply sets already exist. Note also what a matched gain does and does not buy: the two instruments agreed on the delta and disagreed on the ranking of two models from the same vendor family, so a held-out slice validates an improvement claim without validating a leaderboard. - SourceCan the privacy pipeline (no human sees raw queries) be trusted/audited well enough for regulated domains (medicine, law) where the source traffic is most sensitive?
- SourceDoes the calibration-window stability (1-day ≈ 4-week) hold for an agent earlier in development, or one changing faster? The paper explicitly attributes the null result to a "relatively mature development stage" and does not test an early-stage or high-churn agent.
- SourceDo historical caching and adaptive testing transfer across agent families and tolerate a short calibration window the way the deployed fixed subsets do? Both validations in §6 were run only on the difficulty-stratified fixed subset.
- SourceDoes difficulty stratification's win over clustering-based selection generalize beyond one production agent's response history, or is it specific to the narrower response-pattern diversity the paper itself proposes as the explanation for its contradiction of Polo et al. (2024)?
- SourceDoes the two-stage pipeline transfer beyond binary QA? Calibration + sensitivity are demonstrated on binary correct/incorrect factual QA. Do the same probes diagnose reference-sensitivity for graded rubrics, long-form generation, or multi-turn agent transcripts, where "the reference" is a rubric rather than a gold answer? Sharpened rather than answered (2026-08-13): a deployment that needs the answer now exists and does not supply it. multilingual multi agent planning failures (
empirical) runs a frontier judge as a six-way classifier over agent plans in eleven languages down to 0.0008% of Common Crawl — scarcer than the Telugu where the probes here collapse — and validates it against 117 human verifications stratified by category and never by language (κ = 0.860, macro-F1 = 0.906), while the finding it supports is a per-language gradient. So the concrete next form of this question is narrower and free to run: does the C1/C2 calibration gap, measured per language, predict per-language classification agreement on a judge whose reference is visible? The data for the second half of that comparison already exists and has never been cut that way. - SourceIs over-crediting a knowledge gap or a generosity prior? In low-resource Telugu the judge flips the same extracted answer once a reference appears — is the NR generosity driven by insufficient task knowledge (calibration failure) or by a default lean-toward-CORRECT that a reference overrides? The two have different fixes (better judge vs. always supply a reference). Partially answered (2026-08-04) by Zhou (2026): neither, at least on English math — it is candidate anchoring. The same Qwen3-4B judge that accepts 0.91 of wrong answers when scoring a shown candidate solves those problems itself at 0.93 accuracy and drops to 0.012 false positives once required to commit its own answer first, with the candidate still fully visible. So the knowledge is present and a lean-toward-CORRECT is not the mechanism either: conditioning on the candidate is. Corollary 1 turns this into a test any deployment can run — a measured FPR above
1 − solve-acccertifies the verdicts as anchored, and Corollary 2 prices the excess in bits (0.719 against a 0.07 ceiling ⇒ ≥ 1.2 bits of candidate leakage). It does not close the question for this page's setting: Zhou's judges are mid-size open-weight models on exact-matchable grade-school math, where a knowledge gap is implausible by construction, and the low-resource Telugu case — where the judge may genuinely not know the answer — is exactly where anchoring and ignorance are hardest to separate. - SourceHow much does self-/same-family overlap contribute? The design deliberately overlaps generator and judge (Qwen3-32B self-judging; Gemini/Gemma family) and Qwen self-judging is the most reference-sensitive, but the paper does not isolate a self-preference effect from a low-resource effect. When does judge–generator lineage amplify reference-free over-crediting? Partially answered (2026-08-04) by Zhou (2026): lineage amplifies it but does not cause it. The self-judge is the worst cell (FPR 0.906 post-self-play) but errors optimized against it transfer to judges from other families that were never in the loop — Llama-3.1-8B 0.480 → 0.568, Gemma-3-12B 0.764 → 0.918 — and a three-family ensemble still accepts 55%, with acceptances pairwise correlated at φ = 0.29–0.38 (581 unanimous accepts where independence predicts ≈ 497). The residual: this measures transfer of self-play-manufactured errors, not a controlled self-preference ablation, and Kranti's question is about reference-sensitivity on organic responses.
- SourceDoes the effect shrink with stronger or thinking-enabled judges? All judges are ≤ mid-tier at temperature 0 with no reasoning channel. Would a frontier reasoning judge over-credit less in NR, or just flip at different rates? Partially answered (2026-08-04): scale alone does not fix it, and the reasoning half is confounded. Zhou sweeps judge size to 14B (3.5× the policy) and every judge's discrimination collapses under optimization — 14B still accepts 77% of the hacked errors, and the strictest base judge is driven to the highest post-hoc false-positive rate. Reasoning is not cleanly separated: the recompute prompt (solve it yourself, reject when uncertain) leaves FPR at 0.719, but the blind-solve verifier that reaches 0.012 runs with reasoning enabled, so thinking-on and de-anchoring co-vary in the arm that works. A frontier reasoning judge scoring a shown candidate remains untested.
- SourceDoes commit-first survive the loss of an exact-matchable answer? The de-anchoring fix accepts only when the judge's independently committed answer exactly matches the candidate's — which is what makes its false-positive rate provably bounded by
1 − solve-acc. Zhou names extending commitment to open-ended outputs ("committed rubrics, executable tests") as future work, and that is the regime the vault actually cares about: a taste, quality or report-grading reward has no exact match to accept on, and a committed rubric compared for partial agreement reintroduces exactly the graded, plausibility-shaped judgment the fix removes. Does a commitment survive contact with an output that can only be scored by degree, or is the whole result a property of tasks with a checkable final token?
- SourceDoes the two-stage pipeline transfer beyond binary QA? Calibration + sensitivity are demonstrated on binary correct/incorrect factual QA. Do the same probes diagnose reference-sensitivity for graded rubrics, long-form generation, or multi-turn agent transcripts, where "the reference" is a rubric rather than a gold answer? Sharpened rather than answered (2026-08-13): a deployment that needs the answer now exists and does not supply it. multilingual multi agent planning failures (
- SourceDoes the RLHF length-bias hypothesis replicate when tested against base (non-instruct) model variants directly? If verbose generation were primarily pretrained, base-model verbosity differences should match instruct-model differences.
- SourceWhat problem characteristics predict prompt sensitivity? An automated classifier would make scale-specific prompting deployable.
- SourceHow does the overthinking effect interact with tool-using agents? If brevity helps large models but tools require structured reasoning, the optimal prompt is not uniformly brief.
- SourceDo reasoning models (o1, DeepSeek-R1 style) exhibit different overthinking dynamics than instruct models? Their trained behavior is explicitly to generate long CoT — does brevity intervention hurt them? Partially answered 2026-09-22 by thinking hard not smart (Fan et al.,
empirical, seven reasoning models: DeepSeek-R1-Distill-Qwen-7B/14B, Qwen3-8B/14B/32B, DeepSeek-V4 Flash/Pro) on the dynamics half only, and at a different granularity — across questions rather than within one answer. Reasoning models do show a distinct failure: given one shared budget over N scored questions they spend 32% of all reasoning tokens on questions they failed in an independent 40,960-token attempt, and effort–difficulty correlation decays under pressure (+0.33 → +0.11 as N goes 5 → 20), so the overspend is reactive rather than chosen. The brevity-adjacent intervention available there — a "skip hint" granting permission to abandon a question whose cost outweighs its points — is not harmful: it raises coverage by +0.05 to +0.08 and cuts the zero-token rate at every exam length. What remains untested is this question's literal form: nobody has run the brief/direct word-cap intervention on a reasoning model's per-question accuracy. - SourceIs BoolQ's functional-elaboration exception a clean taxonomy boundary, or does every task type have a context-dependent optimal length?
- Skill Lift3 open
- SourceDoes Skill Lift survive an evaluation set the skill's author did not generate? The discriminating experiment is cheap and entirely within NVIDIA's reach: build a held-out task set from product documentation independently of the skill, re-run the same Tier 3 ablation, and publish both numbers. If the lift collapses,
benchmarks.jsonmeasures skill–eval agreement rather than agent capability. - SourceWho or what grades the five dimensions, and has that grader been validated? Nothing in the post names the judge model or reports agreement statistics; LLM-Judge Validation's Minimum Viable Validation Protocol is the bar. Falsifiable by NVIDIA publishing the grading configuration alongside
benchmarks.json. - WaitDoes per-skill lift decay as models improve? The catalog is re-evaluated continuously against a pinned-commit history, so a second snapshot on a newer model generation would show whether skills supplying not-inferable product facts hold their lift while workflow-scaffolding skills erode — the first direct measurement of Harness Shrinkage as Models Improve on context artifacts. Trigger: a
benchmarks.jsonsnapshot on a subsequent frontier model release.
- SourceDoes Skill Lift survive an evaluation set the skill's author did not generate? The discriminating experiment is cheap and entirely within NVIDIA's reach: build a held-out task set from product documentation independently of the skill, re-run the same Tier 3 ablation, and publish both numbers. If the lift collapses,
- WaitIs the 4-month doubling a stable regime or a local steepening? The trend's shape (exponential vs S-curve) is undetermined. Sharpened (2026-07): AISI adds that the doubling rate itself is budget-dependent — the same cyber suite doubles ~60% faster measured at 50M than at 2.5M tokens/task — so the headline rate is undefined without naming the eval budget. The stability question is now entangled with a budget question, not just an exponential-vs-S-curve one.
- Time horizon is measured on task baskets that themselves saturate; what replaces them once weeks-long tasks become measurable — and who builds those tasks? Partially answered / reframed (2026-07): Nadgir et al. argue don't replace — re-instrument: they take CORE-Bench (cited above as saturating in 15 months) and show it still discriminates agents along six non-accuracy axes after accuracy saturates, so "what replaces a saturated basket" can be "keep it and measure differently" rather than "build a harder one." This addresses accuracy saturation, not the length-metric saturation this page's basket faces, and does not answer who builds the next weeks-long tasks — so it reframes the retire reflex without closing the question. Second partial answer, from METR itself (2026-08): Expenditure Horizon is the metric's own authors proposing a third option — neither "build a harder basket" nor "re-instrument the old one," but change the object measured: drop the basket entirely for a single frontier optimization problem humans are still actively improving, and read capability off where the agent's dollar curve crosses theirs. It answers "who builds the tasks" by not needing tasks to be built — the speedrun leaderboard is the instrument and it refreshes itself. But it buys that with two costs the basket doesn't have: it needs continuously-scored, smooth-returns problems (a narrower class than a general task basket), and it needs a per-problem estimate of the returns to human labour, which on NanoGPT took two contributor interviews, an LLM judge over 82 PRs and a bootstrapped correction factor to produce — and still lands on a number METR calls "highly uncertain."
- SourceHow much of the frontier rate reaches a buyer who doesn't switch? Epoch measures ~13×/yr for the cheapest model at each score. Ramp's realized blended per-token price fell about 3×/yr, and that mixes tier shifts with list-price cuts. Falsifiable: keep a fixed agentic workload and a buyer panel's actual model choices, record cost per task every quarter, and compare it with the frontier rate for the same score.
- WaitIs the "fastest at SOTA" decay a durable pattern or an artifact of the fit? It shows up on three of five primary benchmarks. The game benchmarks run the other way. The time curvature often failed to estimate and was then forced to linear. There are no intervals. Falsifiable when Epoch updates the dataset: score levels that first become SOTA in 2026 should fall about 66% per quarter in their first quarter and about half that rate two years on. Trigger: the next data release from github.com/droodman/inference-cost.
- SourceWhy is SWE-bench Verified the slowest-falling benchmark (27.5% per quarter)? Is it rising tokens per task on agentic work, a distorted ceiling from defective and leaky tasks, or slower competition in coding? Falsifiable: rerun the same truncation analysis on SWE-Bench Pro Verified's repaired, isolated task set and see whether the rate moves toward the ~47% average.
- Typed Decision Verifiers5 open
- SourceThis study reports per-system AUC on a shared set but no pairwise error correlation between systems. Does typed-decision-vs-typed-decision error correlation differ from typed-decision-vs-LLM-judge correlation, and is either lower than the judge-to-judge correlation this wiki has measured elsewhere (ρ̄ ≈ 0.2–0.97 across the wiki's judge-dependence studies)? Would need the published per-item scores, which this raw source does not carry.
- SourceThe retry-gate mechanism (precision, not recall, sets deployed value) is measured on one escalation target and one item pool at one budget (20%). Does the ZTC-over-JEV precision advantage hold at other retry budgets, or does it invert the way size does between the 397B and 27B configurations on scientific reasoning?
- SourceThe "pending" entry (re-measurement disagreement the harness cannot yet explain) is unresolved in the source. Which system is it, and does the eventual reconciliation move it across the surface-baseline line?
- SourceThe launch claims RLCD yields calibrated, consistent probabilities; the only independent test (the retry gate above) found Jev's low-score tail barely enriched for errors. Does Jev's calibration hold on a workflow-decision task scored against ground truth rather than a two-LLM reference — e.g. a reliability diagram on TypeSafe's own security-triage workflow with labelled outcomes?
- SourceBrowserbase's Stagehand
act()cascade (Jev picks the action; <0.7 confidence falls back to an LLM) reports only latency (median 1.97 s → 0.46 s). What fraction of actions clear the 0.7 threshold, and does end-to-end task success hold, fall, or rise versus the LLM-onlyact()— i.e. does the retry-gate precision problem measured on the shared board reappear in a production cascade?
- SourceThe synthetic ground truth is Gemini-generated and Gemini-classified. How much of the 22.6% is real capability vs same-family cues, and how much is the classifier being penalized for picking a better fit than the seed? Google says it cannot measure this without inspecting private logs.
- SourceWould running ATLAS's validation battery against the AEI and OpenAI classifiers explain the 2–4× cross-study disagreements, or are the gaps driven by sampling and product mix instead?
- ResolvedHuman raters disagree with each other on 42–48% of 3-digit occupations. Is there a principled way to establish the ceiling a classifier could reach, so accuracy can be reported relative to it rather than to 100%? Answered (2026-08-17) by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators: yes — leave-one-annotator-out agreement on the same items, with accuracy reported as
(observed − chance) / (ceiling − chance). Hold out one rater, predict their label from the others, and score the classifier by the identical procedure — the estimator Yang et al. use to show PandaLM retains headroom (best judge κ = 0.753 against a human ceiling of 0.920) while Judge's Verdict is at or past its noisy ceiling (0.620 against 0.562). ATLAS already collected what this needs and stops one step short: B.3.1's three raters over N ≈ 110–120 give model-vs-plurality κ = 0.83 against human-only pairwise κ = 0.66 / Fleiss 0.68 for SOC Major — but agreeing with a plurality of three is an easier target than agreeing with one drawn rater, so those columns are not comparable; the held-out recomputation costs no new annotation. The sibling instrument, from LLM-Judge Validation's Shopify account, is to measure the ceiling before the classifier exists, on the actual rubric by the actual annotators, with κ ≈ 0.2 as a rubric-rewrite trigger — where ATLAS's ceiling is borrowed from Mellow & Sider (1983) and Mathiowetz (1992). Renormalizing this page's table against those borrowed ceilings: SOC Occupation Title 42.47% → 73–82% of the human ceiling, SOC Major Group 71.57% → 85–94%, both understated because the borrowed ceilings sit at coarser granularities than the levels they are applied to. Three limits are part of the answer rather than gaps in it. (i) The approval rate cannot be the ceiling — 85.8% is an anchored statistic (the label is shown) while 42–48% disagreement is a blind one, and this page's own 11–15pp approval-over-plurality gap prices the difference, with Reference-Free Judge Over-Crediting measuring its extreme at FPR 0.719 → 0.012 under commit-first de-anchoring. (ii) At O\*NET-task granularity no ceiling is estimable at all, and ATLAS proves it (B.3.1): when categories vastly outnumber rated observations,p_eis overestimated, κ underestimated, and the computed value is driven by whichever categories were sampled — so 22.58% has no reportable denominator and will not acquire one by hiring more raters. (iii) Consequently aggregation is the ceiling intervention: rolling tasks into Autor–Thompson types (→ 70.44%, 5 categories, agreement study feasible) and merging distinctions users have no economic reason to disclose (Tier 2 42.9 → 51.7%, Tier 3 23.7 → 32.4%) raise ceiling and accuracy together, because the merged distinctions are exactly the ones a human rater reading the same conversation could not make either. The ceiling is a property of the category system, not of the classifier.
- SourceIs the overconfidence advisory's gain direction-specific? Its text asserts the judge is "especially overconfident when predicting 'False'," which Figure 2 supports on source-grounded faithfulness. On a reference-free correctness task, where judges over-credit, does the same advisory still cut AECE, or does it make it worse? A run of both advisory polarities on a reference-free QA set settles it.
- SourceDoes verbalized confidence match or beat logprob G-Eval on level (mean τ_b against human ratings) and not only on subjectivity slope, with confidence intervals, and with G-Eval run at full top-k where the API allows? The paper's newest datapoint (gpt-5.4) shows G-Eval ahead on level, and its slope-difference curve has four points with no interval.
- SourceIs the generation effect about instruction-following capacity, as the author reads it, or about reasoning training? The pre-2025 cohort contains one reasoning model (o1) and the post-2025 cohort is mostly reasoning-capable. Running the recipe on same-generation small models (the paper's gpt-4.1-mini/nano run only on SummEval) and on a non-reasoning post-2025 model would separate the two.
- SourceDoes the capability gradient survive independent human correctness labels? Execution tests are the only reference so far, and the adjudicator disputes 49% of judge-rejected successes, so a human-labelled audit of the 489-patch blind packet (or any equivalent) would say whether stronger agents' "false" acceptances are partly correct patches that fail over-specified tests.
- SourceDo judges that run the tests or read the full trajectory show the same version-dependent error? All SWE-bench judges here see the patch only. An execution-capable or trajectory-reading judge on the same 20 pairs would say whether differential error is a property of patch-only judging or of LLM judging.
- SourceDo the release-decision and transport results replicate on a second scaffold's version lineage or on private development versions closer together than public releases? The OpenHands follow-up has no old/new pairs, and the paper expects closer versions to make
D_Hsmaller andδdominant.
- Weak-Verifier Ensembling4 open
- SourceWeaver's gains are real and its independence assumption is measurably false. How much of the gain survives once the dependence is priced — i.e. what does the weight-fitting recover that a ρ-corrected aggregation would predict, and does the non-monotonic top-1/5/10 curve match Yang's beta-binomial form? Nobody has run the two together. Sharpened (2026-09-10) by anchor judge error correlation estimator, which prices the pricing. The dependence is not estimable from the pool alone: inter-verifier covariance is
σ_t² + σ_c², signal plus shared error, and splitting it needs an external anchor whose own contamination is a free parameter — identified in closed form with ≥ 2 verifiers and ≥ 2 anchors, not identified at any anchor count when verifiers and anchors are all ordinal, and not identified at all when a residual is shared panel-wide (one prompt template across the pool), which is Weaver's configuration. So the question may have no within-pool answer, and the honest version of it is now two cheap experiments rather than one. (i) A Weaver pool spanning ≥ 2 model lineages already carries the family-block statistic for free: computeK_within − K_crossover the verifier score matrix and see whether the lineage residualσ_f²is non-zero — no anchors, no labels, judge metadata only. (ii) Weaver's filter-and-weight steps consume real labels, so those labels are the anchor; if they were produced by a reference model sharing the pool's biases, the designated-clean result says the failure is silent and confidently zero (ρ̂₂ = 0.502 → −0.002as the trusted anchor's own contamination runs 0 → 0.6) rather than noisy. Neither the beta-binomial half of the question nor the top-1/5/10 curve is touched. Partially answered (2026-09-22) by auditing behavioral dependence llm judges, which is the first source in the corpus to actually build the de-entangled aggregator and report it against the competence-only baseline rather than against majority voting. The answer, in the section above: a quality floor plus a competence weight recovers +3.4pp over majority vote, and adding the dependence penalties recovers +0.1pp more (0.881 → 0.882 accuracy, 0.891 → 0.896 precision, 500 held-out questions). So on the corpus's one measured configuration, pricing the dependence buys approximately nothing beyond pricing competence. Three reasons that is not yet the verdict: the pool is three verifiers, so the redundancy penaltyR_maverages over two others and has nothing to discriminate; the verifiers are prompted judges only, so the kind-heterogeneity question below is untouched; and the entanglement is estimated on the models' answering behaviour and used as a proxy for their verifying behaviour, a step nobody has validated. The sharp follow-up is now a pool-size sweep — report the entanglement-minus-competence delta atM_J= 3, 5, 7, 9 and see whether it rises. Advanced again (2026-09-22) by Kohli, who supplies the denominator the question was missing. Measured against a Condorcet independence ceiling rather than against majority voting, established aggregation closes at most 11% of the gap with oracle access to gold labels: accuracy-weighted voting (this page's filter-and-weight step, stripped to its simplest form) closes under 1% on MNLI, Dawid–Skene underperforms majority vote there, and phi-optimal weighting — inverting the correlation matrix, the most direct possible form of pricing the dependence — is best on two datasets and worse than plain voting on a third. So the answer now has a shape: the recoverable fraction is small and the correlation-aware variants are the unstable ones. Still open on this page's own configuration, for the same three reasons as before plus one new one: Kohli's pool is nine prompted judges on classification, so kind-heterogeneity remains untested, and his panel size is fixed at nine, so it does not run theM_Jsweep either. The reporting discipline it does settle: state an aggregator's gain as a fraction of the independence gap, since a 3pp lift over majority voting can be 1% of what was available. - SourceIs error correlation across kinds of verifier — trained reward model vs prompted judge — materially lower than across judges? Every measurement in this wiki is judge-to-judge, which is the case most favourable to the objection and least representative of Weaver's pool.
- SourceDoes the distilled ~400M scorer inherit the ensemble's robustness or only its accuracy? A single small model reproducing a pool's verdicts has, by construction, no diversity left to lose — which is fine for a static selector and is exactly the object optimization pressure would attack first.
- SourceHossain, Yousefi & Lim's CorrFilter helps under global co-failure and costs 3.3 points under vulnerable-subgroup failure using the same estimated correlation matrix, and their regime router closes little of that gap (non-inferiority against a regime oracle fails in 22 of 24 comparisons). Does Weaver's own weight-fitting — which does not distinguish the two regimes at all — inherit the subgroup-failure cost, or does fitting weights per-verifier (rather than per-item, as CorrFilter does) sidestep it? Untested: nobody has run a vulnerable-subgroup intervention against Weaver's actual pipeline.
- SourceWeaver's gains are real and its independence assumption is measurably false. How much of the gain survives once the dependence is priced — i.e. what does the weight-fitting recover that a ρ-corrected aggregation would predict, and does the non-monotonic top-1/5/10 curve match Yang's beta-binomial form? Nobody has run the two together. Sharpened (2026-09-10) by anchor judge error correlation estimator, which prices the pricing. The dependence is not estimable from the pool alone: inter-verifier covariance is