H
Howardism
Plate IIEvals & BenchmarksHOWARDISM

Version-Dependent Judge Error

A frozen LLM judge's false-accept and true-accept rates move with the agent version it grades, so a judge-only old-vs-new comparison can be biased as well as attenuated. Li (arXiv 2609.34198) on 20 prespecified SWE-bench Verified version pairs, τ-bench and AgentRewardBench: rank correlation 0.71–0.79 yet 8 of 60 judge-pair units declare upgrades execution cannot establish; failed-patch acceptance rises with agent capability (ρ ≈ +0.82, replicated on 8 OpenHands configs at +0.94); transporting old-version calibration multiplies comparison error 5.2×; a paired PPI++ audit is valid but saves ~5% of labels at 80 tasks.

Article metadata
Publication details
Published:October 1, 2026
Filed:Concept
Domain:Evals & Benchmarks
Tags:EvaluationLLM As A JudgeMeasurementRelease DecisionsPrediction Powered InferenceAgents
Reading:18 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Version-Dependent Judge Error

Sources#

Summary#

The release question agent teams actually ask is paired: is v(n+1) better than v(n) on these tasks? The common answer holds an LLM judge fixed and reads the difference in judged success. That reading assumes the judge's error rates do not depend on which version wrote the trajectory. Measurement-error theory calls the assumption non-differential misclassification. Under it, a judged difference is shrunk toward zero (by the judge's Youden index J = TPR − FPR) but keeps its sign. Under differential error it can be biased in either direction.

Jiapeng Li (arXiv 2609.34198, v2 2026-09-29, single author, empirical) measures how much that matters for the release decision itself, on benchmarks with execution or expert reference labels. The paper says plainly that it does not discover the phenomenon (Dorner et al. 2025, Fiedler 2026, GAUGE/Bodhwani 2026 and Huang 2026 precede it). Its contribution is the paired-release estimand made explicit, 20 prespecified public agent-version pairs, jointly resampled uncertainty, a randomized input intervention, and one of its own registered predictions falsified across domains.

This is the fourth invariance the wiki's judge pages have tested. LLM-Judge Validation varies the benchmark (Norman) and the judge's own version (Yang). DRACO varies the judge model. This page holds the judge fixed and varies the thing being judged. Where upgrading the judge proved a small perturbation, upgrading the agent under a frozen judge proved a large one.

Setup#

  • SWE-bench Verified (primary). 35 public mini-SWE-agent submissions × 250 held-out issues. Three judges from three providers: Gemini 3.5 Flash, GPT-5 mini, GPT-6 Sol. They see the issue and the graded patch and nothing else (no tests, no trajectory, no reference patch), on 8,743 aligned cells. A fourth planned judge, Claude Opus 5.5, stopped returning verdicts after 87.5% of cells, so it is reported only on a matched subset. The 20 version pairs were fixed before any verdict existed: 12 model upgrades (GPT-5 → 5.1 → 5.2, Claude Opus 4 → 4.5 → 4.6, GLM-4.5 → 4.6 → 5 …), one reasoning-effort change, two scaffold upgrades, three combined changes and two variants. The reference is the recorded test-based resolution.
  • τ-bench. Claude 3.5 Sonnet (new) vs GPT-4o, airline and retail, four trials each on 155 tasks, all four judges. The reference is the database-state reward.
  • AgentRewardBench. 1,106 expert-labelled web-agent trajectories, four agents, the 15 released evaluators. No new model calls.
  • Post-submission follow-up. Eight OpenHands configurations on the same 250 issues (1,981 three-judge cells), plus one descriptive Agentless submission. It was locked after the original results were known, so it is a new-scaffold check on the same tasks, not a new-task replication.

The plan was frozen in a private repository on 2026-09-23 with file hashes. The paper corrects its own earlier claim of public preregistration, and every deviation is dated in a table (§4.6).

Four identities that organise the results#

With q_v = FPR_v + J_v·p_v (the judged rate as a function of the true rate p_v, the same line Wu et al. write as Ā = ρ₀ + J·Q):

  1. Attenuation plus a differential term (Prop. 1): comparison error e = δ − (1 − J_o)·D_H, with δ = ΔTPR·p_n + ΔFPR·(1 − p_n). Close versions have small D_H, so their error is dominated by δ. A sign reversal needs |δ| > J_o·|D_H|.
  2. Transported calibration amplifies the differential term (Prop. 2): Rogan-Gladen correction of the new version's rate using the old version's TPR/FPR gives error δ / J_o. It removes attenuation and divides the version-dependent part by J_o < 1. This is the version-pair form of Fiedler's shared-calibration 1/J amplification.
  3. Attenuation is set by the discordant tasks (Prop. 4): even a judge whose errors depend only on the task compresses a difference by its discriminability on the tasks where the versions differ, not by its average J.
  4. A judge saves less for a comparison than for a level (Prop. 3): the prediction-powered efficiency gain for a paired difference is the level gain × (1 − r)/(1 − r_ε). Here r is the correlation of the two versions' outcomes across tasks and r_ε that of the judge's errors. Successive versions solve mostly the same tasks, so r is high and the paired gain is small.

Results#

Differential error is present everywhere it could be tested. All three available SWE-bench judges reject within-(task, label) exchangeability (p = 0.0001 each), as do all four τ-bench judges on the cluster-robust test. AgentRewardBench rejects for only 5 of 14 LLM evaluators, short of the preregistered half-of-judges rule. Median false-positive rates across the 35 SWE-bench agents are 65.7% (Gemini) / 66.0% (GPT-5 mini) / 40.0% (GPT-6 Sol), and median false-negative rates 9.6% / 13.1% / 22.2%. Per-agent FPR ranges are 7.5–95.0%, 9.5–90.0% and 2.0–69.0%. A patch-only judge can be accurate on correct patches and still accept most incorrect non-empty ones.

Ranking agreement does not transfer to release decisions. Kendall τ with the reference ordering of 35 agents is 0.71–0.79. On the 60 judge × version-pair units, the differential component e_d is detectable (BH-adjusted) in 32 (53.3%), against ~1.1% under a simulated judge whose errors depend only on the task. Discrete release decisions disagree in 11/60. In 8 of those, a judge-only interval declares an improvement the execution interval cannot establish. Three miss a real improvement, and there is no confident better-to-worse reversal. Two examples: a Claude Sonnet 4.5 scaffold+effort change moves the test solve rate by −0.4 points (−5.2 to +4.4), yet GPT-5 mini judges it +8.0 and GPT-6 Sol +9.6, both intervals excluding zero. All three judges call Claude Opus 4 → 4.5 a 10–16-point improvement against a 4.0-point reference (−1.2 to +8.9). The pooled signed error is +2.8 points and the differential component +6.8 points, so frozen judges systematically flatter the newer version on this cohort. The paper's own framing matters here: the disagreement rate does not exceed the non-differential simulation (18.3% vs ~32.7%). The hazard is confidence the reference does not support, not frequent reversals.

τ-bench supplies the confident reversal. In retail the reference has GPT-4o 9.0 points worse (CI −16.0 to −2.5). GPT-6 Sol calls it 17.8 points better (CI 11.8 to 24.1), so a team using that judge would ship the worse agent. The mechanism is visible in the transcripts. The τ-bench policy forbids calling a tool and messaging the user in the same turn, and the database-state reward ignores that rule. Every Claude conversation contains such turns (7.4 per successful retail conversation), against 0.5 for GPT-4o. 31.7% of GPT-6 Sol's rejections of Claude's successful conversations cite the rule, and none of its rejections of GPT-4o do. The judge was asked whether the agent "followed the policy", so it measured a stricter construct than the reference. A behavioural habit difference between agents became a judge-error difference. The paper's verdict: the strict judge is not wrong by a human policy standard, only as a proxy for the published reward, and which standard applies is a product decision that has to precede calibration.

Transported calibration is worse than no correction. On SWE-bench, correcting the new version with the old version's TPR/FPR is worse than the raw judge in 59 of 60 units. Mean absolute comparison error rises 3.8 → 19.5 points (ratio 5.20), above every draw of both simulated non-differential nulls (means 1.71 and 1.47). Judge Youden indices on these pairs are only 0.19–0.38, and 35% of corrected rates had to be clipped to [0, 1]. The ratio is 1.55 on AgentRewardBench and about 4.0–4.1 on τ-bench (descriptive there). Claude Opus 5.5 on its matched subset is the outlier at 3.5 → 5.2 points, so magnitude depends heavily on the judge. Caveat carried from the paper: 24.6% of bootstrap draws hit a non-positive old-version J and were dropped, so the interval is conditional on the estimator staying defined.

Paired audits are valid, but a judge saves few labels. A random-sample paired audit with power-tuned prediction-powered inference (PPI++) covers the reference difference at 89.5% (20 tasks) → 94.4% (160 tasks). Judge-only intervals cover it in 80.0% of SWE-bench units, 76.7% on AgentRewardBench, and 75% / 50% in τ-bench airline/retail. At 80 audited tasks PPI++ is only ~5% narrower than a classical label-only interval on SWE-bench (13% on AgentRewardBench). Untuned PPI is wider than classical (19.8 vs 14.9 points). The ideal paired efficiency factor is 1.084 against 1.133 for levels, because two versions' outcomes correlate at 0.615 across tasks. The planning arithmetic is sobering: resolving a 5-point difference between versions that disagree on 25% of tasks takes ~217 labelled pairs without a judge and ~202 with one (pool of 300). For future tasks rather than the pool it takes ~778 and ~613, more than either benchmark holds.

Averaging judges helps only when their distortions point opposite ways. On τ-bench the four judges' biases partly cancel and every aggregate reaches the reference decision. The size of the aggregate's comparison error still depends on the tie rule, and counting ties as acceptances exaggerates the retail gap by 8.3 points. On SWE-bench the distortions share a direction. Averaging the three judges' probabilities lowers absolute error to 3.0 points but leaves differential error and decision disagreements in place, and the ensemble's paired efficiency is 1.145 against 1.084 for one judge.

The capability gradient, and the two predictions that failed#

Stronger agents' failures are more convincing (H6, registered from AgentRewardBench before SWE-bench verdicts were read). Across 35 SWE-bench agents, the Spearman correlation of reference resolve rate with the judge's task-conditioned false-acceptance contrast is +0.827 / +0.764 / +0.865 (mean +0.819, p = 0.0001). The correlation with the Youden contrast is mean −0.747. The top half of agents gets a 21–26-point higher false-acceptance contrast than the bottom half on the same tasks. A shuffle within (judge, task, label) strata gives means near zero, so this is not the artefact of agents solving different tasks. The post-submission OpenHands cohort reproduces it at +0.944 / −0.937 (exact eight-agent permutation p = 0.000099 / 0.000397). A blinded GPT-5.5 adjudicator given the maintainer patch finds no capability gradient in reference error (Spearman −0.003), which argues against "stronger agents' test failures are really correct". It disputes the test label on 23.4% of judge-accepted failures and 49.1% of judge-rejected successes, though, so the execution reference is itself noisy. Independent human patch review is still pending: a two-reviewer blind packet of 489 patches exists with no labels.

The solvability gradient reverses sign (H7 falsified). On AgentRewardBench, judges accept a failure more often when other agents solved the task: 12.4% → 23–26%, the "near-miss" reading. On SWE-bench the same slope is negative: false acceptance falls by 38–48 points as the share of other agents solving the issue goes 0 → 1, still −20 to −24 after agent and repository fixed effects. There is no universal rule that failures on solvable tasks are harder to judge.

Showing the judge the agent's own final message does not raise false acceptance (H5, randomized). 2,100 (agent, issue) cells had three randomized arms: patch-only, patch plus own final message, and patch plus a placebo message from another issue. The report-minus-control false-acceptance change is −1.47 / −0.49 / +0.61 points, Holm-adjusted p = 1.0 for all three judges. The result rules out large average effects (upper bounds ≈1–3.5 points) under this patch-only prompt only. It says nothing about full trajectories. It is a direct counterweight to the plausible "confident self-report persuades the judge" story (Advani 2026's "false success"): the capability gradient above is not explained by the self-report channel on this prompt.

Practical protocol (the paper's §6)#

  1. Treat judged differences as screening, not evidence. They decide which comparisons deserve an audit.
  2. Audit the comparison, not the level. Draw tasks at random, label both versions, and estimate the paired difference with PPI++ using the judge as proxy. This is unbiased however the judge's errors depend on version. Use Student-t quantiles and treat ~40 tasks as a rough floor.
  3. Size the audit from the pair's own discordance and ρ_D, and respect the finite pool. A few hundred tasks may not resolve a 5-point gap at all.
  4. Do not transport calibration from the previous version unless invariance is checked on the new version's labels. Checking it requires the labels the paired audit already collects.
  5. Keep fixed anchors, but only as judge-stability monitors. Re-scoring a frozen labelled set cannot, by construction, see error induced by the new version's outputs (§3.4). Anchors certify the judge, not its validity on what is being compared.
  6. Publish per-version class-conditional FPR/FNR next to the judged score.
  7. Test, don't assume, any input-channel effect such as agent self-report.

Evidence notes#

  • Tier empirical, with provenance caveats the reader should weigh. Single author. The preregistration is a private-repo commit, not a public timestamp (the paper says so). Code and row-level outputs are private at posting. The follow-up cohort was locked after the original results were known. Per the paper's own disclosure, GPT-6 Sol substantially assisted design, code, analysis and drafting, and GPT-6 Sol is also one of the three judges.
  • The reference is execution, not human correctness. Test-based SWE-bench labels are noisy in both directions (the adjudicator disputes nearly half of the judge-rejected successes). So part of what is called differential judge error could be differential reference error the adjudicator missed.
  • One scaffold family carries the version pairs. The OpenHands follow-up extends the capability gradient only. It has no old/new lineage, so the release-decision and transport results are not replicated beyond mini-SWE-agent.
  • Judges see the patch only. Judges that run tests or read trajectories may err differently. The 60,000-character patch cap truncated 29 follow-up prompts.

Connections#

  • LLM-Judge Validation — the fourth invariance, completing the page's axis list: benchmark (Norman), judge version (Yang), panel composition, and now evaluated-agent version under a frozen judge. Yang found a judge upgrade a small perturbation. This paper finds an agent upgrade a large one, with median judge FPRs of 40–66% that no judge-level validation would reveal, because the failure lives in the judge–target pair
  • LLM-as-a-Judge — the "rankings safe, thresholds not" split, measured on the decision that matters. Kendall 0.71–0.79 coexists with 8 unjustified upgrade declarations, and the ship/no-ship call is exactly the thresholded use that page's open question asks about
  • Recurring Production Agent Evaluation — the same evolving-system setting, with a judge in the loop instead of deterministic scoring. She & Lin leave release-gate decisions as future work. This paper is that measurement, and its answer is to gate on a paired audit of current outputs
  • Stopping Under a Noisy Verifier — the same q = FPR + J·p line, used for a different decision. Wu et al. need J to place a loop against the repairer's boundary. Here J is the divisor that makes transported calibration blow up, and on close version pairs it is only 0.19–0.38
  • Cross-Model Error Entanglement — a confound for that page's judge-bias association: judge over-acceptance of a target's failures tracks the target's capability at ρ ≈ +0.82 with no lineage variable involved, so an entanglement-vs-bias correlation that does not control for target capability can pick this up. It does not bear on BEI's answerer-side vintage question (see the annotation there)
  • Failures That Look Like Success — the judge-side measurement of the class: stronger agents' failed patches are more convincing to a patch-only judge. But the randomized self-report arm finds the agent's closing message is not the channel on this prompt, so the persuasion lives in the patch
  • Reference-Free Judge Over-Crediting — the SWE-bench judges grade without the maintainer patch or tests, and accept a median 40–66% of failed patches. That is Kranti & Vajjala's over-crediting at agent scale, now shown to vary with the agent
  • Weak-Verifier Ensembling — aggregation as a test of shared direction. Four τ-bench judges with opposite distortions cancel to the reference decision. Three SWE-bench judges with the same distortion do not, so an ensemble removes differential error only when the members' errors are not aligned
  • Benchmark Score Redundancy — the PPI machinery CollabEval uses for levels, priced here for paired differences: valid, but at 80 tasks it narrows an interval by ~5%, and untuned PPI is wider than no judge at all
  • Verbalized-Confidence Soft Scoring for LLM Judges — soft scores do not rescue the paired comparison. Averaging three judges' stated probabilities cuts absolute error but leaves differential error, and individual probabilities do not consistently beat binary verdicts on paired efficiency

Open Questions#

  • Does the capability gradient survive independent human correctness labels? Execution tests are the only reference so far, and the adjudicator disputes 49% of judge-rejected successes, so a human-labelled audit of the 489-patch blind packet (or any equivalent) would say whether stronger agents' "false" acceptances are partly correct patches that fail over-specified tests.
  • Do judges that run the tests or read the full trajectory show the same version-dependent error? All SWE-bench judges here see the patch only. An execution-capable or trajectory-reading judge on the same 20 pairs would say whether differential error is a property of patch-only judging or of LLM judging.
  • Do the release-decision and transport results replicate on a second scaffold's version lineage or on private development versions closer together than public releases? The OpenHands follow-up has no old/new pairs, and the paper expects closer versions to make D_H smaller and δ dominant.

Sources#

  • Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation — Jiapeng Li, Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation, arXiv 2609.34198 (v1 2026-09-28, v2 2026-09-29), 45pp, empirical. Cited for: §1 and §2 (framing, prior work); §3.1–3.4 (Props. 1–4, the anchor argument); §4.1–4.7 (design, deviations); §5.1–5.6 (all results above); §6 (protocol); §7 Limitations; Appendix E (τ-bench rationales). Parse note: docling-derived (2.126.0 / docling-mlx 0.1.1, confidence_grade: excellent, 12 tables). Table 1's flagged table-collapse cells are intended FPR / FNR pairs, not welds. Every Table 1 cell was reconciled against the §5.2 prose: the FPR/FNR medians match, 12+9+11 = 32 nonzero e_d, 4+4+3 = 11 disagreements, and the per-judge transported errors 20.3 / 21.8 / 16.5 average to the reported 19.5. Tables 2–4 match the prose. Appendix B's version-pair table has digit splits (0.58 0, +0.08 0, 1 0 for pair 10). No Appendix B row is cited on this page.
§ end
Cited by 13
Related articles
  • LLM-as-a-Judge

    Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…

  • LLM-Judge Validation

    UC Berkeley's 21-judge / 9-provider / ~541K-judgment audit (Norman et al., 2026): LLM-as-a-judge validation is systemat…

  • Reference-Free Judge Over-Crediting

    Reference answers are a first-order determinant of LLM-judge verdicts: without one, judges systematically over-credit w…

  • Cross-Model Error Entanglement

    Three 2026 audits of whether LLM errors are independent: Kuai et al. measure excess co-failure and same-distractor coll…

  • Same-Model Review Blindness

    Greptile's Rodrigo Caridad on two 500-PR labelled datasets (~1,500 verified high-severity bugs): each frontier model ca…