Sources#
Summary#
An LLM-as-a-Judge that emits only a hard verdict throws away the information a threshold, a ranking or a calibration check needs. There are two ways to get a soft score out of one judge call. You can read it off the token log-probabilities of the rating (G-Eval's expected Likert score, Liu et al. 2023). Or you can ask the judge to state a confidence in words. Pre-2025 evidence favoured logprobs. Verbalized scores were overconfident (Xiong et al. 2024) and clustered on round numbers (Dai & Wang 2026), and G-Eval-style soft scoring beat greedy verbal extraction (Wang, Zhang & Choi 2025). The standing advice was "prefer logprobs."
Yu-Chung Hsiao (Cisco Systems, arXiv 2609.10996, 2026-09-10, single author, empirical) argues that the advice has expired on current closed flagships, for two reasons. First, the logprob interface is disappearing. Claude has never returned logprobs. Gemini withdrew them after 2.5 Pro. OpenAI returns none for o1 or GPT-5, makes them configuration-dependent on GPT-5.2/5.4, and cut the top-logprobs cap from 20 to 5. Second, a two-ingredient prompt recipe makes the stated number good enough to use. The study covers 18 models across three benchmarks. Its useful product is less the recipe than the reporting discipline: soft-score calibration (AECE) and spread next to balanced accuracy. The paper's thesis is that accuracy-only reporting cannot see the effect at all.
The recipe, and what each part does#
All three verbalized protocols ask for Answer: True/False plus Confidence: 0–100 under an eight-band confidence rubric ("95–100: Absolutely certain… Below 25: Consider changing your answer"). The rubric, output order and evidence-quoting instruction are held fixed. Only two blocks vary:
- Rubric (baseline): verdict and confidence under the rubric, with no reasoning.
- Ours-FreeReason (ablation): adds the overconfidence advisory and a free-form explanation with quoted evidence.
- Ours: adds the advisory and a single-call self-debate: "1. True because… 2. False because… 3. Explanation: based on both arguments, determine your final answer."
The advisory is directional, and its exact wording matters: "You have been consistently overconfident in past evaluations. Before assigning a high confidence score, actively look for reasons you might be wrong… ALSO IMPORTANT: You are especially overconfident when predicting 'False'." The second sentence encodes a measured property of the baseline. In Figure 2 (Sonnet 4.5 on AggreFact), the Rubric calibration curve sits far above the diagonal at low predicted P(True): judgments stated at about 15% True resolve True about 43% of the time.
The component ablation (10 models × 5 configurations on AggreFact; Table 3, recovered from the PDF because the markdown table is welded) gives each block a separate job:
| Step (Δ against base) | Era | ΔBA | ΔAECE | ΔSpread |
|---|---|---|---|---|
| +Advisory (vs Rubric) | pre / post | −0.7* / +0.0 | −4.5* / −3.6* | +18.9* / +11.6* |
| +Free-form reasoning (vs +Adv) | pre / post | +0.4* / +0.4 | +0.1 / +0.4 | −2.5* / −3.1* |
| +Self-debate (vs +Adv) | pre / post | −1.0* / +0.6* | −0.0 / +0.5 | +3.0* / +0.5 |
| Full recipe (vs Rubric) | pre / post | −1.8* / +0.7* | −4.5* / −3.1* | +22.0* / +12.1* |
| Self-debate alone (vs Rubric) | pre / post | −0.8* / +0.3 | −0.0 / +0.2 | +6.1* / +2.1* |
* = 95% CI excludes zero, paired model-cluster bootstrap. Five models per era.
- The advisory does the calibration. Per model, it lowers AECE and widens Spread on 10/10 models, all significant, and removing it from the full recipe reverses both on 10/10. One sentence telling the model it is overconfident is the largest single lever in the study.
- Free-form reasoning erodes spread, and self-debate preserves it. Free-form chain of thought significantly narrows Spread on 8/10 models, while self-debate widens it on 6/10. The author's mechanism is that a one-directional argument "reinforces a single stance and concentrates probability mass toward the committed answer," while arguing both sides spreads it. This matches Pre-Reasoning Commitment's finding, from a different model family and task, that chain of thought sharpens a pre-committed distribution rather than searching it.
- Self-debate is also where the accuracy cost lives, and only for older models. The next section covers this.
Four results#
1. Calibration and spread improve on every flagship tested. In Table 2, all 10 AggreFact flagships show significantly lower AECE and significantly wider Spread under the full recipe. Spread is a Bhattacharyya overlap between the confidence histogram and a uniform distribution, used because ECE alone misses the collapse onto a few values that verbal scores are known for. The per-task check is broad: 67 of 90 (model, task) Brier comparisons favour the recipe (binomial p = 1.9 × 10⁻⁶), none favours the baseline at 95%, and the gain shows up on all 9 AggreFact tasks. A maximally disjoint second draw of 300 items per task replicates every significant cell.
2. The soft score becomes less coupled to task subjectivity. On SummEval (16 models × 3 dimensions), the slope of mean Kendall τ_b between the judge's probability and expert ratings, plotted against inter-annotator α, flattens by Δslope = −0.143 (one-sided p < 10⁻⁴). The same analysis on the hard prediction does not separate the prompts (p = 0.46). The author frames this as a less confounded capacity measure: a judge whose agreement with humans falls off steeply as the humans disagree more is partly measuring the task's ambiguity. A finer per-example rater-disagreement measure replicates the result on both signals and on HelpSteer2, an independent crowd-annotated benchmark (Δslope −0.024 and −0.058 on SummEval; −0.023 and −0.036 on HelpSteer2; all p ≤ 0.018).
3. The "compatibility shift": verbalized confidence overtakes G-Eval on GPT-5-era models. This test is restricted to the four OpenAI flagships that still return logprobs (gpt-4o, gpt-4.1, gpt-5.2, gpt-5.4). The difference in subjectivity slope (G-Eval minus Ours) changes sign, from about −0.23 and −0.15 before GPT-5 to about +0.09 and +0.52 after (read off Figure 4a). Mean τ_b for the verbalized protocols rises steadily across the four releases (about 0.32 → 0.46), while G-Eval's moves 0.39 → 0.34 → 0.31 → 0.53 (Figure 4b).
4. The generation effect. The full recipe costs pre-2025 judges −1.8pp balanced accuracy and gives post-2025 judges +0.7pp (cluster-bootstrap means, both significant). On post-2025 models, the gap between oracle-threshold BA and emitted-prediction BA (the Oracle-Prediction Gap) and the BA cost of self-debate relative to free-form reasoning (Debate Stress) both shrink toward zero. Debate Stress goes negative for Sonnet 4.5 and 4.6, where debate helps. The author's explanation is that newer flagships can follow a richer judge prompt without losing their decision. That is why accuracy-only reporting misses the shift: ΔBA looks flat on new models, and ΔAECE alone does not show whether the prediction still carries the signal.
Read these results narrower than the abstract does#
- The G-Eval reversal is a slope result, not a level result. On gpt-5.4, the newest model in the comparison, G-Eval's mean τ_b (≈0.53) is higher than every verbalized protocol's (≈0.45–0.46) in the paper's own Figure 4b. G-Eval agrees better with human ratings there. Its advantage just degrades faster as tasks get more subjective. The abstract's "verbalized confidence is the better signal" holds only for the subjectivity-robustness sense the body defines. The slope-difference curve has four points, one per model, with no confidence interval.
- G-Eval was measured on a degraded interface. On gpt-5.2 and gpt-5.4, logprobs are "config dependent" with a top-5 cap, against 20 on gpt-4o and gpt-4.1. The paper does not say whether truncation limits G-Eval's expected-score readout on a Likert scale. The API change that motivates the paper is therefore also a confound in its central comparison.
- "Post-2025 models gain accuracy" really means "post-2025 models pay no cost." Per model, only sonnet-4.6 (+1.2*) has a significant BA gain in Table 2. gpt-5.2 and gpt-5.4 are at +0.6 (n.s.) and gemini-3.1-pro is at −0.3. At the (model, task) level, post-2025 BA splits 24 better (3 significant) against 20 worse (2 significant), close to a coin flip. Pre-2025 splits 16 (2) against 29 (10). The era asymmetry is real (Fisher one-sided p = 0.013). The positive sign is not established.
- "Generation" is confounded with everything else about the model. Each era is five models, and o1 is the only reasoning model in the pre-2025 cohort. Vendor, reasoning training, post-training recipe and the logprob interface all change at the same boundary. The instruction-following explanation is the author's interpretation, and no experiment isolates it.
- Scope is single-call, pointwise, binary judging on source-grounded or subjective-quality tasks. Tasks with checkable answers (code, math, agent objectives) are excluded by design, as are pairwise and listwise judging and multi-agent debate.
- Two source-internal inconsistencies. The abstract and introduction say Gemini dropped logprobs "at 3.1 Pro," while Appendix A and Table 4 say from 3.0 Pro (2025-11). Figure 2's caption gives Sonnet 4.5's AECE as 8.9% → 3.8%, while Table 2 gives 9.1% → 3.3%. The figure uses 10 equal-mass bins, so the difference is probably binning. This page quotes Table 2.
Where this sits in the wiki#
- It measures what LLM-Judge Validation deferred. Norman et al.'s Minimum Viable Validation Protocol leaves calibration proper (ECE/Brier) out "because most providers don't expose token logprobs." This paper computes AECE and Brier for Claude, Gemini and GPT judges with no logprobs at all. So the missing step can be run on any closed judge today, as a check on the judge's stated confidence. The catch is that the judge's calibration is now partly a property of the prompt: one advisory sentence moves AECE by 1.1–7.1pp. A calibration number is valid only for the exact prompt it was measured under.
- It is the prompt-side version of Yang et al.'s version axis. Yang et al. show that upgrading a judge model is a measurement event that must be re-validated. Hsiao shows the same for the judge prompt: an advisory and a debate block that improve calibration on every model cost significant accuracy on the previous generation and nothing on the current one. A judge prompt tuned on one generation is not safe to carry to the next in either direction.
- It is the closed-model, prompt-only counterpart to readout calibration. Trained Calibration and Confident But Unsure carry Sarfati et al.'s finding that a model's stated probability ranks as well as a probe on its activations (AUROC 0.756 vs 0.758) and is simply worse calibrated (ECE 0.093 vs 0.044): "overconfidence is less a failure of self-knowledge than of self-report." That diagnosis predicts that the verbal channel can be repaired without new information. Hsiao repairs it with no weights and no activations, on models where neither is available, and the lever is exactly a statement about the channel's bias. Kumaran et al. (2026), cited by the paper but not ingested here, report mechanistically on open-weight models that verbal confidence carries information beyond the token probabilities.
- The direction of the error runs opposite to reference-free over-crediting, which fits the setting. Reference-Free Judge Over-Crediting finds judges with no reference answer lean generous and over-credit wrong answers. Here, on AggreFact faithfulness, the baseline judge is especially overconfident when saying False. AggreFact always shows the judge the source document, which is the reference-visible condition, and Kranti & Vajjala find that explicit comparison against a reference makes flips run in both directions, rescuing wrongly rejected answers 30–70% of the time outside English. The two results agree that a judge's bias direction depends on the setting. They also mean the advisory's hard-coded "especially overconfident when predicting False" is a setting-specific claim. On a reference-free task it would push the wrong way.
- It keeps the verbal channel open that Error-Penalized Abstention Training could not test. That paper reads its confidence report through a linear probe and states that the verbalized channel is untested. This is the strongest evidence in the wiki that the verbal report is controllable after training. It is also evidence that the report is highly sensitive to instruction, which is a hazard for any reward built on it.
Connections#
- LLM-as-a-Judge — the primitive; this is the choice of soft signal for its single-call, pointwise form, and the reason a thresholded judge decision can now carry a calibration number
- LLM-Judge Validation — the validation discipline whose deferred calibration step this makes runnable on closed judges, and whose judge-version axis this extends to the judge prompt
- Trained Calibration — the training-side routes to a calibrated report; this is the zero-training, prompt-only route on models whose weights and activations are unavailable
- Confident But Unsure — the self-report failure this recipe targets; the "especially overconfident when predicting False" advisory is a hand-written correction for a measured miscalibration of the stated number
- Pre-Reasoning Commitment — the mechanism behind free-form reasoning narrowing spread on 8/10 judges: chain of thought sharpens an already-committed answer, and only a prompt that forces both stances keeps the probability mass spread
- Reference-Free Judge Over-Crediting — the opposite bias direction, in the opposite reference condition; together they show the advisory's direction is task-specific
- Error-Penalized Abstention Training — its untested verbal-report channel, shown here to be movable by instruction alone
- OpenAI — the only frontier vendor still returning logprobs for a same-family comparison, and with a narrowing interface (top-20 → top-5, config-dependent)
- Anthropic — Claude models have never returned logprobs, so a verbalized protocol is the only soft signal available for a Claude judge
- Version-Dependent Judge Error — where stated probabilities stop helping: averaging three judges' probabilities lowers SWE-bench comparison error to 3.0 points but leaves the version-dependent error in place, and single-judge probabilities do not consistently beat binary verdicts on paired efficiency
Open Questions#
- Is the overconfidence advisory's gain direction-specific? Its text asserts the judge is "especially overconfident when predicting 'False'," which Figure 2 supports on source-grounded faithfulness. On a reference-free correctness task, where judges over-credit, does the same advisory still cut AECE, or does it make it worse? A run of both advisory polarities on a reference-free QA set settles it.
- Does verbalized confidence match or beat logprob G-Eval on level (mean τ_b against human ratings) and not only on subjectivity slope, with confidence intervals, and with G-Eval run at full top-k where the API allows? The paper's newest datapoint (gpt-5.4) shows G-Eval ahead on level, and its slope-difference curve has four points with no interval.
- Is the generation effect about instruction-following capacity, as the author reads it, or about reasoning training? The pre-2025 cohort contains one reasoning model (o1) and the post-2025 cohort is mostly reasoning-capable. Running the recipe on same-generation small models (the paper's gpt-4.1-mini/nano run only on SummEval) and on a non-reasoning post-2025 model would separate the two.
Sources#
- Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models — Yu-Chung Hsiao (Cisco Systems, single author), Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models, arXiv 2609.10996, 2026-09-10, 18pp,
empirical. No COI with the judged vendors: Cisco builds none of the models evaluated. Used for: Method (the three protocols, Figure 1 and the Appendix I verbatim prompts, including the advisory text); Table 1 (the six metrics and the Spread formula); Figure 2 (Sonnet 4.5 calibration curves; opened to confirm the Rubric curve's low-P(True) overconfidence); §Ours / Figure 3 / Tables 5–6 (the SummEval Δslope −0.143 and the per-example SummEval and HelpSteer2 replications); §Compatibility Shift and Figure 4 (both panels opened; the slope-difference and mean-τ_b values quoted as "≈" are read off the figure, and the paper prints no numbers for them); Table 2 (per-model BA/AECE/Spread; clean, verified againstpdftotext -layout); Table 3 (component ablation); Tables 7–9 (per-model ablation grid; the 10/10 advisory results and the 8/10 free-form spread erosion); Table 10 and App. E (the 90-comparison era breakdown); Table 11 / App. G (sampling stability); Table 4 / App. A (logprob availability). Parse warning: Table 3 and Table 10 are welded in the docling markdown (Pre/Post rows fused into one cell holding two values, e.g.30 (14 sig) 37 (24 sig), and Table 3's lower rows lose their base-column alignment). Both were recovered in full frompdftotext -layouton the local PDF, and the values above come from that recovery. Table 3's bolded full-recipe cells (−1.8 / +0.7) also match its caption. Table 10's recovered values are Brier pre 30 better (14 sig) / 15 worse (0 sig), post 37 (24) / 8 (0); BA pre 16 (2) / 29 (10), post 24 (3) / 20 (2). Figures 5a and 6 were not opened, and the Oracle-Prediction Gap claims are quoted from prose.
Cited by 10
- LLM-Judge Validation×4
All judges were run with thinking suppressed. Does reasoning-on flip the consistency–bias paradox,…
- LLM-as-a-Judge×3
How far can the judge's absolute calibration be trusted for thresholded decisions (ship/no-ship,…
- Reference-Free Judge Over-Crediting×2
verbalized confidence llm as a judge compatibility shift — Yu-Chung Hsiao (Cisco Systems, single…
- Trained Calibration×2
Verbalized Confidence Judge Scoring — the prompt-only route for closed models: an overconfidence…
- Confident But Unsure
Verbalized Confidence Judge Scoring — the stated-confidence failure corrected by instruction alone…
- Error-Penalized Abstention Training
Verbalized Confidence Judge Scoring — the verbal confidence channel this paper leaves untested,…
- Evals & Benchmarks
Verbalized Confidence Judge Scoring — Having a judge state a 0–100 confidence next to its…
- Open Questions Backlog
Verbalized Confidence Judge Scoring ×3 (oldest 4d) — Is the overconfidence advisory's gain…
- Pre-Reasoning Commitment
Verbalized Confidence Judge Scoring — the sharpening effect showing up as a measurable calibration…
- Version-Dependent Judge Error
Verbalized Confidence Judge Scoring — soft scores do not rescue the paired comparison. Averaging…
Related articles
- LLM-as-a-Judge
Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…
- Trained Calibration
TML's recipe for making calibration a first-class RL target: proper scoring rules on resolved real-world questions, abs…
- Cross-Model Error Entanglement
Three 2026 audits of whether LLM errors are independent: Kuai et al. measure excess co-failure and same-distractor coll…
- Reference-Free Judge Over-Crediting
Reference answers are a first-order determinant of LLM-judge verdicts: without one, judges systematically over-credit w…
- User Awareness
The model inferring who it is talking to from ordinary harness clues — an account e-mail, a folder name, a MEMORY.md —…
