Sources#
- Separating signal from noise in coding evaluations
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Summary#
A repository-level coding benchmark grades a submission by running hidden tests against it. The task is only well-posed if the prompt and the tests describe the same behaviour. When they don't, a failure no longer means the model couldn't do the task, and a pass no longer means it did. This page is about that defect, the task being wrong. Two neighbouring failures are covered elsewhere: the model reaching the answer key, on Evaluation-Time Answer Leakage, and the model having seen the answer in training, on Benchmark Contamination and Decontamination.
OpenAI's Separating signal from noise in coding evaluations (2026-07-08, empirical, lab voice, unbylined) is the corpus's first prevalence estimate for this defect on a named frontier benchmark. It audits all 731 tasks of SWE-Bench Pro's public split:
- An automated filter flagged 286 tasks (39.1%). The filter read the prompt, model attempts, failure traces and task metadata.
- Two independent review paths ran on those 286:
- A human-supervised agent review judged 200 broken (27.4% of 731). Codex-based investigator agents with repository and environment access ran several independent audits per task, and a researcher made the final call.
- A human annotation campaign judged 249 broken (34.1%). Five trained software engineers reviewed each task independently, and disagreements were escalated.
- OpenAI's stated estimate is ~30% broken. It retracts its earlier recommendation that the community move from SWE-bench Verified to SWE-Bench Pro. That recommendation was made after OpenAI's own audit of SWE-bench Verified found "fundamental design and contamination issues". That earlier audit is not ingested here.
Over the same eight months, frontier pass rates on the split went from 23.3% to 80.3%.
The four defect types, and which way each pushes the score#
OpenAI's taxonomy, with the share of the whole dataset each review path assigned. The two low-coverage figures are stated in the prose. The rest are transcribed from a rendered bar chart, whose pairing of values to series was inferred at ingest (see Sources for the reconciliation problem):
| Defect | What goes wrong | Effect on the score | Agent path | Human path |
|---|---|---|---|---|
| Overly strict tests (formerly "narrow tests") | Tests enforce implementation details the prompt doesn't state, so functionally correct submissions fail | Deflates (false fail) | ≈14.4% | ≈17.8% |
| Misleading prompt | Prompt points toward the wrong behaviour or contradicts the tests | Deflates | ≈6.3% | ≈7.5% |
| Underspecified prompt (formerly "wide tests") | Tests enforce requirements the prompt omits and that can't reasonably be inferred | Deflates | ≈0.6% | ≈0.8% |
| Low-coverage tests | Tests under-check the feature, so an incomplete fix passes | Inflates (false pass) | 4.1% | 9.4% |
| Miscellaneous | — | either | ≈1.9% | ≈1.2% |
The headline mixes errors of opposite sign. Roughly 21–26% of the dataset carries defects that make correct work fail, and 4–9% carries defects that let incorrect work pass. "~30% broken" is the sum of the two, so it says nothing about which way the published scores are biased. SWE-Bench Pro Verified's single "Verified" score has the same property, where a large deflation and a small inflation are reported as one number (see Evaluation-Time Answer Leakage).
The worked example is one character. In OpenLibrary-77c16d5, the prompt specifies TocEntry.to_markdown() output down to the character, with one leading space (" | Chapter 1 | 1"). The hidden test_to_markdown assertions require two. A model that follows the prompt exactly fails. OpenAI files this under misleading prompt. Only this one of the page's four failure-mode examples was captured at ingest.
Where the defects come from. OpenAI's explanation is structural: issues and pull requests "were originally created for human collaboration", so problem statement, merged code and unit tests "do not always line up to form clean, isolated tasks". A PR's tests are "written to validate a specific change, rather than to define an implementation-agnostic standard for solving the task". Mining tasks from commit history inherits the specificity of the commit.
A score above the defect ceiling (inference, not the source's claim)#
Take the deflating categories at face value. Somewhere around a fifth to a quarter of the 731 tasks then penalise a correct, prompt-following submission. Yet the top reported pass rate is 80.3%. Two readings are compatible with that:
- "Overly strict" does not mean "unpassable". A model can match the gold patch's arbitrary choice by convention or by luck, and OpenAI says the tests invalidate "many", not all, correct submissions.
- A model can match the gold patch because it retrieved it. Zheng et al. show that the same 731 tasks leaked the upstream fix at run time, and that closing the channels drops the leader from 89.06 to 62.93. Overly strict tests are exactly the ones that reward reproducing the reference implementation over solving the stated problem. So the two defects compound: a strict test turns retrieval from a shortcut into the only route to a pass.
Neither the audit nor Zheng et al. measures this interaction. It is recorded here as an open question, not a finding.
Two audits of the same 731 tasks#
| OpenAI (July 2026) | Zheng et al., SWE-Bench Pro Verified (September 2026) | |
|---|---|---|
| What was examined | All 731 screened by an automated filter, then 286 flagged tasks reviewed in depth | 119 candidates drawn from public issue reports |
| Output | Prevalence estimate: 200 / 249 broken | Repair count: 102 revised, 17 rejected |
| Taxonomy | Overly strict, underspecified, misleading, low-coverage | Overly narrow test 75, misleading description 22, overly broad test 3, other 2 |
| Remedy | Retract the recommendation; call for new benchmarks written by experienced engineers | Specification-first repair: requirements edited in 90.2% of repairs, test_patch in only 16.7% |
Three things follow from putting them side by side.
- The taxonomies are the same taxonomy. OpenAI's footnotes say "overly strict tests" was previously called narrow tests and "underspecified prompt" was previously called wide tests. Zheng et al.'s "overly narrow test / overly broad test" categories are those earlier names. Their relative shares line up too: strict/narrow dominates, and misleading prompts come second.
- The repair can't have fixed the inflating category. Low-coverage tests are fixed by adding tests. A specification-first policy that touches
test_patchin 17 of 102 repairs, and has no low-coverage category at all, leaves the 4.1–9.4% of tasks that let incomplete fixes pass untouched. "Verified" removes false fails and leaves the false passes. - The two remedies define different constructs. Zheng et al. repair an overly strict test by writing its arbitrary detail into the prompt (the Teleport ordering case). OpenAI's diagnosis is that such tests should be an "implementation-agnostic standard", which means relaxing the test. The first turns the task into "follow this convention exactly". The second keeps it as "solve this problem". Both remove the mismatch, but they measure different things.
Agents as auditors, and how far they agree with humans#
The audit is also a validation datapoint for agent-assisted benchmark QA, and the source frames it that way: evaluation flaws "are easier to detect now than they would have been even a short time ago". The agreement evidence is thinner than that framing suggests.
- The agent path is conservative. It found 200 broken tasks where humans found 249. It was less likely to assign multiple labels, and it undercounted low-coverage tests most (4.1% against 9.4%). Low-coverage tests are the category where the defect is an absence: a check that isn't there. Reading traces and failed attempts is poorly placed to see one, because no attempt fails.
- The overlap statistic is a raw match rate. "Of the categories the agent pipeline flagged, reviewers' judgments overlapped in 74% of cases." It isn't chance-corrected, and the post gives no inter-rater statistic among the five humans. That is the consistency-for-validity gap LLM-Judge Validation measures elsewhere.
- One sentence is ambiguous. "In no flagged task was 'not broken' the most common human label" can't refer to all 286 filter-flagged tasks, because humans called only 249 of those broken. It most plausibly means the tasks the agent path flagged as broken. On that reading, the agent's 200 are largely inside the human 249.
- Both counts are lower bounds on prevalence. The 445 tasks the filter did not flag were never deep-reviewed, so the filter's recall bounds both numbers from above.
Evidence and conflict of interest#
The tier is empirical, kept. These are measured counts from a stated two-path procedure with a named worked example. Four qualifiers travel with it:
- Nothing is released. There's no per-task list and no labels, so the ~30% can't be audited.
- No agreement statistic beyond the 74% overlap.
- The chart can't be fully reconciled. Its ten values sum to 64.0 against a prose total of 27.4 + 34.1 = 61.5, whatever the series pairing. Multi-label counting on the human side is the likeliest reason, but the page doesn't say.
- Structural COI. OpenAI reports scores on external benchmarks with each release, and it is the party declaring this one invalid. SWE-Bench Pro is Scale AI's, not a competitor lab's. The retraction also reverses OpenAI's own earlier advice, which is some evidence against motivated reasoning. The Preparedness Framework rationale ("these results inform OpenAI's deployment and safety decisions") is the lab's stated reason.
Connections#
- The Price of Fixed Capability — SWE-bench Verified has the slowest fixed-performance cost decline in Epoch's 2026 price-trend report (27.5% per quarter against a 47% average). A score ceiling distorted by defective tasks is one candidate explanation
- Evaluation-Time Answer Leakage — the run-time leakage audit of the same 731 tasks, whose task-quality repair half this page compares against OpenAI's prevalence estimate. Defects and leakage compound: an overly strict test is the case where retrieving the reference fix is the only way to pass
- Benchmark Contamination and Decontamination — the training-time channel. OpenAI's retirement of SWE-bench Verified cited contamination and design flaws together. This page is the design half, and it needs no model access to find
- Measuring Beyond Accuracy Saturation — the task-level validity threat that page names ("impossible-to-solve SWE-bench tasks"), here with a prevalence attached. The CORE-Bench v1.1 repair and SWE-Bench Pro Verified are the two maintenance passes, and neither was systematic over the whole task set
- Agent-Generated Test Quality — the same two test pathologies seen from the authoring side. Tests written to validate one specific change over-specify (strict) or under-check (low coverage), whether a human contributor or an agent wrote them, and a benchmark mined from PRs inherits whichever the PR had
- Task Time-Horizon Scaling — the SWE-bench-family saturation curve it cites, now with a third deflator beside leakage and contamination: roughly a fifth to a quarter of one benchmark's tasks penalise correct work
- LLM-Judge Validation — agent auditors are judges too. The 74% agent–human category overlap is a raw match rate with no chance correction, and that page's main finding is that such rates overstate agreement
- OpenAI — the auditor, and the lab that retracted its own benchmark recommendation
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them? — the cluster synthesis, whose "maintenance loses to publicity" residual risk this page supplies the systematic counterpart for
Open Questions#
- Does SWE-Bench Pro Verified's repaired set still fail an OpenAI-style audit, and at what rate? Zheng et al. repaired 102 tasks chosen from public complaints and never examined low-coverage tests. OpenAI flagged 249 by human review but released no list. Falsifiable: run OpenAI's two-path procedure (or any five-reviewer panel) on the Verified release and count broken tasks by category. The prediction from the table above is that low-coverage defects survive at roughly the original rate.
- How much of a frontier SWE-Bench Pro score on the overly-strict subset is retrieval? Overly strict tests reward reproducing the gold patch. Falsifiable: split Zheng et al.'s Baseline → Verified drop by OpenAI's per-task category. If strict-test tasks lose disproportionately under isolation, the two defects compound as argued above. It needs OpenAI's per-task labels, which are unreleased.
- Is the agent path's low-coverage undercount structural? The prediction is that an auditor reading failures misses defects that produce no failure. Falsifiable: seed known low-coverage tasks (delete assertions from a clean task) and measure the agent path's recall against the other three categories.
Sources#
- Separating signal from noise in coding evaluations — OpenAI (unbylined), Separating signal from noise in coding evaluations, openai.com, 2026-07-08,
empirical. The headline counts (286 / 200 / 249, 27.4% / 34.1%, ~30%), the 23.3% → 80.3% trajectory, the four-way taxonomy and its footnoted former names, the methodology, the 74% overlap, the low-coverage 9.4% / 4.1% gap, the OpenLibrary-77c16d5 example, and the retraction are all from prose. Chart warning: the "Share of Dataset Flagged by Issue Type" values were transcribed from renderedinnerText, with the series pairing inferred from label order. Only the low-coverage pair is confirmed by prose. The agent column sums to 27.3 (≈27.4 with rounding), but the human column sums to 36.7 against a stated 34.1, and all ten values sum to 64.0 against 61.5, so no re-pairing reconciles it. Chart cells other than low-coverage are marked ≈ and should not be quoted as exact. Only the first of the page's four failure-mode tab examples was captured. - SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents — Zheng et al. (ECNU / Shanghai AI Lab / Fudan), arXiv 2609.08149, 2026-09-08,
empirical. Cited here for the Table 2 defect mix, the Table 9 field-modification shares and the 89.06 → 62.93 leader drop. Full treatment on Evaluation-Time Answer Leakage
Cited by 11
- Evaluation-Time Answer Leakage×4
separating signal from noise coding evaluations — OpenAI (unbylined), Separating signal from noise…
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them?×2
And maintenance loses to publicity, then to an adversary. The first such pass observed under load…
- Measuring Beyond Accuracy Saturation×2
Task-level threats — the headline metric does not faithfully measure the intended capability. This…
- The Price of Fixed Capability×2
Benchmark Task Defects — one possible reason SWE-bench Verified's cost trend is the slowest measured
- Task Time-Horizon Scaling×2
And part of it runs through broken tasks. OpenAI's own audit of the same 731 SWE-Bench Pro tasks…
- Agent-Generated Test Quality
Benchmark Task Defects — the same two test pathologies, overly strict and low-coverage, found in…
- Benchmark Contamination and Decontamination
Benchmark Task Defects — the design half of the SWE-bench retirements. OpenAI dropped SWE-bench…
- LLM-Judge Validation
Benchmark Task Defects — agent auditors are judges, validated here the way this page says not to.…
- Evals & Benchmarks
Benchmark Task Defects — Benchmark instances whose prompt and hidden tests disagree, so a pass or a…
- Open Questions Backlog
Benchmark Task Defects ×3 (oldest 4d) — Does SWE-Bench Pro Verified's repaired set still fail an…
- OpenAI
Benchmark Task Defects — its July 2026 audit of SWE-Bench Pro: ~30% of the 731 public tasks broken…
Related articles
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Evaluation-Time Answer Leakage
The channel by which an agent retrieves the reference solution *during* a benchmark run — residual Git objects, hidden…
- Production-Sourced Evaluation
Building benchmarks from de-identified real production usage rather than synthetic or hand-authored tasks; DRACO's cent…
- Benchmark Score Redundancy
Zeng & Papailiopoulos: an 84-model × 133-benchmark public score matrix is effectively rank-2, so BenchPress matrix comp…
- How Much Signal Do Public Benchmarks Still Carry — and What Replaces Them?
Synthesis of the 2026 eval-science cluster: public benchmark suites carry far less independent signal than their count…
