H
Howardism
Plate IIAI Coding Practice中文HOWARDISM

Agent Review Comment Resolution

Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex in 341 Python GitHub repos — first large-scale look at the review loop running the OTHER way: the agent reviews, the human decides. ~71% resolved (Copilot 72.9%, Cursor 67.2%, Codex 54.8%), core developers do most of the resolving (78.1% of Copilot's), an inline suggestion is the strongest predictor (OR 1.62), longer comments fare worse — but AUC 0.58, so most of what decides adoption is not in the comment. Card sorting 470 unresolved-but-argued discussions: modal failure is project context the agent cannot see (23.8%), then confident false positives (63); hallucination 4 of 470; 24.3% acted on but never marked resolved. Contradicts Goldman et al. (ASE 2025) — 4,000 RovoDev comments, 1,007 Atlassian repos, 39.3% — a construct gap, not population (reconciled 2026-09-22)

Article metadata
Publication details
Published:August 12, 2026
Filed:Concept
Domain:AI Coding Practice
Reading:42 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Agent Review Comment Resolution

Sources#

Summary#

Cynthia, Widyasari, Roy, Zhang & Lo (University of Saskatchewan / Singapore Management University / Monash, arXiv 2607.21997, July 2026) run the first large-scale study of what happens after an AI agent posts a review comment. They mine 54,713 agent-generated review comments across 341 Python GitHub repositories and ask three questions: what fraction gets resolved, who resolves it, and what properties of a comment predict that it will be.

Read the direction carefully — it is the opposite of every other review page in this vault. Review as the Control Point, Security Debt of Agent-Generated Code and Efficiency Debt of AI-Generated Code all measure humans reviewing agent-authored code, and their efficacy question is "does the reviewer catch the defect." This paper measures an agent reviewing a pull request, and its efficacy question is "does the human act on the agent's finding." The two loops share vocabulary and almost nothing else. A resolved comment does not mean a defect existed; it means a project collaborator agreed enough to close the thread. Note also that the paper never establishes who wrote the code being reviewed — these are PRs in active Python repos that a review bot commented on, human- and agent-authored alike, never separated.

Evidence note. empirical, confirmed on full read, with four qualifications that travel with every number. (1) "Five agents" is the collection figure, not the analysis. 54,791 comments across 342 repos were collected from Copilot, Cursor, Codex, Devin and Claude — but Devin (50 comments) and Claude (28) were dropped as too sparse for inference, leaving 54,713 from three agents across 341 repos. Nothing here characterizes Claude or Devin as reviewers. (2) It is largely a Copilot study. Copilot is 45,668 of 54,713 comments (83.5%; the paper says 86% in one place), the authors state the pooled regression "largely mirrors Copilot's interaction patterns," and per-agent models for Cursor and Codex fail to converge on sparsity and separation. (3) A layer of the analysis is an LLM judge. All 15-way comment categories, the seven explanation types, and the relevance/clarity/conciseness scores come from Llama-3.1-70B, chosen over GPT-4o and Qwen3-8B against a 100-comment human gold set (kappa 0.74 for categories against a 0.86 human-human baseline; Jaccard 0.90 for multi-label explanation types, retaining only labels with confidence >= 0.9). Substantial agreement, not ground truth. (4) The resolution construct is GitHub's isResolved thread flag — the authors say so in their own threats section, and their card sort proves it cuts both ways (see below).

RQ1: roughly seven in ten comments get resolved#

(Table III, verified against the PDF.)

AgentCommentsResolvedResolution rate
Copilot45,66833,26572.9%
Cursor6,7784,55467.2%
Codex2,2671,24254.8%
Pooled54,71339,06171.4%

Fix the abstract's arithmetic before quoting it. The abstract reads "Copilot accounting for the majority of resolved comments (72.9%)," which welds two different quantities. 72.9% is Copilot's resolution rate; Copilot's share of all resolved comments is 33,265/39,061 = 85.2%, and that is a statement about market share in this corpus, not about quality. Both numbers are real and neither is the other.

The 18-point spread between Copilot and Codex is the paper's headline, and its own explanation is that the spread is not a property of the comments. When the regression controls for comment characteristics the agent-level gap persists, so the authors attribute it to "differences in how agents are integrated into the review workflow, as well as agent integration and developer familiarity" — i.e. to invocation mode and habituation rather than to review quality. Treat the ordering as an adoption measurement with heavy confounds (no controls for repository composition, task mix, or how each agent is triggered).

What each agent comments on differs more than how often it lands. Copilot's resolved comments spread across categories — solution approach 21.8%, documentation 19.8%, functional defect 18.1%, visual representation 11.3% — while Cursor and Codex are near-monomaniacs about functional defects (88.0% and 88.6% of their resolved comments respectively; Codex's runner-up category is resource issues at 2.9%). A general-purpose reviewer and a defect detector are being compared on one axis, and the defect detector has the lower resolution rate.

RQ2: core developers do the resolving, peripheral developers do the defects#

Restricting to cases where the resolver is also the PR author (Table IV), core developers — the top 20% by authored-plus-reviewed closed PRs, computed per repository rather than globally — account for 78.1% of Copilot resolutions, 54.6% of Cursor's, and 58.4% of Codex's, against a population of 1,538 core and 2,554 peripheral developers.

By category (Figure 4), core share runs 70.5% to 79.8% across all fifteen categories — core developers resolve the majority of everything. The variation is in where peripheral developers show up most: functional defects (29.5% peripheral, n=8,481) and visual representation (28.1%, n=3,167) at the top, solution approach (20.2%, n=6,262), organization of code (20.3%) and naming (20.9%) at the bottom. Design and evolvability feedback needs project familiarity to act on; a defect fix does not.

This is Returns to Expertise in Agentic Coding on the receiving end of review: the ability to use an agent's design feedback is the same project-specific knowledge that lets you overrule it.

The ten patterns behind non-resolution — and the ones the parse hid#

Of 15,652 unresolved comments, only 1,056 (6.7%) received any reply, 655 of them from PR authors. From the 431 core-replied and 224 peripheral-replied comments the authors drew a stratified sample of 470 discussions (284 core, 186 peripheral) and open-card-sorted them into 10 categories and 13 sub-categories, at inter-rater kappa 0.85.

(Table V below is recovered from the PDF; the docling parse of this table is damaged — see Sources. Every row reconciles: sub-rows sum to parents, core plus peripheral sums to the total, the per-agent breakdowns sum to the total, and 466 mapped plus 4 unmappable equals 470.)

PatternDiscussions (Copilot/Cursor/Codex)CorePeriph.
Accepted Agent Feedback114 (66/31/17)5460
— Accepted Suggestion783543
— Addressed Afterwards361917
Intentional Design Decision112 (55/48/9)8131
— Context-Specific Implementation644420
— Expected Behaviour372611
— Developers' Preference11110
Incorrect Suggestion67 (36/19/12)3235
— Factually Wrong or False Positive632934
— Agent Hallucination431
Challenging Agent Feedback39 (22/12/5)2514
— Disagreement with Suggestions271413
— Questioning Agent's Claim12111
Needs Further Discussion36 (19/10/7)2016
Delegation of Work31 (22/6/3)247
— Agent Task Delegation23185
— Delegating to Developer862
Suggestions Deferred25 (12/9/4)169
— Future Improvement21156
— Deferred to Future PR413
Acceptable Trade-Offs19 (9/10/0)136
Missed Existing Fix12 (9/2/1)93
Dismissed as Low-Value11 (7/4/0)83

Four things in that grid matter more than the taxonomy:

1. A quarter of "unresolved" is a measurement artifact. Accepted Agent Feedback — the developer applied the change and never marked the thread resolved — is the largest category at 114/470 (24.3%), and it is the one category that skews peripheral (60 vs 54). GitHub's isResolved flag is a hygiene habit, and less-embedded contributors have it less. So the 71.4% resolution rate is a floor on adoption, not an estimate of it. (A useful cross-check: RQ3 rebuilds the label from scratch — excluding AI-only resolutions and folding in explicit acknowledgments like "fixed in commit...", "done", "updated" — and lands at 37,512/53,086 = 70.7% useful. Two differently-constructed denominators, effectively the same number, which suggests the acknowledgment correction is small in aggregate even though it dominates this argued sample.)

2. The modal genuine rejection is context, not error. Intentional Design Decision is 112/470 (23.8%) and heavily core-skewed (81/112, 72%): the code is deliberate and the agent could not see why. Its dominant sub-pattern is Context-Specific Implementation (64) — "No, the idea is to use add logs to the docker stream" — followed by Expected Behaviour (37). The agent is not wrong about the code; it is wrong about the project.

3. Hallucination is nearly absent. Confident wrongness is not. Incorrect Suggestion is 67, but the split inside it is 63 factually-wrong / false-positive against 4 hallucinations — 0.9% of the sampled discussions. The popular account of AI review failure ("it makes things up") is, at this scale, the rarest failure mode measured. The real one is an agent that reads real code and draws a wrong conclusion about it. Note that this sub-category is exactly what the damaged parse concealed: docling welded the 4/3/1 row into its neighbours, inflating hallucination into a 39-discussion category. This single row is why the table had to be recovered rather than cited.

4. Rejecting bad agent feedback does not require seniority. Incorrect Suggestion is the one major category that leans peripheral (35 vs 32), and the authors read it the right way: spotting that a suggestion is wrong "may not require deep project knowledge," whereas knowing that an implementation is deliberate does.

Agent-level, the two dominant patterns split cleanly along the RQ1 result: Intentional Design Decision is disproportionate for Cursor (32% of its discussions vs Copilot 21.4%, Codex 15.5%) — a context-awareness gap despite its functional focus — while Incorrect Suggestion is elevated for Codex (20.7% vs Copilot 14.0%, Cursor 12.6%), consistent with Codex's lowest resolution rate.

The sampling frame is the taxonomy's hard limit, and the paper does not say so. The 470 discussions are drawn from the 1,056 unresolved comments that got a reply. The other 14,596 unresolved comments (93.3%) were met with silence and are unrepresented in every percentage above. This taxonomy explains why developers argue with agent review comments. It says nothing about why they ignore them — which is the majority behaviour by an order of magnitude, and the behaviour a reviewer-fatigue or noise hypothesis would predict.

RQ3: actionability wins, and the model still barely predicts anything#

A logistic regression over six comment characteristics, with a comment counted useful if a human resolved it or explicitly acknowledged the fix (37,512 useful vs 15,574 non-accepted, n = 53,086):

PredictorOR (all)FunctionalEvolvability
Inline code suggestion1.6171.6721.459
log comment length0.9260.855—
has explanation0.593— (n.s.)0.578
explanation: Rule1.144—1.102
explanation: Benefit1.0871.066—
explanation: Example1.080——
Relevance / Conciseness / Clarity1.046 / 1.059 / 1.0221.044 / — / —1.051 / 1.077 / 1.033

Univariately, comments carrying a GitHub suggestion block resolve at 75.5% against 64.6% without one (Chi-squared p < 0.05, Cramer's V = 0.12 — a small effect), and the gap holds in both issue groups. Non-accepted comments are longer (mean 807 characters, median 405) than useful ones (mean 617, median 356), significantly but with a negligible effect size; the length penalty is real only in the functional group (non-accepted mean 1,240 vs useful 966) and vanishes in the evolvability group. Writing-quality scores separate the groups by at most one median point (conciseness 7 vs 6) with negligible effect sizes throughout.

The deflator the paper puts in its threats section: AUC = 0.58. Every odds ratio above is statistically significant on 53,086 rows and the model as a whole barely beats a coin flip at ranking which comment will be acted on. The authors' own conclusion — "unobserved factors such as reviewer authority or project context may influence outcomes" — is the same finding as the card sort, arrived at statistically: what decides adoption mostly is not a property of the comment. Anyone tuning a review agent's prose against this table is optimizing the small residual.

One internal tension worth not inheriting. has explanation carries OR 0.593 — explaining yourself hurts — and the paper reports it as a headline ("concise, action-oriented comments are more readily adopted"). But its own Table VII debunks the univariate version of exactly this: the no-explanation group resolves at 83.7% and is 89% surface-level categories (documentation 234, visual representation 178, naming 75), which the authors correctly say means "the observed effect is driven by category characteristics rather than the absence of explanations." That caveat is not carried into the regression, which controls for the functional/evolvability split but not for the fifteen-way category. The defensible reading is the type result, which survives: rule-, benefit- and example-based explanations are positively associated, scenario- and issue-based ones are not. Explanation style is the finding; explanation presence is probably composition.

Usefulness by explanatory richness peaks at two explanation types (71.7%) and declines beyond (three 69.3%, four-plus 65.0%) — the same shape as the length penalty, and the paper's reading is that excessive detail reduces actionability.

Reconciled: this paper against the one it cites#

The introduction cites Goldman et al. (ASE 2025) for the claim that "approximately 60-70% of LLM-generated comments remain unresolved." This paper measures 71.4% resolved — a roughly two-fold disagreement on the central quantity between two empirical studies a year apart, cited without reconciliation. The vault holds neither Goldman's population nor its resolution construct. (superseded 2026-09-22 — Goldman was ingested to settle it, and it does: the two papers never measured the same event.)

Goldman, Lin, Pasuksmit, Thongtanunam, Tantithamthavorn et al. (The University of Melbourne + Atlassian + Monash, arXiv 2510.05450, ASE 2025) draw 4,000 LLM-generated review comments across 3,746 pull requests in 1,007 Atlassian internal repositories, over a two-month window (June–July 2025), posted by Atlassian's own RovoDev Agent (via Claude 3.5 Sonnet) and classified by a GPT-4.1 judge into a five-category taxonomy. Their resolution construct, verbatim:

"We consider a given code review comment to be resolved if a subsequent commit modified the exact line where the comment was placed."

(Figure 5, recovered from the page image. The design and no-issue fractions appear nowhere in the prose, and the overall rate appears nowhere in the paper at all; the five denominators sum to exactly 4,000, which is the reconciliation.)

CategoryResolved / commentsRate
Code readability586 / 1,35443.3%
Code bugs399 / 95241.9%
Maintainability483 / 1,33336.2%
Code design101 / 35328.6%
No issue2 / 825.0%
All comments1,571 / 4,00039.3%

The comparable pair is therefore 39.3% (Goldman) against 71.4% (this paper) — a 32-point gap on a common scale, which is the thing that needs explaining. Goldman's stated "many of the LLM-generated comments are not resolved by developers (60%-70%)" is a loose band around the 60.7% its own figure gives exactly.

Construct carries the gap; product carries a bounded part; population pushes the other way#

Population pushes the wrong way, so it cannot be the carrier. This page's RQ2 is that core developers do the resolving — 78.1% of Copilot's, a majority in all fifteen categories — and that peripheral contributors both act less and mark less. Goldman's population is 1,007 internal repositories staffed entirely by paid employees under a mandated review process: core developers, all of them, with none of the drive-by-contributor dilution that depresses the number here. If population were driving the gap, the industrial figure should be the higher one. It is 32 points lower. Population may dampen the gap; it cannot generate it in the observed direction.

Product is real but bounded at roughly a third of the gap. The strongest single predictor on this page is a GitHub inline suggestion block — 75.5% resolved with one against 64.6% without, OR 1.617 — though at Cramér's V = 0.12, a small effect. Copilot, Cursor and Codex post into GitHub's native review UI with one-click apply; RovoDev posts line-anchored comments and no applicable diff is described anywhere in the paper. That channel is worth on the order of 11 points, and only as a ceiling, since not every comment counted here carries a suggestion block either.

The construct difference is a difference in kind, and it takes the residual. Goldman requires a code diff landing on one specific line in a later commit. This paper requires a thread flag to be flipped, by anyone, for any reason. These count different events, and their errors run in opposite directions:

  • Goldman scores as unresolved a developer who agreed and fixed the problem anywhere other than the anchored line — at the call site, by extracting a function, by restructuring — and one whose anchored line was deleted outright, and one whose PR merged with no further commit. It also states no follow-up horizon: the pool is a June–July 2025 window, so late-window comments are right-censored by however soon the data was pulled. Every one of those mechanisms is a one-way under-count of agreement.
  • This paper scores as resolved a thread closed by dismissal with no code change at all, and misses acceptance that was never marked — its own card sort puts 24.3% of argued-unresolved discussions in Accepted Agent Feedback.

The sharpest evidence that the construct is load-bearing sits inside Goldman's own category ordering. Readability (43.3%) > bugs (41.9%) > maintainability (36.2%) > design (28.6%) is also, exactly, the ordering of how line-local a fix is. The paper reads the design deficit behaviourally — design issues "involve more complex, architectural considerations that require deeper understanding and broader changes, making them harder to resolve" — without noticing that broader changes are precisely what an exact-line proxy is built to miss. Its headline category result is confounded with its measurement instrument, and only holding the two papers side by side makes that visible.

The honest limit on this attribution. It is a reading of two published constructs, not a re-measurement of one construct across two populations. Nobody has run Goldman's exact-line rule over GitHub threads, or GitHub's flag over Atlassian's corpus, and either would decompose the gap properly. What is now safe to say is stronger than what the vault had: neither number is an adoption rate, and the two bracket one rather than estimate it. 39.3% is a floor that demands a diff at an anchor; 71.4% is a disposition count that admits "won't fix". Quoting either as the resolution rate for agent review comments is a category error, and averaging them is worse.

What the same source says about what an LLM reviewer chooses to comment on#

Goldman's RQ1 is a separate measurement from the resolution rate, carrying its own caveat: it compares the same RovoDev agent against human reviewers on two corpora, and the two disagree, so neither pattern generalizes on its own.

CategoryAtlassian internal — LLM / humanOSS (CuRev) — LLM / human
Code bugs18.2% (128/702) / 6.5% (30/465)20.1% (253/1,256) / 18.1% (181/1,000)
Maintainability26.8% (188/702) / 19.1% (89/465)37.9% (476/1,256) / 23.6% (236/1,000)
Code readability33.2% (233/702) / 46.0% (214/465)23.7% (298/1,256) / 24.5% (245/1,000)
Code design21.2% (149/702) / 19.4% (90/465)15.7% (197/1,256) / 23.6% (236/1,000)
No issue0.6% (4/702) / 9.0% (42/465)2.5% (32/1,256) / 10.2% (102/1,000)

(Figures 3 and 4, read as images; the prose supplies only six of these twenty cells. Every column sums to its denominator.)

On Atlassian's own code the LLM reviewer files 2.8× the human share of bug comments (18.2% vs 6.5%) and 1.4× the maintainability share, while humans put nearly half their comments on readability. That is the paper's complementarity claim — and the scope restriction matters: it is Atlassian-internal only. On OSS the bug advantage collapses to 20.1% vs 18.1% and the design relationship inverts, with humans filing more design comments (23.6% vs 15.7%). Design is near-parity at Atlassian too (21.2% vs 19.4%), so even in the corpus where complementarity holds it rests on bugs and readability, not on design.

The finding that survives both corpora is negative and about the agent: the LLM reviewer under-produces bug comments relative to its own other categories — 17.8 points below maintainability on OSS, 15.0 below readability at Atlassian — which is the paper's actual claim, not a claim that it files fewer bug comments than humans do. Set beside this page's RQ1, where Cursor and Codex put 88% of their resolved comments on functional defects and still resolve lower than generalist Copilot, the corpus now holds two industrial-scale measurements whose category mixes differ by an order of magnitude on the same axis and whose adoption rates do not order with them. Category mix looks like a product decision, and it does not predict adoption in any consistent direction.

One caveat the paper does not draw, and it weakens the link between its two RQs. The RQ2 pool is composed differently from the RQ1 pool: design is 21.2% of the 702 RQ1 Atlassian comments but only 353/4,000 = 8.8% of the RQ2 sample, and no-issue falls from 0.6% to 0.2%. Different window, different draw. So the resolution rates are measured over a mix that is not the mix the complementarity claim describes, and the two results should not be multiplied together into an "expected value of an LLM reviewer".

What this does and does not settle about review efficacy#

The vault carries a standing question, on Review as the Control Point and Security Debt of Agent-Generated Code, about whether review coverage buys catch-rate: agent-PR review coverage is converging toward the human baseline while measured efficacy on leaked credentials sits at 18.9%. This paper does not answer that question, and the resemblance is a trap. It measures the other loop — the agent as reviewer, the human as decider — so its 71.4% is an adoption rate for the review layer's output, not a catch rate for defects in agent code.

What it does contribute is the missing half of a two-sided picture of why the automated review layer under-delivers, and the two halves fail differently:

  • Non-detection (Security Debt of Agent-Generated Code): on hard-coded credentials, the smell class with seven purpose-built commercial detectors present, bots and humans together commented on 18.9% of genuine live credentials. The layer mostly does not fire.
  • Non-relevance (here): when the layer does fire, roughly seven comments in ten are closed, and the modal reason the rest are not is that the agent misread project context (23.8%) or was confidently wrong about real code (13.4%). Only 11 of 470 argued cases were dismissed as low-value noise.

Neither failure mode is human inattention. That is the finding that should update the vault's rubber-stamping prior most: in the corpus of unresolved-and-argued comments, developers are reading closely enough to catch the agent being wrong, and doing it disproportionately when they are peripheral contributors with the least project context to lose. This does not contradict the one over-time measurement of the prior. Reviewer Habituation on Agent Pull Requests finds approval of agent-authored PRs tilting upward with each reviewer's exposure. That is the other direction of the loop, and the tilt is small (d = 0.25), which is compatible with reviewers who are still engaged.

The convergence with Efficiency Debt of AI-Generated Code is the sharpest one available. Tran et al. diagnose AI authoring failures as "generalist LLMs lacking the highly specific context of internal enterprise monorepo structures" and prescribe injecting repository context at prompt time. The dominant failure of AI review measured here is the same missing context wearing the other hat — an agent flagging as a defect what the team decided on purpose. One context gap, two symptoms, one prescription.

Practitioners name noise as the problem; this study finds almost none (2026-10-01). Orosz's survey (practitioner-opinion) states that "noise is a big problem with AI code reviews". WeTravel turned AI review down on noise, and a June 2026 re-evaluation still fell short. Here, only 11 of 470 argued cases were dismissed as low-value noise. The two do not conflict, because they count different things. The 470 are comments a developer argued with. A comment dismissed as noise is usually ignored rather than argued, so noise would sit among the 14,596 unresolved comments (93.3%) that got no reply and were never sampled. The empirical source is the better evidence on why argued comments fail. It is not evidence that noise is rare. The same survey's preferred loop keeps this page's direction but narrows the human's job: Weaviate's CTO has an adversarial agent review, a human make "critical scope decisions" on its findings, and an agent implement them. That turns the per-comment resolve-or-reject decision into a scope call, which is where the 23.8% project-context failure would be caught.

Connections#

  • Closed-Loop AI Review — the population these 54,713 comments are a sample of, and the half of the loop that never reaches a human. Selvanayagam & Ghaleb count the review side of public GitHub at 8,560,237 attributed review-comment entries and 4,141,107 review entries, on 248,641 agent-authored PRs with at least one AI-attributed review. Two things compose with this page. The scale says the adoption measurement here is of a layer that is large and growing two orders of magnitude a year rather than a niche. And the composition complicates the construct: 208,145 of those PRs were reviewed by the product that wrote them, a configuration in which "resolution" may be an automated product workflow rather than a developer's decision, which is not separable in either dataset. Neither study measures correctness
  • Same-Model Review Blindness — the other half of automated review's efficacy, and the two numbers must not be pooled. This page measures adoption of an agent reviewer's output (71.4% of comments resolved, with the annotating judge's kappa published); Greptile measures its recall against a bug ground truth (52–62% of high-severity bugs named, with no judge validation and a vendor-built label set, case-study). A comment that lands is not a bug that was caught, and a bug that was caught is not a comment that lands — held together they price the layer end to end for the first time, roughly half the serious bugs named and roughly seven in ten of the resulting comments acted on, across different corpora, different agents and evidence tiers two apart. The failure mechanisms are complementary rather than rival: the dominant non-adoption cause here is a reviewer missing project context, and the dominant recall deficit there is a reviewer sharing the author's model priors
  • Review as the Control Point — the mirror direction, and the first sizeable measurement of that theory's second moderator. Its automated-reviewer-capability moderator is here given an adoption number (54.8-72.9% by agent) and a mechanism for where capability actually binds: project context, not correctness or prose. It confirms neither P8 (throughput) nor P9 (quality/security), because this study measures neither — but it does establish that the pooled agent-level spread survives controlling for comment characteristics, which is a moderator effect rather than a message-quality effect
  • Security Debt of Agent-Generated Code — the two failure modes of the automated review layer, measured from opposite ends: that page finds it does not fire (18.9% comment rate on genuine live credentials), this one finds that when it does fire ~71% of the output is acted on and the residue is context error rather than human neglect. Together they relocate the review layer's weakness from attention to detection-and-relevance
  • Risk-Tiered Auto-Approval — independent convergence on comment design. StampHog's approval is a bare GitHub approval with no line comments and its refusal is 1-2 sentences plus a risk level and next steps; this paper measures why that shape is right, with an inline suggestion the strongest resolution predictor (OR 1.62) and length carrying a penalty concentrated in functional feedback (OR 0.855). It also validates demoting the model to a veto rather than a commenter: a veto has no resolution rate to lose
  • Efficiency Debt of AI-Generated Code — the same missing-project-context diagnosis on the authoring side. Tran et al. attribute AI's imperative bias to "generalist LLMs lacking the highly specific context of internal enterprise monorepo structures" and prescribe context injection; here that same gap surfaces as an agent reviewer flagging deliberate design as a defect in 23.8% of argued discussions. It is also the counterweight on evidence class: that study has a human control cohort and this one has none, so "is an agent reviewer worse than a human reviewer" remains unmeasured
  • Acceleration Whiplash — Faros AI records agentic review going 0% to 25% of PRs, faster than agentic authoring; this measures whether that layer's output lands. On adoption it does (~71%), which is a mild counterweight to the whiplash framing — but the pooled figure is 83.5% Copilot, and neither study connects review-comment adoption to any downstream quality outcome
  • AI as Primary Author — the oversight layer automating faster than the authoring layer is the precondition for this page existing at all; this supplies the first population-level number for whether that oversight layer's output is acted on
  • Verification as the New Bottleneck — an adoption datum for "how far do you push fully automated reviews": the output of a deployed automated reviewer is acted on roughly seven times in ten, the strongest lever on that is an applicable diff rather than better prose, and an AUC of 0.58 says most of the decision lives outside the comment
  • LLM-as-a-Judge — a published calibration on a fifteen-way code-review classification: open-weight Llama-3.1-70B at kappa 0.74 against a human gold set beat GPT-4o at 0.70 and crushed Qwen3-8B at 0.38, with the human-human ceiling at 0.86; multi-label explanation typing hit Jaccard 0.90 after retaining only confidence >= 0.9. A judge chosen on measured agreement rather than brand, with the measurement published — and, since 2026-09-22, half of a matched pair: Goldman runs the same task with GPT-4.1 at kappa 0.42 against a human-human baseline of 0.80/0.86 in the same sanity check, and deploys it on 4,000 comments on the judgement that this is "sufficient". Two code-review-comment judges a factor of ~1.8 apart in chance-corrected agreement, both reported as adequate
  • Returns to Expertise in Agentic Coding — the receiving end of review: acting on design and evolvability feedback is core-developer work (79.8% of solution-approach resolutions), while functional-defect fixes are where peripheral contributors show up most (29.5%). The project knowledge that lets you use an agent's design suggestion is the same knowledge that lets you overrule it
  • Optimizer–Evaluator Decoupling — the deployed instance of the split at population scale: a reviewer agent that never authored the code, whose comments carry no authority to merge and must persuade a human. Its measured failure is exactly the cost of that decoupling — the evaluator lacks the author's project context, and 23.8% of argued rejections are about precisely that
  • Post-Acceptance Edit Behavior — the two behavioral instruments for how humans dispose of machine output, at opposite granularities and landing on the same diagnosis. Here ~71% of agent review comments get resolved and the modal genuine rejection is project context the agent could not see (23.8%); there, the AI completions most likely to be deleted are the ones that "subtly do not align with a developer's intent or programming context." One context deficit, two artifacts. The shared methodological limit is worth naming too: both measure adoption, never correctness — a resolved comment is not a caught bug, and a retained completion is not good code
  • Deterministic Engineering for Agent Code Review — a third construct joins adoption and recall, and it is the one most exposed to the ground-truth-match trap this page already warns about. OpenCodeReview measures precision against AACR-Bench's expert-verified reference set (7.23%–37.80% across twelve configurations) — a different quantity again from this page's 71.4% resolution rate and Same-Model Review Blindness's 52–62% recall, and all three must stay unpooled. Its own §6 concedes exactly this page's construct warning in reverse: a genuinely useful comment matching no ground-truth item scores as a false positive there, the mirror of this page's "only 11 of 470 argued discussions dismissed as low-value." The two converge hardest on mechanism — the modal reason OpenCodeReview's reflector and comments miss is project context outside the diff, the same failure this page finds behind 23.8% of argued non-resolutions — which is exactly what OpenCodeReview's bounded file_read_diff and per-file SubAgents are built to recover, and, on this evidence, still don't fully
  • Open Source Under Agent Contributions — the same inverted loop at the repository gate rather than inside a PR: DHH has agents triage every incoming PR and returns only a merge/don't-merge decision
  • Writer/Reviewer vs Agent-to-Agent Review — this page supplies the adoption half of a four-axis comparison of in-session Writer/Reviewer against PR-posted agent review, and its card sort is the corpus's only false-positive-cost data (63 factually wrong, 4 hallucinations, 11 of 470 dismissed as noise). The synthesis's constraint on anyone comparing the two patterns: 71.4% resolution and 7.23–37.80% reference-match precision are an order of magnitude apart and neither substitutes for the other, so a head-to-head reporting one metric alone is uninterpretable
  • Context Smells — the 23.8% Intentional Design Decision category is Houck's lost in the details smell measured (the code was deliberate and the rationale was not visible to the agent). It is the strongest empirical support for his context-not-model thesis in the vault
  • Reviewer Habituation on Agent Pull Requests — the other direction of the review loop, measured over time: humans reviewing agent authors approve more as their exposure grows. The drift is small enough to coexist with the engaged reviewers found here, so the two are not in conflict

Open Questions#

  • The 470-discussion taxonomy is drawn only from unresolved comments that received a reply — 6.7% of the unresolved population. What explains the silent 93.3%? Reviewer fatigue, comment volume, triage, or the same context error just not worth arguing about? The taxonomy's shape (only 11 of 470 dismissed as low-value) would look very different if noise dominated the silent majority, and the two hypotheses are distinguishable by sampling the silent set directly. Context from What Types of Code Review Comments Do Developers Most Frequently Resolve? (2026-09-22), which is not an answer: its 4,000-comment pool applies no reply filter, so its 60.7% non-resolution covers the whole unresolved population including the silent majority — establishing that the silent majority is an industrial phenomenon and not a GitHub-etiquette artifact. But its only account of why is an unmeasured discussion-section conjecture ("the lack of clarity, relevancy, and simplicity of the comments"), with no card sort, no reply analysis and no sampling of the silent set. The question survives intact.
  • Resolution is adoption, not correctness. Does a resolved agent comment correspond to a defect that would otherwise have shipped — and does an agent reviewer change any downstream outcome (escaped defects, incidents, revert rate) against a no-agent-reviewer baseline? No study in the corpus has run this on either side of the review loop; it is the same missing outcome measurement that leaves Risk-Tiered Auto-Approval's throughput figures unattached to safety. Unchanged, and now doubly so (2026-09-22): What Types of Code Review Comments Do Developers Most Frequently Resolve? measures the same loop at industrial scale with a construct that is further from correctness, not closer — an exact-line modification in a later commit scores as resolution whether or not the comment caused it, so unrelated churn on a hot line counts as agreement — and it measures no downstream outcome either. Two industrial-scale studies, two incompatible constructs, zero correctness ground truth on either.

Resolved Questions#

  • Two empirical studies a year apart disagree by roughly a factor of two on the central quantity: Goldman et al. (ASE 2025) report 60-70% of LLM-generated review comments unresolved, this study reports 71.4% resolved. Is the gap population (industrial vs open-source GitHub), product (in-house pipeline vs shipped agents with one-click suggestion blocks), or construct (whatever Goldman counted vs a GitHub thread flag that this paper's own card sort shows under-counts by ~24%)? Answered 2026-09-22 by What Types of Code Review Comments Do Developers Most Frequently Resolve? — the gap is construct. On a common scale it is 39.3% (Goldman, 1,571/4,000, recovered from his Figure 5 since the paper never states an overall rate) against 71.4% here: 32 points, not a factor of two. Population is ruled out by direction — Goldman's 1,007 internal repos are staffed entirely by core-equivalent paid developers, and this study's own RQ2 says core developers resolve more, so an all-core population predicts the higher number and Goldman has the lower one. Product is bounded at roughly 11 of the 32 points — the inline-suggestion effect measured here is 75.5% vs 64.6% at Cramér's V = 0.12, and RovoDev ships no applicable diff. The residual is the construct, and it is a difference in kind: "a subsequent commit modified the exact line where the comment was placed" is a code-diff-at-an-anchor event that misses every off-line fix, every deleted anchor and everything right-censored by an unstated follow-up horizon; GitHub's isResolved is a disposition flag that admits dismissals and misses unmarked acceptances. Confirmed from inside Goldman's own data: his category ordering (readability 43.3% > bugs 41.9% > maintainability 36.2% > design 28.6%) is the ordering of how line-local a fix is, and he explains the design deficit by saying design needs "broader changes" — exactly what an exact-line proxy cannot see. Caveat carried forward: this is a reading of two constructs, not a re-measurement of one construct on two populations, so the split between the product and construct shares is argued rather than estimated. The durable conclusion is that the two figures bracket an adoption rate rather than estimating one, and neither may be quoted as "the" resolution rate. Full working in Reconciled: this paper against the one it cites above.

Sources#

  • "Go Home Copilot, You're Drunk": Understanding Developer Responses to Agent-Generated Code Review Comments — Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy, Ting Zhang & David Lo (University of Saskatchewan / Singapore Management University / Monash, arXiv 2607.21997, 2026-07-24), empirical. §III (data collection, agent identification by reviewer login pattern, the Devin/Claude exclusion, core/peripheral definition, the LLM-annotator selection and its kappa/Jaccard validation, the usefulness construct), §IV + Table III + Figure 3 (resolution rates and per-agent category distribution), §V + Table IV + Table V + Figure 4 (core/peripheral split and the ten-pattern discussion taxonomy), §VI + Tables VI-VIII (usefulness by code suggestion, length, quality scores and explanation type; the logistic regression), §IX (threats to validity — the AUC 0.58 admission and the resolution-proxy concession). Parse warning: this raw is docling-derived and verify: warn with table-collapse, and the flag is a true positive. Table V (the discussion taxonomy) is damaged two ways at once — the label Agent Hallucinated is welded into the following row's label (Agent Hallucinated Challenging Agent Feedback) while its three values are welded into the preceding row's cells (| Factually Wrong or False Positive | 63 4 | 29 3 | 34 1 |), and the Suggestions Deferred parenthetical migrated into the Future Improvement row. Read as parsed, the Agent Hallucination sub-category disappears and its 39-discussion "parent" is an artifact. The table quoted on this page is recovered via pdftotext -f 7 -layout on and independently reconciled: sub-rows sum to parents, core plus peripheral sums to each total, per-agent triples sum to each total, and 466 mapped plus 4 unmappable equals the stated 470 — and every per-agent percentage in the §V prose (Cursor 32%, Copilot 21.4%, Codex 15.5%, and the Incorrect Suggestion trio) reproduces from the recovered grid. The other seven tables were cross-checked the same way: Tables III, IV, VI and VII are faithful and arithmetically self-consistent; Table I is welded cosmetically (docling shows 12 rows plus an orphan software cell where the PDF has the 15 categories the prose describes — descriptions preserved, so no number is affected); Table VIII's two section-header rows (Functionability issue, Evolvability issue) are splattered across all four columns while its data rows are clean. Four internal inconsistencies belong to the paper, not the parse, and were confirmed against the PDF: the abstract's "majority of resolved comments (72.9%)" conflates a rate with a share; §III states analysis was "restricted to PRs created up to December 2023" while the repository filter requires a PR after December 2024 and every cited example PR is from 2024-2025; the RQ3 answer box gives code suggestion as OR 1.609 against Table VIII's 1.617 and the prose's 1.62, and reports an "OR range: 0.77-0.134" that is not a range; and Table VI's evolvability no-suggestion cell reads 67% where 6807/9873 is 69.0%. None is load-bearing for anything quoted here. Figures 3, 4, 5, 6 and 7 carry per-category distributions that appear in the prose only as selected values; the Figure 3 and Figure 4 numbers quoted above are taken from the pdftotext -layout axis labels rather than read off the rendered chart.
  • What Types of Code Review Comments Do Developers Most Frequently Resolve? — Saul Goldman, Hong Yi Lin, Patanamon Thongtanunam (The University of Melbourne); Jirat Pasuksmit, Kla Tantithamthavorn (also Monash), Zhe Wang, Ray Zhang, Ali Behnaz, Fan Jiang, Michael Siers, Ryan Jiang, Mike Buller (Atlassian, Australia); Minwoo Jeong, Ming Wu (Atlassian, USA). arXiv 2510.05450 v1, 2025-10-06, 6 pages, ASE 2025. empirical — tier kept and earned: 4,000 real LLM-generated comments over 3,746 real PRs in 1,007 production repositories, with resolution computed from actual subsequent commits. §III-B/C (the five-category taxonomy and the GPT-4.1 judge), §IV RQ1 + Figures 3-4 (human-vs-LLM comment-type distributions for Atlassian internal and OSS/CuRev), §IV RQ2 + Figure 5 (the resolution rates and the verbatim resolution construct), §VI (threats — the RovoDev proprietary-architecture disclosure and the 13-OSS/1,007-internal scope statement). COI, disclosed and load-bearing. Twelve of the fifteen authors are Atlassian employees, the LLM reviewer under study is Atlassian's own RovoDev Agent, and the paper carries an explicit disclaimer that its outcomes "should not be constructed as an assessment of the quality of products offered by Atlassian." This is first-party measurement of a first-party product; what protects it is that the headline finding is unflattering (60.7% of its own agent's comments go unresolved). RovoDev's architecture and prompt are stated to be confidential, so the intervention is not reproducible. The classification layer is a GPT-4.1 judge at Cohen's kappa 0.42 (moderate) against a two-annotator gold set whose own human-human agreement in the same check was 0.80 and 0.86 — roughly half the agreement its annotators reached with each other, accepted with the one-line justification that it is "sufficient". The taxonomy underneath it has no inter-rater agreement at all: six Atlassian engineers card-sorted 336 comments collaboratively in one session, and the paper says outright that a kappa "could not be calculated." Parse note: the docling parse is clean on text (6pp, confidence excellent, canary-recall 6/6, one table) but the figure captions are mis-associated, and the error is silent. The raw markdown places image_000001 and image_000002 both under the "Fig. 3" caption and image_000003 under "Fig. 4", while the "Fig. 5" caption has no image beneath it at all. The true mapping, established by reading all four images: image_000001 = Figure 3 (Atlassian RQ1), image_000002 = Figure 4 (OSS RQ1), image_000003 = Figure 5 (RQ2 resolution). A reader trusting the markdown would attribute the resolution chart to Figure 4 and conclude Figure 5 was lost. The mapping is confirmed arithmetically rather than by eye: every column sums to its stated denominator (702, 465, 1,256, 1,000, 4,000), and the two prose deltas reproduce exactly from the images — "17.8% less than maintainability for OSS" is 37.9 − 20.1 from Figure 4, and "15% less than code readability for Atlassian's internal projects" is 33.2 − 18.2 from Figure 3. Figure 5 is the load-bearing recovery: it supplies the design fraction (101/353) and the no-issue row (2/8) that the prose omits, and its five rows sum to 4,000, which yields the overall 1,571/4,000 = 39.3% resolution rate that appears nowhere in the paper and is the number this page's reconciliation turns on. Table I (the taxonomy) parsed faithfully, carries definitions only, and no number is taken from it. One paper-internal error, confirmed against both figures and not a parse artifact. The OSS results paragraph states "The LLM reviewer generated less no issue comments than human reviewers, with only 0.6% falling into no-issue compared to 9.0% for humans." Those are Figure 3's Atlassian values (4/702 = 0.6%; 42/465 = 9.0%). Figure 4's actual OSS values are 2.5% (32/1,256) and 10.2% (102/1,000) — the direction of the claim survives, the magnitudes quoted in it belong to the other corpus. The table on this page uses the figure values.
  • AI-to-AI Code Reviews of GitHub Pull Requests — Selvanayagam & Ghaleb (ÉTS Montréal / Trent), arXiv 2608.21311, 2026-08-21, 15pp, ESEM 2026 Emerging Results track, empirical. Cited here only for scale and composition: §3.2's attributed review streams (4,141,107 review entries, 8,560,237 review-comment entries across 12 reviewer-side agents), §4.1's 248,641 agent-authored PRs with at least one AI-attributed review, and §4.3 + Table 2's same-product majority. It measures comment counts and self-declared category headers, never adoption and never correctness, so it neither replicates nor contests the 71.4% or the 39.3%. Tables reconciled against pdftotext -layout at compile. Full treatment on Closed-Loop AI Review
  • What is happening with code reviews? — Gergely Orosz, The Pragmatic Engineer, 2026-09-08, practitioner-opinion, free portion only. Approach 1: the "noise is a big problem" claim, WeTravel's rejection and June re-evaluation, and Dilocker's adversarial-review/human-scope-decision loop. Testimony only, no counts
§ end
Cited by 20
Related articles
  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Review as the Control Point

    Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…

  • Same-Model Review Blindness

    Greptile's Rodrigo Caridad on two 500-PR labelled datasets (~1,500 verified high-severity bugs): each frontier model ca…

  • Security Debt of Agent-Generated Code

    Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…

  • Agent-Vendor Heterogeneity

    Kraishan (Texas Tech, arXiv 2609.17598): 37,623 provenance-labelled PRs across 2,807 repos with a same-repo human basel…