Sources#
- 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
- Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review
Summary#
Reviewer habituation is the hypothesis that a human who repeatedly reviews AI-agent pull requests gradually lowers her scrutiny and approves more readily. It is the individual-level version of the rubber-stamping worry that runs through Review as the Control Point (its P1/P4 and the vicious loop), AI Brain Fry and Security Debt of Agent-Generated Code, and it borrows its mechanism from the human-factors literature: automation complacency (Parasuraman & Riley; Lee & See) and vigilance decrement (Mackworth).
The only direct measurement in the vault is Haoran Yu, Lifei Liu & Danping Zhang, Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review (arXiv 2609.06213, 2026-09-05). It is a longitudinal cut of the AIDev corpus (public GitHub repos with ≥100 stars; Copilot Autofix, Devin, Codex CLI, Cursor, Claude Code; January–July 2025): 400 repeat reviewers with ≥10 agent-PR reviews each, 11,429 reviews over 207 days, and 10,104 human inline comments. It tests four research questions, and it frames them as a list of what the authors got right and wrong. The approval half of the hypothesis replicates. The language half does not. The predicted time order reverses. And a detector works only once the hand-crafted metrics are replaced with embeddings.
Evidence note.
empirical, observational. It uses one public corpus, there is no preregistration, and the three authors are independent researchers (two unaffiliated, one at Nanchang Hangkong University) with no COI declared. The authors state the decisive limit themselves: "We cannot rule out that AI-agent code genuinely improves over time", and post-merge defect rates, the clean test, "lie outside the AIDev release." So everything below measures approval drift, not lost scrutiny. Whether the extra approvals are wrong is unmeasured. The effect is also small (Cohen's d = 0.25) and confined to open source.
RQ1: approval drifts upward, gradually#
Each reviewer's reviews are split at her temporal midpoint. Approval rises from 30.5% to 36.6% (+6.1 pp; paired Wilcoxon p = 8.6 × 10⁻⁸; d = 0.25). The sign distribution is the more informative number: 52% of reviewers approve more, 28% approve less, and 20% are unchanged. So this is a population tilt, not a universal slide.
The decile series is a different cut and should not be conflated with the headline. It covers only the 173 reviewers with ≥20 reviews (8,481 reviews), pooled into within-reviewer experience deciles, and it runs 27.9% (decile 1) to 42.8% (decile 10), the "+14.9 pp gradient." A line fits it (R²adj = 0.47, about +1.5–1.6 pp per decile). On 50 bins a linear model beats PELT's piecewise-constant one by ΔBIC = +12.0. That is the paper's argument against the "phase-transition" framing of earlier preprints: habituation, if that is what this is, is a slope, not a threshold. The caption's "monotonically rising" is looser than its own numbers. Deciles 2, 4 and 7 each dip below the one before (27.1, 27.8, 32.5).
Two controls are reported. With calendar date as a per-reviewer covariate, the experience coefficient stays positive (+0.11, same direction for 65% of reviewers) while the calendar coefficient is negligible (−0.007). PR size does not remove the effect. The agent-specific effect points the same way where n allows (Copilot n = 209, Devin n = 88). The human-PR control (Fig. 6) is meant to rule out "reviewers becoming generally lenient". The text says human-PR approval shows no directional change, but the figure shows it falling from about 37% to about 11% across 2025-05 to 2025-08, on human-review counts that collapse to near zero in the last two months. The control therefore still argues against global leniency, because leniency would lift both lines, but the figure does not support the "flat" description. Read it as thin in the months that matter.
RQ2: the lexical decline does not replicate#
Four indicators come from the usefulness literature: MTLD lexical diversity, Shannon entropy, technical specificity (identifiers, paths, line references, code spans per word) and constructive actionability (imperative verb, question or code fence, after Bosu et al.). None declines across experience deciles. Spearman ρ values are +0.13, −0.53, +0.12 and −0.04, all p ≥ 0.11, and every Bonferroni-corrected D1-vs-D10 Mann-Whitney p is ≥ 0.36. Within reviewers (339 with ≥6 comments) the paired p ≥ 0.29 for every metric, and the split between reviewers who decrease and reviewers who increase is roughly 50/50 on each one. The result survives length control: the ≥15-word subset (n = 4,094) and the interquartile [6, 22]-word band (n = 5,239) are equally flat. The authors' phrasing is "the decline is absent", not "hidden by length". The practical conclusion: a review dashboard that trends MTLD or technical specificity will not flag habituation. In this corpus those metrics did not move.
The authors give three reasons the metrics fail. The comments are very short (median 11 words; 24% have ≤5 words), and short text saturates lexical statistics. Review style is idiosyncratic: one reviewer attends to types, another to naming, another to tests. And an effect of d = 0.25 may sit below what population lexical statistics can resolve.
Where the drift is: per reviewer, in embedding space#
Population-level embedding probes are also mostly flat. Mean distance to the decile-1 centroid moves only from 0.677 to 0.686 (p = 0.26), and dispersion and PCA drift are null. The exception is the k-means topic mix (k = 8), which shifts across deciles (χ² = 135.3, p = 3.3 × 10⁻⁷). The signal appears within individuals. A reviewer's late-half comment centroid (all-MiniLM-L6-v2) sits at a mean cosine distance of 0.444 from her early-half centroid, against 0.330 under a random permutation of her own comments (paired Wilcoxon p < 10⁻³, n = 342). The authors put it this way: "detectable within-individual, invisible across-individual averaging." Each reviewer drifts in her own direction, and the directions do not add up to a population trend. That is Aggregate Cancellation in a human population, with the page's remedy applied: replace the pooled mean with a per-unit distance against the unit's own null.
The classifier (RQ4) predicts whether a 5-comment window falls in the late half of a reviewer's sequence. It uses 5-fold reviewer-grouped CV over 8,039 windows from 237 reviewers, and 55% of windows are positive. Label position is used deliberately, so that neither approval nor embeddings leak into the label.
| Features / model | Precision | Recall | F1 |
|---|---|---|---|
| Majority class (always "late") | 0.55 | 1.00 | 0.71 |
| Hand-crafted, 10 dims, LR | 0.524 | 0.458 | 0.485 |
| Hand-crafted, 10 dims, MLP | 0.555 | 0.630 | 0.590 |
| Embedding stats, 3 dims, MLP | 0.638 | 0.884 | 0.741 |
| Combined, 13 dims, MLP | 0.633 | 0.707 | 0.667 |
The headline F1 = 0.74 is only +0.03 over the always-late baseline. The real gain is precision (0.55 → 0.64) and balanced accuracy. The text reports balanced accuracy as 0.61 in one sentence and 0.63 in the next. Specificity is only about 0.38, so the model leans toward "late." The hand-crafted features lose to the majority baseline. Adding them to the embedding features makes the MLP worse (0.67), which suggests the hand-crafted features contribute mostly noise. The authors call 0.74 a lower bound, since the encoder is off-the-shelf with no code-aware fine-tuning. Their monitoring proposal is cheap: one 384-d encoder (about 1 GB RAM), unsupervised, and above the permutation null for about 60% of reviewers. But it is a detector of change in a reviewer's comments, not of a missed defect.
RQ3: behavior moves first, language follows#
The paper uses 50 experience bins, first-differenced, with bivariate Granger tests at lags 1–4. Approval-rate change predicts later technical-specificity change at every lag (p = 9.6 × 10⁻⁴, 5.0 × 10⁻⁴, 8.1 × 10⁻⁴, 1.5 × 10⁻³). The reverse direction is significant only at lag 2 (p = 0.036) and does not survive Bonferroni correction across 48 tests. No other metric, the composite, or the embedding distance Granger-causes approval at any lag. Cross-correlation puts the linguistic composite about 4 bins behind approval (|r| = 0.15).
This answers the design question an early-warning monitor depends on, and the answer is unfavorable. Comment language is not a leading indicator of approval drift; it trails it. The method note generalizes: a continuous signal gives PELT finer per-bin variation than a binary one, so running PELT on both can make language look like it leads when Granger says the reverse. Claims of temporal precedence between a continuous and a binary behavioral series should be backed by Granger or cross-correlation, not by comparing breakpoints.
The earlier claim this supersedes#
The vault first met this line of work secondhand. Review as the Control Point's CMU paper cites "Habituation at the gate: Rising approval and declining scrutiny in human review of AI agent code" (H. Yu, L. Liu et al., arXiv 2606.22721) for within-reviewer habituation "where oversight weakens over time as approval rises and commenting falls". The 2606.22721 preprint has not been ingested. The 2609.06213 paper does not cite it by ID. It refers only to "earlier preprints" that claimed a phase transition and a language-leads-behavior order, and it does not explain how the two author lists relate. Read with that caveat, the later paper keeps the approval half, disconfirms a lexical-decline reading of "declining scrutiny", and reverses the time order. It does not test comment volume, so the "commenting falls" clause is neither confirmed nor refuted here. The vault's secondhand summaries are marked accordingly on Review as the Control Point and The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence.
How it sits against the rest of the oversight evidence#
- Against CMU's convergence (Review as the Control Point): the share of agent PRs merged with no review falls toward the human baseline, while individual reviewers approve more. These are different units and different quantities. Review coverage is whether anyone looked. Approval rate is what the looker decided. Both can be true at once. The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence already proposed that reading. This paper moves one of its two terms from "counted" to "counted and tested", but it cannot say whether the added approvals are wrong.
- Against Agent Review Comment Resolution's engaged reviewers: humans in that corpus catch and argue with agent reviewers, which that page reads as evidence against plain inattention. There is no conflict. That page measures humans responding to agent reviewers, and this one measures humans reviewing agent authors, and a d = 0.25 tilt is compatible with engaged reviewers.
- Against Security Debt of Agent-Generated Code: 81.1% of genuine leaked credentials drew no reviewer comment. That is outcome-level evidence in the rubber-stamping direction with no over-time component. This paper has the over-time component and no outcome. Neither closes the gap alone.
- For The Committed-Artifact Chain's point that "a commit log records that a human accepted, never that a human read": this is the first instrument in the vault that reads something about attention from what the reviewer wrote. Per the RQ3 result, it lags the acceptance signal rather than leading it.
- Mechanism: the human-factors cousins are AI Brain Fry (fatigue within a session) and the complacency literature the paper cites. The paper tests exposure, not load, so it does not distinguish habituation from calibration: a reviewer who has learned that a given agent's PRs are usually fine should approve more.
Connections#
- Review as the Control Point — the theory this measures one node of: rubber-stamping in the vicious loop, and P4's surface plausibility. It also carries the secondhand citation of the predecessor preprint superseded above
- Security Debt of Agent-Generated Code — the outcome-level counterpart: missed credentials without an over-time series, where this page has an over-time series without outcomes
- Agent Review Comment Resolution — the opposite direction of the review loop, and the "developers are engaged" finding that a d = 0.25 approval tilt does not contradict
- AI Brain Fry — the within-session fatigue mechanism; habituation is the across-months exposure version, and this paper cannot tell exposure from load
- Risk-Tiered Auto-Approval — the design answer that takes approval out of a habituating human's hands for low-risk classes. If approval drifts upward with exposure, a fixed written gate at least does not drift
- The Committed-Artifact Chain — its "corrector who stops correcting" gap: acceptance is logged, attention is not, and this is the first attention proxy, which lags
- Unknowns as the Agentic Bottleneck — the quiz gate's "comfortable equilibrium" worry, now measured on the human side as a gradual drift that no fixed lexical threshold would catch
- Aggregate Cancellation — the statistical shape of RQ2: roughly 50/50 up/down per reviewer on every lexical metric, a flat pooled mean, and a per-unit distance against the unit's own permutation null as the remedy
- Verification as the New Bottleneck — the hub: human review is the scarce verifier, and this is evidence that its calibration moves with exposure
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — the oversight synthesis; this is the first direct over-time measurement of the "rubber stamp" half
- The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence — where the "direction of oversight is contested" claim lives
Open Questions#
- Is the approval drift habituation or calibration? Rising approval is predicted both by lowered scrutiny and by correct learning that an agent's PRs are usually fine. The paper names post-merge defect rates as the clean test and does not have them. Linking these AIDev reviews to verified follow-up fixes (the Follow-Up Fixes on Agent PRs construct, same corpus family) and asking whether late-half approvals are fixed forward more often than early-half ones would separate the two.
- Does the per-reviewer embedding-drift signal track anything a team cares about? It is validated only against sequence position, a label chosen to avoid leakage. It has not been tested against missed defects, reverted merges, or a reviewer's own report of reviewing less carefully. A monitor that fires on 60% of reviewers is useless if drift and harm are uncorrelated.
- Does the slope persist or flatten beyond 207 days, and does it appear in enterprise review with required approvers? The corpus is seven months of open source, where a reviewer can simply stop reviewing. The linear fit predicts continued drift, and a longer panel or enterprise telemetry would falsify it.
Sources#
- Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review — Haoran Yu, Lifei Liu & Danping Zhang (independent / Nanchang Hangkong University), arXiv 2609.06213, 2026-09-05, 14pp,
empirical. Sources cited: the abstract; §3.1 cohort (400 / 173 / 52 reviewers; 11,429 and 8,481 reviews; 10,104 comments); §4.1 RQ1 (30.5% → 36.6%, 52/28/20 split, ΔBIC); §4.2 and Table 2 RQ2 (Spearman and Mann-Whitney values, within-reviewer 50/50, length-controlled subsets); §4.3 and Fig. 3 embedding drift (0.444 vs 0.330); §4.4 and Fig. 4 Granger (p values by lag, Bonferroni); §4.5 and Table 3 classifier; §4.6 and Fig. 6 controls; §5.2–5.4 monitoring, causal-framing note, limitations. Parse warning: Table 1 (approval by decile) is collapsed in the raw parse into one row whose cells each hold ten values. It was reconciled at compile againstpdftotext -layouton the local PDF, and the order is decile 1 to 10 (27.9 27.1 30.0 27.8 33.9 35.5 32.5 34.2 38.3 42.8), matching the prose and Fig. 1. Tables 2 and 3 parsed cleanly. Numerals in the body are spaced ("30. 5%"). Figure-read, not stated by the paper: Fig. 6's human-PR approval line falls from about 37% to about 11% over 2025-05 to 2025-08 on near-empty monthly counts, against the text's "does not exhibit a directional change"; Fig. 1's fit legend reads +1.51 pp/decile against the text's +1.6. - 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse — Agarwal et al. (CMU, arXiv 2607.07980), §II-B: the secondhand citation of the predecessor preprint arXiv 2606.22721 ("rising approval and declining scrutiny"), superseded in part above.
Cited by 12
- Security Debt of Agent-Generated Code×4
The over-time half now exists, on the same corpus family (2026-10-01). Yu, Liu & Zhang (empirical,…
- The Committed-Artifact Chain×3
That matters more here than in a single-gate design, because the artifacts are chained: spec.md is…
- Agent Review Comment Resolution×2
Reviewer Habituation — the other direction of the review loop, measured over time: humans reviewing…
- Aggregate Cancellation×2
Human reviewers drifting in their own directions (2026-10-01). Reviewer Habituation measures four…
- Review as the Control Point×2
The no-review rate converges, not diverges. The share of merged agent PRs receiving no human review…
- Unknowns as the Agentic Bottleneck×2
The quiz gate is self-administered and self-graded (by the model, on the model's own work). What…
- AI Brain Fry
Reviewer Habituation — the across-months exposure version of the same oversight decay: on 400 AIDev…
- Follow-Up Fixes on Agent PRs
Reviewer Habituation — the same AIDev family's over-time approval drift, which has no outcome…
- AI Coding Practice
Reviewer Habituation — Yu, Liu & Zhang (arXiv 2609.06213, AIDev, empirical, observational): the…
- Open Questions Backlog
Reviewer Habituation ×3 (oldest 4d) — Is the approval drift habituation or calibration?
- Risk-Tiered Auto-Approval
Reviewer Habituation — a reason the fixed gate has value beyond throughput: a human approver's rate…
- The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence
Even the direction of oversight change is contested among studies. CMU's project-level convergence…
Related articles
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- AI as Primary Author
Faros 2026: the assistant→author threshold crossed without a deliberate decision, marked by AI-code acceptance rising 2…
- Review as the Control Point
Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…
- Risk-Tiered Auto-Approval
Gating review by risk tier instead of reviewing everything. PostHog's StampHog auto-approves PRs passing four ordered f…
