H
Howardism
Howardism · Vol. 03Plate II · No. 02

AI Coding Practice, in order.

Notes57DomainAI Coding PracticeOpen Qs119Newest1 Oct 2026Oldest6 May 2026

Workflow, review discipline, and verification practice for AI-native teams.

Map of Content for the ai-coding-practice domain — 45 concepts. How humans and teams practice AI-assisted software work: workflow techniques, SDLC telemetry, review and verification as the bottleneck, and division of labor. Curated entry point; see Home for all domains.

  • Acceleration Whiplash — Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality down (bugs +54%, incidents/PR +243%, review time 5x), gap widening with adoption and hitting even high-maturity orgs; Faros's own Sept-2026 successor reports per-change risk decelerating (incidents/PR +14.5%) while aggregate strain accelerates (monthly incidents +125.4%, QA time +300.6%, PR size +71.8%) — the whiplash relocating from the change to the system
  • Agent-Generated Test Quality — Two AIDev cuts on whether agent code is tested. Jhanglani et al. (204K test files): a trade, not a deficit — agents double human edge-case variety and match assertion strength but carry higher flakiness-candidate rates; three method defects cut into the numbers. Dipongkor et al. (4,882 PRs, ICSME 2026) measure tests against the diff instead: 50.4% of code-changing PRs carry no test change, existing tests execute 61.5% of agents' changed lines in Java and 27.0% in Python (64.8% of Python PRs zero), agent-written tests raise coverage in only 35.9%/22.5% of Code+Tests PRs, and error-handling constructs miss up to 86.0%. Breadth is intrinsic, targeting is relational, and only the second is a safety net — neither study has a human baseline
  • Agent Review Comment Resolution — Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex in 341 Python GitHub repos — first large-scale look at the review loop running the OTHER way: the agent reviews, the human decides. ~71% resolved (Copilot 72.9%, Cursor 67.2%, Codex 54.8%), core developers do most of the resolving (78.1% of Copilot's), an inline suggestion is the strongest predictor (OR 1.62), longer comments fare worse — but AUC 0.58, so most of what decides adoption is not in the comment. Card sorting 470 unresolved-but-argued discussions: modal failure is project context the agent cannot see (23.8%), then confident false positives (63); hallucination 4 of 470; 24.3% acted on but never marked resolved. Contradicts Goldman et al. (ASE 2025) — 4,000 RovoDev comments, 1,007 Atlassian repos, 39.3% — a construct gap, not population (reconciled 2026-09-22)
  • Agent-Vendor Heterogeneity — Kraishan (Texas Tech, arXiv 2609.17598): 37,623 provenance-labelled PRs across 2,807 repos with a same-repo human baseline — the spread between coding-agent vendors is wider than the spread between agents and humans on every outcome measured. Codex reverts at 6.1% (OR 0.50 vs the human 11.5%), Devin at 14.5% (OR 1.31), and the other three are statistically indistinguishable from humans; security-smell presence runs 1.6% (Cursor) to 9.5% (Claude Code) around a 4.6% human rate; hardcoded-credential rates spread fourfold; first-human-review latency spreads from ~1 h to 12.6 h. Pooling agents into one category averages away exactly the differences a maintainer is choosing between — but vendor is confounded with PR size and with task self-selection, and the paper cannot separate them
  • Agentic Coding Work-Composition Shift — Anthropic's 400K-session telemetry, Oct 2025→Apr 2026: as models improved, the share of sessions fixing broken code fell 33%→19% (debugging nearly halved), while operating software (14%→21%) and writing+data-analysis (~10%→~20%) grew; estimated task value rose ~25–27% — usage moving from firefighting toward end-to-end agentic work
  • Agentic Work Systematization — OpenAI Codex study's 'systematization' margin: the shift from ad-hoc agent use (describe task → agent does it → done) to reusable workflow infrastructure via skills and plugins; skill use rose 5.4%→26.6% of weekly-active users (Mar→Jun 2026) and is near-universal at OpenAI (96.2%); custom skills concentrate where shared conventions exist — but the measured post-adoption lifecycle is a one-time copy (53% of reused skills never modified, maintenance 2.7:1 additive), so systematization compounds only under a maintenance discipline most adopters skip — and on the maintained minority (Shen & Hruschka, five public vendor repos) that discipline is a human-governed, AI-assisted loop with no measured transfer benefit yet
  • AI as Primary Author — Faros 2026: the assistant→author threshold crossed without a deliberate decision, marked by AI-code acceptance rising 20%→60% (65% in Faros's Sept-2026 successor); 'not an assistant, the author'; humans move from creation to oversight, making it an authoring problem not a review problem — and the oversight layer keeps automating ahead of the authoring one, agentic review on 50–80% of PRs against agents opening 13–14% even at the leading edge
  • Building Is Cheap, Arguing Is Expensive — "In technical debate, code wins": generate three PRs vs whiteboard; prototype over design doc; reduce design docs
  • Checkpoint-Gated Convergence — Shopify's Helix (September 2026, case-study, nothing measured): an agent rebuilds the 300+-screen React Native app in Swift/Kotlin one small checkpoint at a time, and each checkpoint must clear four ordered gates it may retry but never override — CLI behavior tests generated from the reference app, a Gemini visual-equivalence review with an INVALID verdict, two context-isolated adversarial reviewers checking documented architecture (union of findings, stricter verdict wins), then engineer approval whose feedback enters memory so later checkpoints need less oversight. The design bets on reliable convergence over a correct first attempt, and it builds its migration oracle rather than inheriting one
  • Closed-Loop AI Review — Selvanayagam & Ghaleb (ÉTS Montréal / Trent, arXiv 2608.21311, ESEM 2026 emerging results): the population-scale census of AI reviewing AI on public GitHub — 2,830,284 signature-attributed agent-authored PRs, 248,641 (8.8%) with at least one AI-attributed review, split 45,269 cross-product (1.6%) to 208,145 same-product and growing >2 orders of magnitude across 2025. Composition is sharply non-uniform: Copilot PRs are 95.7% self-reviewed, Cursor's and Google Jules's 0.0%. Reviewer output moves with the pairing — CodeRabbit's refactor share runs 9.7% to 35.0% by author, and three of four dual-role bots emit 58-65% MORE comments on their own product's code, the opposite sign from same-model blindness on a different quantity. Attribution is product-level, not model-level; no defect ground truth, and all counts are lower bounds
  • Code as Source of Truth — Docs go stale at high coding throughput; check specs/skills into the repo; onboard via Claude; spec-drift verification
  • The Code-Quality Payoff Is Token-Indexed — DHH's argument that the 25-year economic case for beautiful, coherent architecture was premised on humans doing the modifications — with agents doing them, the only surviving justification he can name is token scarcity, which makes code craft a moment-indexed economic bet rather than a permanent engineering virtue
  • The Committed-Artifact Chain — Anthropic's Applied AI SDLC playbook makes every stage end by committing a machine-readable artifact the next stage reads — intent.md → spec.md → plan.md → diff+tests → PR findings → incident record — so handoff becomes a merge event and the commit log doubles as the audit trail; the strongest version of markdown-as-interface in the corpus, and entirely prescriptive: every 'how to measure it' is an indicator to collect, not a result
  • Compute Allocator — The human's evolving role: deciding what's worth spending compute on; ~1% of generated tokens ship, 99% is scaffolding invested in alignment/communication; abundance mindset
  • Configurable Human Participation — HAS-Bench (Wu et al.): human participation as a configurable benchmark variable (five-level agency scale × three interaction channels × personas, 397 tasks) — equal partnership (A3) beats full automation by +8.4 Pass@1 and recovers 65% of autonomy-failed tasks, but returns are configuration-dependent and diminish beyond A3; LLM-simulated human caveat.
  • Context Smells — Brian Houck's (DX) vocabulary for recurring agent-context failures, by analogy to code smells: confident hallucination, specification ambiguity, stale guidance, lost in the middle, lost in the details, weekend runaway. The thesis is that AI did not create bad context. It removed the silent human repair loop that used to absorb it. The proposed check is a new-colleague test. Practitioner opinion from a context-measurement vendor, and the vault's measurements back it only in part
  • Design by Selection — Nate Parrott's Claude Design practice: when generating a candidate is nearly free, the designer's labor migrates to the two ends — deciding intent away from the keyboard, then hand-editing the last mile — while the middle becomes 'ask for ten options, then remix the two that work.' Left undirected the model collapses to a recognizable house aesthetic, so explicit aesthetic direction is the load-bearing input, and fidelity itself becomes a control knob (wireframe first when visuals would distract)
  • Design Concept Grilling (hub) — Matt Pocock's grill-me skill; reach Brooks "design concept" before any plan; counter to specs-to-code; PRD as destination doc, Kanban as journey doc
  • Disposable Micro-Apps — Throwaway custom UIs built per-task to edit a plan ("micro-software on top of micro-software"); copy-back-to-markdown; rational under the abundance mindset
  • Efficiency Debt of AI-Generated Code — Tran et al. (Google, arXiv 2608.06640): 3.52M changes over 12 months in one production C++ monorepo, with a human-written control cohort — AI-generated C++ writes ~2x the explicit loops and 30-40% fewer standard-library calls, and that source-level imperative bias shows up in production as ~5% relative compute and ~8% relative memory overhead. The reliability result splits: build failures ~1.3x and sanitizer findings ~1.3x above human, but revert rate ~0.9x BELOW. Review friction is real (blocking threads 1.92x) yet review depth does not predict which inefficiencies survive — the paper's own null, and its argument that this debt has to be caught upstream of review. Taxonomy-informed feedback cuts targeted findings 11.1%, leaving the regenerated functions still net-regressive against the human original
  • Follow-Up Fixes on Agent PRs — Takerngsaksiri, Duong & Barnett (Deakin, arXiv 2609.26847): merged agent PRs draw a verified follow-up fix within 30 days at 1.62x the odds of same-repo human merges (3.68% vs 2.34%), 69.6% of those fixes come from the same agent and 89.1% of fixes to human merges from humans. This is the first controlled post-merge measure in the vault to put agents above human parity, because it counts forward fixes, the channel revert-by-message studies miss. Most of the excess arrives on the day of the merge, and 'self-fix' means same product, not same session
  • HTML as the New Markdown — Thariq Shihipar's thesis: as models improve, thousand-line markdown plans overwhelm the human; HTML artifacts (visual, interactive) keep humans in the loop. The model-facing harness shrinks while this human-facing harness grows
  • Human-Governed Skill Maintenance — Shen & Hruschka (Megagon Labs, arXiv 2609.05677): the first longitudinal measurement of who maintains public SKILL.md files and what maintenance consists of — 254 substantive edits across five purposive AI-tooling repos, every one authored or merged through a named human account, 62% carrying an AI co-author trailer in a sharply bimodal per-repo split (93% / 92% / 16% / 5% / 0%); 60% enhancement vs 38% correction, deprecation 1 of 254, skills stable-to-additive at a 5-day median cadence; and a pre-registered, powered transfer-task test finding the maintained version no better than the earliest (−0.09 on a 1–5 scale, CI [−0.28, +0.10])
  • Impose Values, Not Disciplines — Robert C. Martin's distinction between the outcomes a practice is meant to produce and the human-ergonomic ritual that produces them: a lifelong TDD advocate refuses to make agents do TDD because red-green-refactor is an adaptation to human working memory, not a property of good code — agents get the same values (coverage, bounded complexity, tested branches) with thresholds moved (CRAP under 4 for humans, 6 for agents, maybe 8) and are left to reach them their own way, which they do by reverting to write-function-then-test no matter what the prompt says
  • Layered Supervision — Stolze & Strässle (ESEM 2026 SEIP, 5 interviews + a 50-person indicative survey): as generation outruns review, supervision stops being one control point and distributes across three layers — preventive guardrails (architectural intent externalized into steering files and specs, which nothing checks), executable guardrails (lint/test/CI promoted from quality tooling to the build-as-arbiter), and human oversight re-scoped from line-by-line reading to concurrent supervision plus 'operational explainability'. The layers are distinguished by mechanism, not sequence: preventive shapes ex ante only if the generator consults it, executable verifies ex post regardless. No layer suffices alone, and the paper measures none of them
  • Living Design System — design_system.html extracted from repos as a portable, human- and machine-readable source of truth; component playgrounds; bridges engineering ↔ non-technical stakeholders
  • LLM-Assisted Grey-Literature Theory Building — Agarwal et al.'s secondary contribution (arXiv 2607.07980): a scalable template for constructing grounded theory from thousands of practitioner documents instead of a few dozen interviews — LLMs do the mechanical, quote-anchored open coding (38,709 docs collected → Gemini relevance judge at κ=0.75 → 3,100 coded with the multi-agent Thematic-LM under three deliberately-polarized coder lenses → 4,838 codes / 109,951 quotes at ~$0.35/doc) while humans keep the interpretive axial/selective coding; automating that back half FAILED (a bottom-up pass yielded 15,029 shallow, redundant causal statements), so the codes→theory step stayed a manual, LLM-as-search-engine process — a division-of-labor lesson about what LLMs can and can't do in qualitative research
  • Open Source Under Agent Contributions — When contribution supply goes free and unbounded, the maintainer's scarce resource stops being contributors and becomes attention: DHH reports 1,000+ merged PRs in three months on Omarchy with the backlog doubling weekly, agents doing first-pass triage, and rejection turning socially cheap because no human wrote the patch — against the maintainer-burnout reading of the same influx
  • Outsource Your Thinking, Not Your Understanding — "You can outsource your thinking but not your understanding"; understanding as the non-delegable human bottleneck; knowledge bases as understanding-tools
  • Planning / Execution Division of Labor — Anthropic's 400K-session telemetry: in a typical Claude Code session humans make ~70% of planning decisions (what to do) while Claude makes ~80% of execution decisions (how to do it); each prompt sets off ~10 actions (8 when the user keeps execution control, ~16 when Claude controls planning) — 'people decide what to build, the agent decides how'
  • Post-Acceptance Edit Behavior — Liang et al. (CMU, arXiv 2607.25130): DECODE, 53.6K in-IDE edits of accepted AI completions from 1,141 developers — the first measurement of what humans DO to AI code after accepting it, pre-commit rather than at PR level. Half of all edits land within 50 minutes and the volume collapses after 15; retention is bimodal (median 63% survives, but the mass sits at 0% and 100%); 31% of trajectories contain a removal edit, and the developer who first tries to CUSTOMIZE a completion is the likeliest to delete it next (23.4% vs 12.2% after a functionality change). Which of 20 models wrote it barely matters (eta-squared 0.002–0.007). The prediction half is weaker than its abstract: fine-tuned 3B models beat their own base by +0.23 F1 but the best frontier baseline by only +0.08, and the dominant edit type — changing functionality — tops out near 0.49 Levenshtein similarity for every model tried
  • Review as the Control Point — Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded practitioner documents — review is the control point through which a coding agent's effect on software is decided, and AI does NOT fix the sign of that effect; the team sets it through reviewer expertise, disposition, and how it adapts the review process (three moderators). Central core is review depth + reviewer skill, threaded by comprehension-debt feedback loops; the paper's own non-vendor GitHub telemetry (2.5M+ PRs) finds agent PRs reviewed less / merged several× faster / discussed less, but the trends flip direction under defensible analysis choices and the no-review rate CONVERGES toward the human baseline over time
  • Reviewer Habituation on Agent Pull Requests — Yu, Liu & Zhang (arXiv 2609.06213, AIDev, empirical, observational): the same human reviewer approves agent PRs more often the more of them she has reviewed — 30.5% to 36.6% early-vs-late half across 400 repeat reviewers (Wilcoxon p = 8.6e-8, d = 0.25), a gradual linear rise rather than a phase shift — while four canonical comment-quality metrics (MTLD, entropy, technical specificity, actionability) stay flat. The drift is visible only per reviewer, in sentence-embedding space, and Granger analysis puts approval change before the language change, reversing the 'language degrades first' claim of the authors' own earlier preprint. No defect data: a rising approval rate is equally consistent with agent code getting better
  • Reviving Impractical Quality Tools — Robert C. Martin's mechanism for why agents change code quality: CRAP score and mutation testing were sound ideas around 2000 that he abandoned because a human had to pay for their output — the tools did not change, the labor did, and an agent that does not care how boring the work is turns an overnight run plus weeks of remediation into a 30-minute loop; the general form is that any quality technique whose cost sat in remediation rather than detection is now worth re-auditing
  • Risk-Tiered Auto-Approval — Gating review by risk tier instead of reviewing everything. PostHog's StampHog auto-approves PRs passing four ordered fail-closed checks (state, blast-radius deny-list, diff ceiling, LLM veto) and stamps ~1 in 3 merged PRs, with no defect rate reported. Other tiering keys: reversibility of the response (security ops: 13% read-only to 0% full autonomy, self-reported) and consequence of the output for non-code workflows. Also why checkpoints only accumulate, a three-question audit to remove them, and Duckbill's before/after (merged PRs +94%, no defect count)
  • Same-Model Review Blindness — Greptile's Rodrigo Caridad on two 500-PR labelled datasets (~1,500 verified high-severity bugs): each frontier model catches fewer bugs in code authored by its own model family than in the other family's code — Opus 4.7 53.7% same-model vs 60.0% cross-model, GPT 5.5 50.5% vs 62.0%. The crossover is a pure interaction (both reviewers average ~56% overall and the two datasets differ by 2.6pp), but the post's offered mechanism — that a model misses the bug categories it produces most — reproduces only ~7% of its own headline when the category table is reweighted by the bug mix, so the blindness operates WITHIN category, not through composition. Vendor-built ground truth with an unspecified labelling procedure, one arm's prompt tuned against the outcome metric, no released artifact — filed case-study, not empirical
  • Security Debt of Agent-Generated Code — Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4,022 AIDev agentic PRs — 38.9% of agent PRs carry ≥1 security smell, supply-chain integrity (mutable action/image tags, unpinned installs) is 82.3% of them, GitHub Actions + Dockerfiles hold 87.6%, hard-coded credentials are 99.6% of critical smells, and flagging climbs with PR size from 16.2% to 53.6%. The two RQ2 surprises invert the usual story: humans, not agents, committed 67.6% of the 74 genuine leaked credentials, and 81.1% of them reached integration with no comment from any bot or human reviewer. There is no human-PR control group, so it measures the security posture of agent-assisted workflows, not an agent-vs-human delta
  • Spec-Driven Development as the New Waterfall — Robert C. Martin reads the 2026 spec-driven-development movement as the 1970s big-design-up-front temptation returning under new branding, and reports his own attempts failing the same way: the plan cannot anticipate what the agents hit, so the human stops them, rewrites, restarts — his answer is the agile one (a story or two, then look at the architecture), justified by the cost of change falling near zero, with specs kept ephemeral and never committed because the finished artifact is the specification
  • Telemetry vs. Survey Measurement — Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (instrument agreement tracks shared data source, not construct), randomization as the only causal instrument, the survey arm's counter-case (shadow AI is invisible to telemetry), and the Ramp payment-rail aperture; first cross-family convergence: Anthropic passing OpenAI mid-2026; and the vendor-survey rule stated twice — third-party fielding and a stated n do not move the tier when the conclusion is the product and the method is gated.
  • The Three Loops of AI-Native Building — Andrew Ng's nested-loop taxonomy for 0-to-1 products: the agentic coding loop (minutes, agent-closed), the developer feedback loop (tens of minutes to hours, human-closed), and the external feedback loop (hours to weeks, market-closed); loop engineering has been optimizing only the innermost one, and the human's remaining job is a context transfer that lives in the outer two
  • Unknowns as the Agentic Bottleneck — Thariq Shihipar's map-vs-territory thesis: the gap between what you told the agent and what the work actually requires is unknowns, and with Fable-class models the human's ability to surface them — not the model's capability — sets output quality; the Rumsfeld 2×2 applied to prompting, plus a phase-ordered catalog of elicitation techniques
  • The Verifiability Thesis (hub) — LLMs automate what you can verify as computers automate what you can specify; RL verification rewards → jagged peaks; "verifiable + labs care"; everything eventually verifiable
  • Verification as the New Bottleneck (hub) — Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax; PR-cycle-time funnel analysis
  • Vertical Slice Tracer Bullets — Pragmatic-Programmer tracer-bullet pattern applied to agent task decomposition; vertical slices > horizontal layers; Kanban-with-blocking-edges over numbered phase plans
  • Vibe Coding vs. Agentic Engineering — Vibe coding raises the floor (anyone builds); agentic engineering preserves the quality bar while going faster; ">10x and widening"; hire on big projects, not puzzles

Derived#

  • What the Agent-PR Oversight Numbers Can and Cannot Say — Six-question synthesis of the agent-PR oversight cluster. (1) 'Acceptance' of an agent-applied change is non-reversion over a horizon, a survival measure that carries oversight meaning only when stratified by who examined the diff, now four classes (nobody, the invoker, an independent human, an AI reviewer); retired. (2) The 58-65% same-product comment gap is already reviewer-fixed by construction, but with the reviewer fixed the author product varies (86% of Copilot's cross-product PRs are Codex-authored), and nobody has matched on size. (3) Codex and Devin revert at OR 0.50 vs 1.31 on near-identical median PR sizes (63 vs 61 lines), so the page's headline vendor gap is not a size artifact at the median, while Claude Code's anomalies are the size-sensitive ones. (4) 'How far to automate review' gains a 'by whom' axis: the ecosystem default is self-review 4.6:1, the configuration the only recall data disfavours. (5) Ng's QA relief credits agents testing their own code, which the diff-coverage data says mostly does not happen in open-source PRs. Every residual asks for a number the wiki lacks
  • Opinions on Using AI Tools & the Future of the Software Engineering Role — Debate map of four stances on using AI tools (bullish-insider / pragmatist-practitioner / skeptic-governance / architecture-thesis) + synthesis on the future SWE role: coding→deciding/verifying, role convergence, what stays human, which moats survive, honest caveats
  • The HTML Artifact Lifecycle: Where Plan History Lives, and When Disposable Becomes Durable — Two-question synthesis on the lifecycle of human-facing HTML artifacts. (1) The diff/version problem dissolves once the artifact is recognized as a compiled view, not a record: version the content layer (markdown/config/repo — the copy-back round-trip and the extract-from-code pattern already do this), regenerate the presentation on demand, and let review reattach to decisions rather than diffs (the plan-ordered-by-likelihood-of-change technique puts the reviewable delta at the top). What's genuinely lost — blame history of the presentation itself — is acceptable precisely because presentation is regenerable at abundance prices; the moment a presentation choice is load-bearing, it graduates to durable tooling. (2) The templating question resolves by naming the correct reuse unit: not the artifact but the generator — a recurring micro-app becomes a skill that regenerates a fresh, fitted app each time (keeping disposable's per-task fit while gaining reuse's consistency), which is exactly the systematization move measured in the wild (skills 5.4%→26.6% of weekly-active users). An artifact itself graduates from disposable to durable only under recurrence + sync-pressure + audience (the design_system.html profile), at which point it stops being free: it inherits maintenance cost, sync cadence, and a seat under the artifact-sprawl bloat ceiling. The failure mode is the un-chosen middle — ad-hoc apps kept around unmaintained, which is sprawl plus rot with neither fit nor consistency
  • Human-in-the-Loop Boundaries — Humans belong at allocation, understanding, design-concept, risk, and accountability boundaries; they slow the system down as manual executors, universal reviewers, or ceremonial approvers
  • Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — Five-question synthesis of the oversight cluster. (1) 'Acceptance' at 60% blends three acts of different evidentiary weight — affirmative adoption, reviewed non-reversion, and bare non-reversion — and only the third is the rubber-stamp class; the construct, not the data, is why Faros and CMU can headline opposite trends. (2) The allocator-rubber-stamp risk is documented at three evidence layers (brain-fry error rates +11/+39%, 31.3% no-review telemetry, P2–P4 practitioner discourse): volume plus surface plausibility push the human past engagement, so the framing survives only with structural countermeasures (quiz gate, sample-based depth, risk-tiered gating) that make understanding rather than signature the merge condition. (3) 'How far to automate review' is a partition, not a dial: automate mechanical verification fully, keep human depth on a sampled/high-stakes slice — because the binding constraint isn't defect-catching (contested P9) but ownership, skill growth, and comprehension debt, which accrue regardless of who catches bugs. (4) Faros-vs-DORA is partly a category error — surveys measure felt productivity, telemetry measures system outcomes, both true at their layer — but the maturity-protection disagreement is substantive and unresolved. (5) Ng-vs-Faros is both-and: the 0-to-1/production scope split is real and does most of the work, while documented optimism bias means Ng's self-reported QA relief can't be read as measurement
  • Learning to Co-Work with AI: A Software Engineer's Field Guide — Field guide for software engineers in the AI era: 6 skill clusters (taste, harness, alignment-first planning, agent-friendly architecture, verification, strategic positioning), daily practices, anti-patterns, 90-day plan
  • Is Persistence the Line Between Prompting and Spec-Driven Development? — No — persistence is an observable proxy for the property that actually matters, which is authority: whether later work is judged against the artifact. The corpus breaks the persistence test in both directions (CLAUDE.md, WORKFLOW.md, REVIEW.md, tickets and Martin's own dependency-rule file are persistent and nobody calls them specs; the grilling transcript, Carey's why-conversation and Martin's deleted plans are ephemeral and do the whole specification job), and both cited positions turn out to be misread — Pocock's sentence contains two clauses, persist and return to, and only the second survives; Martin does keep a specification file in the repo, scoped to constraints a checker enforces. Authority classifies every corpus case correctly; persistence is what makes authority cheap. Cluster is almost entirely practitioner-opinion, and nothing measures persistence-per-se
  • Rationale as a Dated Record: Where the Why Lives for the Next Reader — Three-question synthesis on the orphaned why. (1) Rationale survives only where it is committed as a dated decision record beside the work, and in that form it does not share code-as-source-of-truth's staleness problem: staleness is a read-back-as-current harm, and a record of why B beat A on a given date stays true as history, provided its authority is revoked when superseded (Pocock's closed issue). The corpus now has this in practice: OpenAI's Codex team checks execution plans with decision logs into the repo, and Anthropic's playbook commits intent.md carrying the why. (2) A post-hoc decision record does not reintroduce the PRD, because what made the PRD a PRD was authority: later work was judged against it. A record written after the choice holds none (intent.md, which binds the spec generated from it, does not qualify). (3) Partial: the why can live in the repo after all, so Fung's carve-out shrinks to records whose authority sits outside it (regulator-accepted systems, other teams' dependencies) and to tacit knowledge. The mechanisms for keeping that slice current (CI regeneration, doc-gardening agents, same-session write-back) are prescribed but unmeasured, and org strategy is covered by no source. Evidence is vendor-claim and practitioner-opinion throughout
  • The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence — Resolves acceleration-whiplash's open question: the Faros-vs-CMU under-review 'divergence' is mostly measurement artifact — a vendor's adoption-depth delta in unreviewed-PR count over enterprise all-PRs vs a non-vendor calendar-time share of unreviewed agent PRs in open source — and both fit one story: total unreviewed output rises with volume while the share of agent PRs merged unchecked falls as teams learn risk-triage; the volume-concentration clause is supported (median per-project no-review ≈0%, pooled >50%; triage by PR type), and the residual disagreement is a forecast — whether triage discipline survives agentic authoring crossing from <1% to double digits
  • When Does Verification Quality Determine Whether AI Automation Works? — Verification-quality ladder from Lean/formal proof search through software CI and vulnerability reproduction, plus the rung below it where no executable verifier exists at all and structured output plus intra-repository peer comparison stand in; autonomy should rise only to the level the verifier can support
  • Verifying Without a Compiler: Cowork's Harness vs Claude Code's, and Why the Slice Verifier Stays — Two-question synthesis on verification where no mechanical checker exists. (1) Cowork and Claude Code share primitives (skills, MCP, sub-agents, computer use) but sit on opposite ends of the verifier ladder, so the harness weight redistributes: Claude Code leans on a deterministic post-hoc verifier stack (tests, compiler, diffs, spec-drift checks) that both catches errors and bounds damage before merge; Cowork's outputs have no such rung, so its harness substitutes judgment-encodings for mechanical checks — the loaded design system as the nearest thing to a style linter, evals and LLM-judges for quality, human review concentrated at decision checkpoints — while the pre-action classifier gate becomes load-bearing because errors ship directly into live SaaS state with no red test in between. Failure modes split accordingly: loud (build breaks) vs silent (a polished deck that reads fine — the failures-that-look-like-success class), which is why accountability redesign matters more for Cowork, not less. (2) The planner needs the horizontal-slice verifier by design, not just empirically through 4.7: 'every slice produces end-to-end feedback' is a mechanically checkable invariant (does it touch schema+service+UI?), and checkable invariants belong in the deterministic checker regardless of model trust — the verifier is a constraint (doesn't compound, costs nothing to keep, catches the training-prior regression toward horizontal layering), while 'please slice vertically' is a behavior request, the transient form of the same discipline. Trust-the-model applies to prompt lines; verifiers are the durable class
  • Writer/Reviewer vs Agent-to-Agent Review — The two patterns are one architecture differing on venue, so the corpus's evidence attaches to four decomposed axes rather than to either brand. Reviewer lineage is the only axis with a code-review number (Greptile, case-study: cross-model 61.0% vs same-model 52.1% high-severity recall, an 8.9pp crossover surviving both main effects and NOT explained by bug composition) — which falsifies the Writer/Reviewer rationale as written, since the own-code bias is a property of the model family and clearing context does not clear it. The information-boundary axis has the better design argument (Bun's diff-only reviewer, OpenCodeReview's falsification-only reflector, Cursor's swept-but-unreported lens sweep) and its only number comes from an adjacent task (Anthropic monitor recall 92% → 48% under long benign context). Gate position has one block rate (Ouroboros 63.5%) and no outcome measurement. The outcome axis is fractured across two incommensurable constructs: 7.23–37.80% reference-match precision on AACR-Bench against 71.4% developer resolution of 54,713 real agent review comments. Codex's shipped /review is the high-precision/very-low-recall point (266 comments over 200 PRs, 4.92% recall) and Claude Code's pinned /code-review the opposite (5,980 comments, 28.90% recall, 7.23% precision) — a difference in comment-volume policy, not in pattern. Nothing measures the two patterns head to head

Open questions 119 open

    • SourceFaros's own deferred question: do the bug/incident increases persist when normalized for PR size, or do larger PRs account for most of the quality deterioration? (If the latter, hard PR-size limits are the highest-leverage fix.) Partially answered by Security Debt of Agent-Generated Code (empirical, non-vendor): on the security axis, PR-level flagging rises monotonically with change size — 16.2% for 1–9-line PRs to 53.6% for 1000+-line PRs, a 37.4-point spread — which is the size-stratified evidence this question asks for and supports hard PR-size limits as a real lever. Two gaps keep it open: it measures smells introduced, not the bugs and incidents Faros counts, and being a cross-sectional association it can't say whether capping size lowers density or merely re-partitions the same changes across more PRs. Further partially answered 2026-08-12 by Tran et al. (empirical, with a human control cohort): AI changes there are indeed larger (median 89 lines vs 33, 3 files vs 2), and the downstream comparisons are stratified on change size among other covariates — so the ratios that survive stratification are not the size effect. What survives is split by outcome: blocking threads 1.92x and build failures ~1.3x stay above parity, revert rate ~0.9x stays below. So size does not account for the deterioration, and the deterioration does not have one sign. A hard PR-size limit therefore addresses review burden rather than production stability, which is a narrower case for the lever than this bullet originally assumed. A third partial, 2026-09-22, on a public-GitHub population: not all agents are equal five coding agents post merge (Kraishan, empirical, non-vendor, same-repo human control) runs the size question two ways and both say size is not the story. Its size-stratified sensitivity check finds the pooled agent-human security-density difference concentrated entirely in the XL PR bucket (δ = −.08, p = .025) with no difference in the four smaller buckets — which is the opposite reading of the same shape the security study gives: there, larger PRs carry more smells; here, larger PRs are where agents beat humans. And its post-merge churn measure is size-normalized by construction (follow-up commits ÷ changed lines), with four of five vendors at or below the human level after normalization; the unnormalized ranking is different (Copilot lowest of any group, Codex slightly above humans), so normalization changes who wins but does not create an agent penalty at either setting. What it still cannot supply is the quantity this bullet actually wants: it counts reverts and churn, not the bugs and incidents Faros counts, and its revert detector works by commit message and misses silent rewrites. The lever is now supported by two sources on review burden and by none on production stability. Faros's own successor was asked this question and did not answer it (2026-09-22): The Speed Trap re-runs the same panel six months on and still relates incidents per PR to PR size nowhere in the post — no size-stratified incident figure, no normalization statement, nothing. What it does supply is a suggestive non-answer: across the two windows the per-change incident growth collapsed from +242.7% to +14.5% while PR-size growth accelerated from +51.3% to +71.8%, so the two quantities this bullet asks to be related moved in opposite directions. If size were driving the incident rise they should track. That is evidence against the strong form of the size hypothesis on Faros's own instrument, and it is emphatically not a normalization: two uncontrolled period-over-period deltas on an already-high-adoption panel, no joint model, no control cohort, and the methodology is in a form-gated report. The trigger is unchanged — incidents per PR cut by PR-size bucket, from anyone.
    • SourceCode churn +861% is genuinely ambiguous (Faros lists three explanations: rework of AI code, productive legacy refactoring, or accelerated polish). The cross-customer metric can't resolve it — a real gap, not a finding. A partial proxy added 2026-07-29 by circleci q2 pulse 2026 (vendor-claim): CircleCI's Merge Efficiency Ratio — validation cycles a feature branch needs before it lands on main (median 3.9, top-5% 2.6, elite cohort 1.3) — counts a pre-merge form of the same rework, and it is countable per team rather than pooled cross-customer. It narrows the ambiguity from one side only: cycles spent failing validation before merge are hard to read as "productive legacy refactoring," so a high MER is closer to unambiguous rework than churn is. It does not decompose Faros's metric, because the two measure different things — MER counts attempts, churn counts lines-deleted-to-added, and a clean refactor that passes CI first try is invisible to MER while dominating churn. A second partial from the other side of merge, 2026-09-22: not all agents are equal five coding agents post merge measures post-merge churn with a same-repo human control — follow-up commits on the PR's most-changed file within 90 days, divided by changed lines. Four of the five agents sit at or below the human level (Claude Code δ = −.33, median 0.5 against the human 5.3; Copilot −.16; Cursor −.08; Codex and Devin at parity). That does not decompose +861% either, and the constructs are again different — commits-per-changed-line over 90 days against lines-deleted-to-added within a short window, and this one is measured on one file per PR as a cost-driven proxy for whole-PR churn. What it does is eliminate the first of Faros's three explanations in this population: if the churn were rework of AI code, agent-authored changes should draw more subsequent commits per line than human-authored ones in the same repositories, and they draw fewer. The remaining two explanations (productive legacy refactoring, accelerated polish) are both about work that is not attributable to a specific PR, which is precisely what a per-PR design cannot see. The metric was re-measured by its own publisher and it does not discriminate either (2026-09-22): The Speed Trap reports the same quantity — now named "code deletion ratio" rather than "code churn" — growing +71.6% against the prior window's +861%. A growth rate that falls by more than an order of magnitude is consistent with all three of Faros's explanations (transition rework subsiding as teams learn the tools; a legacy-refactoring backlog being worked off; polish normalizing), so it discriminates among none of them, and the report offers no decomposition. Two things it does settle, both narrow. The +861% was a transient of the adoption transition, not a new level — the runaway-and-compounding reading of this bullet is out. And Faros uses two different names for what is presumably one metric across two reports with no definition published for either, which is an independent reason the construct cannot be decomposed from the outside. What would settle it is unchanged: a per-team decomposition of deleted lines by whether the deleted code was itself recently AI-authored. One adopter rebuilt the construct, 2026-10-01: ai changed how spotify builds quality at higher velocity (Spotify, case-study) replaced churn with an age-weighted rework rate that separates genuine rework from new work and legacy refactoring. It reports churn rising industry-wide (citing this report) with no corresponding rise in its own rework rate. Taken at face value, that rules out the first explanation (rework of AI code) at Spotify and leaves the other two, which is the closest any source has come to the decomposition above. Two limits keep it a partial. Nothing quantitative is published: no rework figure, baseline or age threshold. And the metric weights the age of the changed code, not its authorship, so it detects rework of recent code, not of AI-authored code. That is a proxy for the decomposition, not the decomposition itself.
    • How much of the "maturity doesn't protect" claim survives the vendor incentive to argue exactly that (i.e., "your existing practices won't save you — you need our platform")? Partially answered by Review as the Control Point (non-vendor, empirical): its whole thesis is the opposite — AI doesn't fix the sign; team expertise and process do — which leans toward DORA's "foundations protect you" and against Faros's determinism. But it argues the moderators exist rather than measuring a maturity effect, so the vendor-incentive question isn't closed, only counterweighted by a non-vendor source that disagrees with the framing. A third position, 2026-09-22 — layered supervision ai assisted software engineering (case-study, 5 interviews, non-vendor, academic) refuses the axis rather than taking a side on it. Its five teams varied along three named dimensions — governance posture, system criticality and homogeneity, and team composition — and the authors decline the word maturity explicitly, "since it suggests a linear progression that our data does not support." That is a direct challenge to the shared premise under both this claim and DORA's, which is that organizations can be ordered on one scale at all; if the space is genuinely multi-dimensional, "maturity doesn't protect you" and "foundations protect you" are both answers to a malformed question. It does not close the vendor-incentive question and cannot: refusing an ordering on five qualitative points is a modelling choice, not a measurement, and the survey behind those dimensions was convenience-sampled and unpiloted. What it does is put a non-vendor source on record that the ladder is the thing in doubt. A weak datum from the incentive's own side, 2026-09-22: in The Speed Trap Faros reports its own headline alarm metric — incidents per PR — decelerating from +242.7% to +14.5%, and states plainly that per-change quality "is no longer collapsing." A vendor selling on "your practices won't save you" publishing a seventeen-fold softening of its scariest number is the direction that costs the seller something, which is mild evidence against reading the maturity claim as pure narrative construction. It is mild: the alarm did not go away, it moved to metrics the same product addresses (QA time +300.6%, monthly incidents +125.4%, unreviewed merges), and the successor is the same instrument with the same undisclosed method, so it cannot audit the prior report's framing from outside. The question still needs a non-vendor measurement of a maturity effect, which nothing in the corpus supplies. A mature adopter's own account, 2026-10-01: ai changed how spotify builds quality at higher velocity (Spotify, case-study, self-graded) splits the claim in two. Its foundations did not stop the volume of change from outrunning review, testing, rollout and observability, which is the maturity claim in its system-level form. But its incident retros found no AI-authored code as a material direct contributor, which is the per-change form the claim was originally argued in. It is non-vendor with respect to Faros but not disinterested, since the company is grading its own AI rollout, so it adds a data point without closing the incentive question.
    • ResolvedFaros reads under-review as a widening crisis; CMU's non-vendor GitHub telemetry finds the agent no-review rate converging toward the human baseline (>50%→~14%) as orgs learn to review agent code. Is the divergence real (enterprise vs open-source populations, adoption-depth cross-section vs calendar-time trend) or does the whiplash's under-review pressure only surface where PR volume is highest? Answered: The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence — mostly not real: Faros's +31.3% is a delta in unreviewed-PR count across adoption depth (enterprise, all PRs) while CMU's is a falling share of unreviewed agent PRs over calendar time (open source) — a falling rate and a rising count coexist under Faros's own volume growth. The volume clause is supported (median per-project no-review ≈0% vs pooled >50%; triage by PR type), and Faros's own risk-tiered-gating remediation is the triage behavior CMU observes emerging. The residual disagreement is a forecast: does triage discipline survive agentic authoring crossing from <1% to double digits — untested in both datasets.
    • SourceThe 470-discussion taxonomy is drawn only from unresolved comments that received a reply — 6.7% of the unresolved population. What explains the silent 93.3%? Reviewer fatigue, comment volume, triage, or the same context error just not worth arguing about? The taxonomy's shape (only 11 of 470 dismissed as low-value) would look very different if noise dominated the silent majority, and the two hypotheses are distinguishable by sampling the silent set directly. Context from goldman code review comments developers resolve ase 2025 (2026-09-22), which is not an answer: its 4,000-comment pool applies no reply filter, so its 60.7% non-resolution covers the whole unresolved population including the silent majority — establishing that the silent majority is an industrial phenomenon and not a GitHub-etiquette artifact. But its only account of why is an unmeasured discussion-section conjecture ("the lack of clarity, relevancy, and simplicity of the comments"), with no card sort, no reply analysis and no sampling of the silent set. The question survives intact.
    • SourceResolution is adoption, not correctness. Does a resolved agent comment correspond to a defect that would otherwise have shipped — and does an agent reviewer change any downstream outcome (escaped defects, incidents, revert rate) against a no-agent-reviewer baseline? No study in the corpus has run this on either side of the review loop; it is the same missing outcome measurement that leaves Risk-Tiered Auto-Approval's throughput figures unattached to safety. Unchanged, and now doubly so (2026-09-22): goldman code review comments developers resolve ase 2025 measures the same loop at industrial scale with a construct that is further from correctness, not closer — an exact-line modification in a later commit scores as resolution whether or not the comment caused it, so unrelated churn on a hot line counts as agreement — and it measures no downstream outcome either. Two industrial-scale studies, two incompatible constructs, zero correctness ground truth on either.
    • ResolvedTwo empirical studies a year apart disagree by roughly a factor of two on the central quantity: Goldman et al. (ASE 2025) report 60-70% of LLM-generated review comments unresolved, this study reports 71.4% resolved. Is the gap population (industrial vs open-source GitHub), product (in-house pipeline vs shipped agents with one-click suggestion blocks), or construct (whatever Goldman counted vs a GitHub thread flag that this paper's own card sort shows under-counts by ~24%)? Answered 2026-09-22 by goldman code review comments developers resolve ase 2025 — the gap is construct. On a common scale it is 39.3% (Goldman, 1,571/4,000, recovered from his Figure 5 since the paper never states an overall rate) against 71.4% here: 32 points, not a factor of two. Population is ruled out by direction — Goldman's 1,007 internal repos are staffed entirely by core-equivalent paid developers, and this study's own RQ2 says core developers resolve more, so an all-core population predicts the higher number and Goldman has the lower one. Product is bounded at roughly 11 of the 32 points — the inline-suggestion effect measured here is 75.5% vs 64.6% at Cramér's V = 0.12, and RovoDev ships no applicable diff. The residual is the construct, and it is a difference in kind: "a subsequent commit modified the exact line where the comment was placed" is a code-diff-at-an-anchor event that misses every off-line fix, every deleted anchor and everything right-censored by an unstated follow-up horizon; GitHub's isResolved is a disposition flag that admits dismissals and misses unmarked acceptances. Confirmed from inside Goldman's own data: his category ordering (readability 43.3% > bugs 41.9% > maintainability 36.2% > design 28.6%) is the ordering of how line-local a fix is, and he explains the design deficit by saying design needs "broader changes" — exactly what an exact-line proxy cannot see. Caveat carried forward: this is a reading of two constructs, not a re-measurement of one construct on two populations, so the split between the product and construct shares is argued rather than estimated. The durable conclusion is that the two figures bracket an adoption rate rather than estimating one, and neither may be quoted as "the" resolution rate. Full working in Reconciled: this paper against the one it cites above.
    • SourceThe paper's own two-stage protocol was never completed: does the 0.44 vs 0.30 candidate-rate gap survive dynamic confirmation, or do agent tests contain flakiness indicators without being measurably flakier under repeated runs? The specified experiment (1,000 sampled tests × 100 runs per cohort) would settle it directly.
    • SourceDoes the edge-case-breadth advantage survive data-flow analysis? The literal-only detector may be measuring "agents pass literals where humans pass fixtures" rather than a real coverage gap — a re-run with variable resolution, or a matched pytest-aware parser, is the discriminator.
    • SourceWhat is the survival rate of agent-authored tests? The paper's own future work names the missing quantity: how often agent tests are deleted, rewritten, or @skip-marked over subsequent months. Coverage breadth bought at the cost of a suite people learn to ignore is negative value, and nothing here measures the maintenance side. Not answered, but the first deletion figure lands nearby (2026-08-12): Dipongkor et al. find that in non-improving Java Code+Tests PRs agents delete more tests than they add — 82 deleted against 31 added, 2.6×, with a further 51.2% editing only existing test bodies. That is the opposite direction of this bullet (agents removing pre-existing tests within a single PR, not agent tests decaying over months) and it comes from 64 Java PRs, but it is the corpus's first measurement of agentic test deletion in any form, and it makes the longitudinal version cheaper to ask: the same repositories already carry the history.
    • SourceNeither AIDev study has a human baseline, and the missing comparison is now the same one twice. Agents include a test change in 49.6% of code-touching PRs and their tests raise diff coverage in 22.5–35.9% of Code+Tests PRs — but nothing establishes whether human-authored PRs in the same repositories do better. Diff coverage is computable retroactively from any merged PR, so a matched human cohort over the same 44 instrumented repos is a tractable study rather than a wish, and it would settle simultaneously whether the 0.62-vs-0.32 edge-case gap here survives a targeting-aware metric. Partially answered on the method half only, 2026-09-22. not all agents are equal five coding agents post merge (Kraishan, arXiv 2609.17598, empirical) is the first AIDev study to actually build the matched human cohort, and it publishes the recipe: keep human PRs only from the 810 repositories that also contain agent PRs, restrict them to the agents' own activity window (2024-12-24 to 2025-07-30), and cap the count per repository under a fixed random seed so no project dominates — yielding 4,027 human PRs beside 9,750 agent PRs in the same repos, with the remaining ~24K agent PRs left to carry within-agent power. So the baseline is no longer a wish on this corpus; it has a published construction and a cost. It measures no test-related quantity whatsoever — its metric families are security smells, structural maintainability, post-merge churn, reverts and review behaviour — so neither the presence gap nor the diff-coverage gap nor the edge-case gap moves an inch. What it does add as a caution for whoever runs the test version: static analysis in that study reached only 23.7% of PRs (Python/JS/TS diffs under 5,000 added lines), and a diff-coverage cohort restricted to instrumented repos will be thinner still. See Agent-Vendor Heterogeneity.
    • SourceVendor is confounded with task mix by the paper's own admission — repos and users self-select which agent to use and for what. Does the Codex-vs-Devin revert gap (OR 0.50 vs 1.31) survive conditioning on task type? The discriminator is per-PR task-mix labels, which AIDev does not carry; the authors name it as the missing variable. Until it exists, every "which agent is better" reading of this page is a reading of a product-and-its-users bundle.
    • SourceRevert-by-commit-message misses silent rewrites, and the paper says so. What fraction of agent PRs are functionally undone without a revert-shaped commit? A diff-level check — how much of a merged PR's added code survives at 90 days — is computable from the same cached payloads and would say whether the sub-parity revert rate is a real survival advantage or an artefact of how agent code gets unwound. Partially answered 2026-10-01 by who finishes the job follow up fixes agent prs (empirical, AIDev-pop, ≥500 stars). The forward-fix channel is measurable and runs the other way: 3.68% of agent merges draw a verified fix within 30 days against 2.34% of human merges, within-repo MH OR 1.62 [1.10–2.39]. So at least part of how agent code gets unwound is forward fixes that a revert detector never sees, and the sub-parity revert rate cannot be read as a survival advantage on its own. Not settled: that paper counts fix PRs, not the fraction of added code that survives, and its 30-day window, star threshold and human cohort differ from this page's, so the two rates share no denominator. The diff-level survival measure is still unrun. See Follow-Up Fixes on Agent PRs.
    • SourceClaude Code's median PR is 495 changed lines against 52–96 for every other group, and its three anomalous results (9.5% smell presence, heaviest structure, 12.6 h review wait) all co-move with it. Is there any per-vendor outcome left once PR size is matched? Size-matched resampling within the existing corpus is enough to test it, and a null would collapse most of this page into a statement about PR size. Partially answered 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — Table 1 against Tables 2–4 splits the results in two. The headline, Codex vs Devin reverts (OR 0.50 [0.44, 0.57] vs 1.31 [1.11, 1.54]), compares groups whose median PRs are 63 vs 61 changed lines, so it is not a size artifact at the median. Copilot's commented code against Codex's uncommented code sits on 76 vs 63. Claude Code's three anomalies are the size-sensitive ones: its 9.5% smell presence sits beside a negligible per-line density (δ = +.05), and the paper itself ties its structure and review wait to PR size. A null would therefore collapse the Claude Code results, not most of the page. It is not settled: median matching is not distribution matching, revert odds plausibly rise with size, and no size-matched resampling exists in the wiki. Retagged to #oq/source: it needs that re-analysis.
    • SourceThe window is seven months and the value proxy is coarse/relative. How much of the +27% is genuine task-complexity growth vs. classifier/marketplace-matching drift?
    • SourceThe study excludes headless/SDK/IDE usage — a "substantial share," and likely the most automated/end-to-end. Does including it accelerate or reverse the composition shift?
    • SourceIf "fixing" keeps falling, is that because models break less, or because broken-code work is migrating to non-interactive pipelines this study doesn't see?
    • WaitThe 5.4%→26.6% curve is three months. Is this a durable behavior change or a novelty spike following a Codex skills-feature push? (Cf. the OpenAI-internal training campaigns the paper notes.)
    • SourceCustom skills encode org-specific context — but who maintains them as the codebase and conventions drift? Systematization could itself become a debt surface (Agentic Technical Debt) if skills rot. Partially answered: skill registry to repository lifecycle (empirical, 2026-07) measures the drift directly on public GitHub — 53% of reused skills are never modified after adoption, 40.2% of never-updated copies sit on a changed upstream, and local maintenance is 2.7:1 additive (6.1:1 for locally authored skills), with rename/tooling-substitution chasing the single largest evolution activity (24.3%). So skills do rot and the ratchet is real. Still open: the consequence side — no study yet links skill staleness to degraded agent task outcomes, and the sample (public repos ≥10 stars) excludes the high-complement orgs where the adoption curve is steepest. Extended 2026-09-22: who maintains agent skills longitudinal (empirical, Shen & Hruschka, COLM 2026 workshop) answers the who clause directly for public repositories — every one of 254 substantive edits across five vendor-maintained skill repos is authored or web-merged by a named human account, 62% with an AI co-author trailer in a bimodal per-repository split (93% / 92% / 16% / 5% / 0%) that tracks disclosure culture as much as AI use — and advances the rot clause: the maintained minority is corrected in 38% of edits, failure-triggered in 24–63% depending on coding instrument, stable-to-additive in size (32 grow / 7 shrink / 81 stable of 120), and pruned almost never (deprecation 1 of 254, consolidation 10). Not retired, for the reason the question was asked: its premise is org-internal custom skills drifting against a private codebase, and this sample is five public repositories from AI-tooling vendors selected for sustained activity, where the skill's subject is the vendor's own product rather than one team's conventions. What would settle it: a study of private or org-internal skill directories — a vendor telemetry cut of custom-skill edit histories in the Codex population above, or an org case study — reporting who edits, at what cadence, and whether edits track codebase changes.
    • SourceDoes systematization cause deeper delegation or merely correlate with already-intensive users? The paper shows the association, not the direction. (Not addressed by skill registry to repository lifecycle — it mines artefacts and their diffs, never observing the user or the delegated task, so it can say nothing about direction.)
    • SourceDoes a stale skill measurably degrade agent task outcomes, or do models increasingly route around outdated instructions (Harness Shrinkage as Models Improve)? Gao et al. establish that staleness is widespread and argue stale guidance is "executed rather than read", but never measure the downstream effect; SkillsBench-style evaluation could settle it. Partially answered 2026-09-22: who maintains agent skills longitudinal runs the nearest experiment yet and finds a bounded null in the opposite direction from the question's premise — maintenance buys nothing measurable, so staleness costs nothing measurable either: for 13 public skills with ≥6 substantive edits, the latest version scores −0.09 (95% CI [−0.28, +0.10]) against the earliest in-window version on 143 transfer tasks, null under a length control and under a weaker non-reasoning solver (so not a strong model routing around the old version). The interval admits small effects, the judge panel's ICC (0.52) fell below its 0.6 gate, and the tasks are author-and-model-constructed rather than native repository workflows — so it bounds the pilot-sized effect, not every effect. Open: the same v_old/v_new pairs on an independently sourced task set.
    • WaitIf agentic authoring crosses from <1% toward double digits, does the whiplash become unmanageable before context-engine tooling matures — or does the tooling mature because of the pressure? Partially answered 2026-09-22, and the honest reading is "neither half, but the whiplash changed address": Faros's own successor report (vendor-claim) is the first source to observe the trigger — agents opening 13–14% of PRs at a leading edge — and reports the two halves of "manageable" moving in opposite directions. Per change, it became more manageable: incidents per PR decelerated from +242.7% to +14.5%, and Faros says outright that per-change quality is no longer collapsing. In aggregate, it became less: monthly incidents +125.4%, average task time in QA +300.6% (the largest deterioration in the dataset), PR size +71.8%, and the unreviewed-merge metric growing +76.3% against a prior +31.3%. So the whiplash did not become unmanageable in the way this bullet assumed — it relocated from the individual change to the system throughput around it. Four reasons this cannot retire: the 13–14% and the <1% have different denominators (leading-edge slice vs pooled panel), so the crossing is not established for any one population; "unmanageable" was never given a threshold here and no number supplies one; the report measures no context-engine adoption at all, so the second half of the either/or is simply untested; and the mechanism Faros offers for the new dominant strain — restarts +66.7%, read as agents lacking sufficient context — is the same diagnosis pointing at the same product recommendation as the prior report, i.e. the vendor's prior, not a measurement of whether the tooling matured. What would settle it: a panel-wide agentic-authorship share on one denominator across two windows, paired with any adoption measure for context-provisioning tooling.
    • SourceWhat are the shares of the four examiner classes for agent-applied changes (unexamined, invoker-only, independent-human, AI-reviewed) on one population, over one stated survival horizon? The pieces exist only on incommensurable corpora: 40.1% invoker-only review on Agents in the Wild, 8.8% AI-reviewed on CodAGE (with AI review not excluding a human), and Faros's merged-unreviewed trend on its customer panel. So the rubber-stamp class, bare non-reversion, has no measured size anywhere. It is settled by one dataset carrying reviewer identity, reviewer type and post-merge survival for the same PRs; see What the Agent-PR Oversight Numbers Can and Cannot Say.
    • ResolvedThe 60% figure aggregates very different tools and modes (autocomplete acceptance vs. agent-applied diffs). What does "acceptance" mean when the agent applies the change directly and the human's "acceptance" is not reverting it? Sharpened by Review as the Control Point: on GitHub, agent PRs are most often examined only by the developer who invoked the agent (author-only review 40.1% vs 21.5% for human PRs). Whether that counts as review at all is a definitional choice (agent-as-author ⇒ a second set of eyes; agent-as-tool ⇒ self-review) that literally flips the sign of the trend — so "acceptance" and "review" blur into the same unresolved construct. Partially answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — the construct resolves into a three-way partition by who acts and who looks: affirmative adoption (human applies a suggestion), reviewed non-reversion (agent applies, independent human examines), and bare non-reversion (agent applies, nobody looks — the only rubber-stamp class). The 60% blends all three, which is why it can't answer whether oversight is real; the proposed metric is the partition itself, with only bare non-reversion read as oversight erosion. Still unmeasured: the partition's actual shares in any dataset. Partially answered again, from below (2026-08-12): DECODE measures the affirmative adoption class — the one the partition treats as unambiguous — and finds it is not an endpoint. Of trajectories that begin with a developer accepting a completion, 31% contain a removal edit, retention is bimodal, and the median completion has lost roughly a third of itself within the hour. That does not supply the partition's shares, and it does not touch the two non-reversion classes at all (its unit is a suggestion a human applied, not a diff an agent applied). What it settles is narrower and useful: acceptance is a point on a trajectory, so any partition of it needs a time horizon attached, and the instrument that can see the trajectory is pre-commit editor telemetry rather than anything at PR level. Answered 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — for an agent-applied diff, "acceptance" is non-reversion: a survival measure, not an act, and it carries oversight meaning only with two parameters attached. A horizon (DECODE shows even affirmative acceptance keeps moving; Kraishan's 90-day revert window is what a defined one looks like, and revert-by-commit-message misses silent rewrites, so it is an upper bound on survival). And an examiner class, now four: nobody, the invoker, an independent human, or an AI reviewer (the closed-loop census adds the fourth: 8.8% of attributable agent PRs, 83.7% of them reviewed by the authoring product itself). Faros's 60% publishes none of these, which is why it cannot be read as oversight. The residual is a measurement question, the four classes' shares on one population, and is filed as its own #oq/source question below.
    • SourceWhen does "generate three and compare" become wasteful — at what decision weight is a real argument (or a design doc) still cheaper than three implementations?
    • ResolvedIf design discussion lives in PRs/prototypes, where is the rationale recorded for future readers — does the "why we chose this" knowledge survive, or does it share the staleness problem of Code as Source of Truth? Answered: Rationale as a Dated Record: Where the Why Lives for the Next Reader — it survives only where it is committed as a dated decision record beside the code (execution-plan decision logs, as OpenAI's Codex team checks in; an ADR), so build-to-decide needs one added step: write the comparison verdict into the log when the winner merges. Such a record makes no present-tense claim about the code, so code drift cannot falsify it; its residual hazard is being read as still binding after a later decision overturns it, handled by marking it superseded (close, don't delete). Earlier partial answer: Where Does the Why Live? (left in merged PR threads, the why is orphaned).
    • SourceDoes the Gemini UI gate agree with human designers? It is the gate the post credits with near 1:1 output, and it is unvalidated. The discriminating measurement is cheap from Helix's own archive of UI reviews: a designer independently labels differences on a sample of matched screenshot pairs, and the gate is scored on precision (false blockers force pointless fix cycles) and recall (missed differences reach the engineer), with the INVALID rate reported alongside.
    • SourceDoes oversight actually fall as memory accumulates? The falsifiable form is engineer feedback items per checkpoint, plotted against checkpoint index within a screen and across screens. A second test is the rejection rate at later human review of autonomous-mode checkpoints against approved-mode ones. A flat curve would mean the autonomy is granted, not earned.
    • SourceHow many gate cycles does a checkpoint take, and what share never converges? With unbounded retries and three bar-raising rules (every fixable visual difference blocks, union of reviewer findings, stricter verdict wins), the distribution of cycles per checkpoint, and the gate that sends work back most often, would show whether "reliable convergence" is a property of the loop or of the engineer who steps in when it stalls.
    • SourceThe 8.8% review rate divides a near-complete numerator (0.01% reviewer-side quarantine) by a demonstrably incomplete denominator (38.0% author-side quarantine, 96.9% of it branch-only). What is the AI-review rate on the 1.73M quarantined branch-only PRs? A hand-labelled sample of a few hundred would bound it, and the replication package plus the quarantine files with recorded reasons make it a day's work. Until then no rate on this page may be quoted as an ecosystem rate rather than a rate over the signature-attributable population. Partially answered 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — the rate is reported nowhere in the paper or the wiki, so it needs the labelled sample. The synthesis adds that the quarantine is not the whole hole: §7 says unsigned PRs "are not quarantined but simply never enter our population," so a sample of the 1.73M would bound the review rate only on the branch-prefixed part of the missing population. An ecosystem rate needs a second sample drawn from outside the signature framework. Retagged to #oq/source: it needs new labelling work, not synthesis.
    • SourceSame-product and cross-product are nearly confounded with reviewer identity: the same-product arm is 80% Copilot, the cross-product arm is heavily populated by Cursor, Google Jules and Claude Code, which have no reviewer of their own and so could not have self-reviewed. Does the 58–65% same-product comment-volume gap survive holding the reviewer bot fixed and matching on change size? The Codex row already hints not (0.91 vs 0.89 on 42,732 PRs, the one bot with substantial population on both sides), and the data to test it is released. Partially answered 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — the reviewer-fixed half already holds by construction. Each row of the comment-volume table is one reviewer bot comparing its own product's PRs with others', and the gap is present for three of four bots. So the Codex row is the one bot whose own gap is flat, not a hint that the pooled gap is confounded. With the reviewer fixed, the confound moves to the author: 18,114 of Copilot's 21,022 cross-product PRs (86%) are Codex-authored, so Copilot's 58% is close to "Copilot-authored vs Codex-authored PRs, same reviewer," carrying every size and task difference between the two. The size half is unanswerable here: the paper names change size as differing and does not match on it, and AIDev's medians cannot stand in. Retagged to #oq/source: it needs a re-analysis of the released data.
    • SourceThe paper's own first future-work item is the experiment this vault most wants: review functionally equivalent PRs from known models under fixed configurations with source blinding, scoring defect recall and false positives across same-model, related-model and independent-model reviewers. Does the product-pairing effect measured here survive when the model behind each product is known and held fixed? Nothing observable on public GitHub can answer it — the records do not carry model identity — so this is a controlled-experiment question, not a mining one.
    • SourceWhat knowledge genuinely can't live in the codebase (org strategy, the "why," cross-team context) and therefore still needs a durable doc — and how do you keep that small slice current? Partially answered: Where Does the Why Live? — the "why" is the clearest such slice, and every candidate home fails or only partly works; a compiled knowledge base outside the staling code is the least-bad option. Partially answered: Rationale as a Dated Record: Where the Why Lives for the Next Reader — the "why" can live in the repo after all, as a dated decision record with no drift problem, so the slice shrinks to records whose authority sits outside the repo (regulator-accepted systems of record, other teams' dependencies) and tacit knowledge; the keep-current mechanisms (CI regeneration, doc-gardening agents, same-session write-back to the legacy record) are prescribed but unmeasured, and no source covers org strategy. Needs an external account of drift in an externally-authoritative record under agent write-back.
    • SourceIf onboarding is "ask Claude," what happens to the tacit knowledge that was previously transferred socially in deep-dives — is it captured anywhere, or quietly lost?
    • SourceIs 1% a Thariq-specific number or a regime? For larger, more code-heavy projects the production residue is presumably higher; what sets the ratio?
    • SourceAllocation quality is hard to measure — what's the feedback loop that tells an allocator they spent compute badly (vs. just spending a lot)?
    • ResolvedDoes treating humans as "compute allocators" risk the oversight-fatigue / accountability failure modes the HBR research flags, where the human nominally decides but actually rubber-stamps? Answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — yes, it is the role's central failure mode, documented at three evidence layers (brain-fry error rates +11%/+39% and the under-engagement swap; Faros's 31.3% no-review telemetry; the CMU theory's P1–P4 load-and-plausibility mechanisms), with two allocator-specific aggravators: allocation quality has no feedback loop, and the human's retained 70% planning share is exactly where rubber-stamping is transcript-invisible. The framing survives only with structural countermeasures — understanding-gated merges (quiz gate), sample-based depth concentrated on high-stakes points, risk-tiered gating — that make understanding rather than signature the merge condition.
    • SourceThe "human" is an LLM simulator (GPT-4.1) that also judges — how much of the configuration-dependent structure is a property of human-agent collaboration versus an artifact of GPT-4.1 modeling both sides? The Appendix C.6 simulator-swap is the only check, on a subset.
    • SourceA4 is instantiated as a fixed one-shot proactive intervention. Real proactive humans time their input adaptively — does the "premature/distracting intervention breaks tasks" result hold, worsen, or vanish under a human who chooses when to interject?
    • SourceThe optimal channel is shown to be pattern-specific, but pattern is an oracle label assigned at construction. Can an agent infer which pattern it is in (and thus which channel to solicit) at runtime — the actual deployed skill the paper says agents lack?
    • WaitDoes the "more channels adds coordination overhead" penalty shrink as the backbone improves (a capability gap), or is it a structural cost of mixed-initiative interaction that persists?
    • SourceDo the six smells predict failure? Houck promises a measurement framework (CAFE(S)). Falsifiable: rate context for the six smells before delegation and test whether flagged tasks fail at a higher rate than unflagged ones, controlling for task difficulty.
    • SourceDoes the new-colleague test actually "catch a surprising amount"? No source measures how many agent failures a pre-delegation outsider check would have prevented, or its cost in delegator time against the failures avoided.
    • WaitIs "make the last mile manual" durable or transient? Parrott gives two reasons (tokens, and eyeballing beats describing); the token argument dies with cheaper inference, the bandwidth argument shouldn't. Which one is actually load-bearing is testable by watching whether direct-manipulation use falls as models improve.
    • SourceIs the default-aesthetic collapse fixable by context (brand files, moodboards) or is it the novelty ceiling from Why AI Lags at Design reason 3 — i.e. does a fully specified design system still produce house-style output underneath the palette? Partially answered: astryx agent ready design system test (n=1, single practitioner) says fixable where the design system has a slot for the decision, untouched where it doesn't — the model read the brand correctly throughout and only regressed on components with no customization surface. That reframes the residue as a coverage gap rather than a novelty ceiling, but it tests one vendored system, not the underlying question of whether house style persists on fully covered surfaces.
    • SourceTen-options-then-remix assumes the human reliably recognizes the good one. Where does selection break down — does discrimination degrade when all ten candidates are competent, and is there a candidate count past which review cost exceeds authoring cost? Partially answered (adjacent domain, single-candidate case): Post-Acceptance Edit Behavior (empirical, 53.6K in-IDE edits from 1,141 developers) is the first dataset measuring what humans actually do to AI output they already chose. It says selection is revisable and frequently revised — retention is bimodal (kept essentially whole or discarded, rarely in between), 31% of trajectories contain a removal edit, and the sequence identifies when recognition fails: the customize-first path leads to deletion roughly twice as often as the change-functionality-first path. What it cannot answer is this question's actual ask. It observes code completions accepted one at a time, not a ten-candidate set, so it measures neither discrimination-under-uniform-competence nor the review-cost-versus-authoring-cost crossover — and retention records what the developer did, never whether they were right, so nothing in it grades the discrimination itself.
    • WaitCan grilling be run AFK against another agent that holds the user's preferences? Pocock's answer in 2026 is "no, this part has to be human-in-the-loop" — but the question is open as agents get better at modeling their principal.
    • SourceHow does grilling change for team work where multiple humans need to align? Pocock's hint: pair-program with the agent in the room, treat it as a third interlocutor.
    • NowWhere's the line between a disposable micro-app and tool sprawl? If every edit spawns a bespoke UI, does the workflow fragment?
    • SourceDoes the copy-back-to-markdown round-trip generalize beyond config-shaped data (rules, tables) to richer artifacts?
    • ResolvedCould these micro-apps be templated/reused rather than regenerated — and at what point does that defeat the "disposable" framing and turn into durable tooling? Answered: The HTML Artifact Lifecycle: Where Plan History Lives, and When Disposable Becomes Durable — the correct reuse unit is the generator, not the artifact: a recurring micro-app pattern becomes a skill that regenerates a fresh, fitted app each time, keeping disposable's per-task fit while gaining reuse's consistency (the measured systematization move — skills at 5.4%→26.6% of weekly-active users). An artifact itself graduates to durable only under recurrence + sync-pressure + audience (the design_system.html profile), at which point it inherits maintenance cost, sync cadence, and a seat under the artifact-sprawl bloat ceiling. The failure mode is the un-chosen middle — apps kept around unmaintained: sprawl plus rot with neither fit nor consistency. The discipline is binary: regenerate it, or maintain it; never merely keep it.
    • SourceDoes the imperative-bias and library-avoidance pattern generalize beyond C++? The paper's own closing question names Rust, Go and Java, and the answer decides whether this is a property of generated code or of C++'s particular idiom space (where the "right" call is a specific absl/std API a generalist model has weak priors on).
    • SourceIs the sub-parity revert rate a property of AI-generated code or of the gates around it? Every reliability figure here comes from a review-gated monorepo with mature static analysis and presubmit CI, and the paper's own reading is that the gates catch the fatal errors. The discriminator is the same measurement in an org with weaker presubmit gates — where Faros's incident numbers come from. Largely answered 2026-09-22, and the answer is "a property of the code, or of something upstream of the gates": not all agents are equal five coding agents post merge (Kraishan, Texas Tech, arXiv 2609.17598, empirical) runs the nearest available version of that measurement on public GitHub — 37,623 provenance-labelled PRs across 2,807 repositories, with human PRs drawn only from the 810 repositories that also contain agent PRs, same window, capped per repo. Those repositories have nothing like a centralized presubmit stack. The 90-day revert rate is 11.5% for humans against 6.1% for Codex (OR 0.50), 14.5% for Devin (OR 1.31), and Copilot / Cursor / Claude Code indistinguishable from humans; pooled, ≈7.7%, OR ≈0.64 (arithmetic over Table 2, not reported by the paper). So sub-parity reverts survive the move out of the monorepo, and survive it with more room — which is evidence against the gates explanation, at least for the gate-poor direction. Three reasons this is "largely" and not "fully". The revert construct is different and weaker: revert-style commit messages on the PR's most-changed file, which the authors state misses silent rewrites, against Google's internal revert record. The population is not gate-poor so much as gate-heterogeneous — >100-star public projects still run CI and human review, and every PR counted here already cleared one. And the paper's own alternative mechanism is not about gates or code: "the humans steering them may assign bounded, well-specified tasks," a selection effect that would produce this result at any gate strength. The result it adds that this page cannot produce: the vendor spread. Revert odds run 0.50 to 1.31 against one baseline, which brackets this page's ~0.9× from both sides — so Google's single AI cohort, pooled across whatever assistants its engineers used, may be reporting a mixture rather than a property. Qualified 2026-10-01: "sub-parity reverts" is not the same as sub-parity post-merge failure. Takerngsaksiri et al. (arXiv 2609.26847, empirical, ≥500-star AIDev repositories) count verified forward fixes, which revert detection does not see, and find agent merges fixed at 1.62× the same-repo human odds [1.10–2.39]. The gates question above is unaffected. What changes is that a revert ratio by itself, Google's ~0.9× included, understates how often AI code needs post-merge work.
    • SourceStage 3's mean R_eff of 0.385 sits below the 0.5 parity mark, so the best-mitigated regeneration is still on average worse than the human original it replaced. Does any feedback regime — the RLEF pipeline the authors propose, or context injection from the monorepo — push mean R_eff above 0.5, or is sub-parity efficiency a floor for generalist models on performance-sensitive code?
    • SourceIs agent self-fixing ownership, or repository tool monoculture? The 69.6% (and Copilot's 95%) has no null model. The test is the share of all fix PRs in the same repository and window that come from that agent product, compared against its share of fixes to its own merges. The replication package's linked pairs are enough to compute it.
    • SourceIs the agent excess mostly same-day follow-up, rather than defects that surface later? Fig. 3 puts about 1.7 of the ~1.9-point verified gap on day 0. A split of verified fixes by time-to-fix (same day and same author vs later) would show whether the 1.62 measures latent defects or agent work-splitting.
    • SourceDoes the 1.62 survive task-type conditioning? Agents are assigned bounded tasks (the self-selection Kraishan names), and fix-typed follow-ups may be likelier on some task types. The within-repo stratification does not control task mix.
    • SourceDoes this generalize past one expert practitioner, or does it require Thariq-level fluency with Claude to be worth the overhead?
    • ResolvedDoes the human-facing harness keep growing without bound, or does it hit its own bloat ceiling (an HTML plan too elaborate to read, like the markdown it replaced)? Answered: Does the Human-Facing Harness (HTML Artifacts) Hit Its Own Bloat Ceiling? — yes; HTML raises and reshapes the human-attention ceiling but can't remove it, and the bloat relocates from document-length to artifact-sprawl/rubber-stamping.
    • ResolvedHTML is heavier to diff and version than markdown — what happens to plan history and review when artifacts are single-file websites? Answered: The HTML Artifact Lifecycle: Where Plan History Lives, and When Disposable Becomes Durable — the artifact is a compiled view, not a record: version the content layer (the copy-back round-trip and the extract-from-code pattern already do this), regenerate the presentation on demand, and reattach review to decisions rather than diffs (the plan-ordered-by-likelihood-of-change technique puts the reviewable delta at the top). Presentation history is deliberately discarded — regenerable at abundance prices, it isn't worth versioning — and a presentation choice that becomes load-bearing has by definition graduated to durable tooling with real versioning obligations. Residual (tracked above): no source yet documents team-scale multi-author HTML-plan review.
    • SourceDoes the maintenance-benefit null survive tasks the skill's authors did not construct? The 95% CI [−0.28, +0.10] admits small positive effects, the judge panel's ICC(2,1) = 0.52 sits below the pre-registered 0.6 gate, and the tasks are author-and-model-generated transfer scenarios rather than the repositories' native workflows. The released harness makes the test cheap: re-run the same 13 skills' v_old/v_new pairs on an independently sourced task set (SWE-Skills-Bench-style real repository tasks) and report the skill-level delta with a judge panel that clears its gate.
    • SourceIs the bimodal trailer split AI usage or disclosure culture? getsentry and trailofbits trailer 92–93% of edits, anthropics 5% and cloudflare 0% — yet Anthropic reports >80% of its merged production code authored by Claude. The paper cannot disentangle actual involvement from trailer conventions and squash workflows. Falsifiable by pairing each repository's trailer rate with its merge settings (squash on/off) and the agent tooling its maintainers use, or by asking the maintainers.
    • SourceCan any instrument reliably code whether a skill edit encodes a reusable rule? The abstract three-way axis reached κ = −0.02 same-family and 0.17 cross-family; a narrower binary operationalization (a specific added line states a rule) reached κ = 0.43. If no operationalization clears 0.6, the "generalize rules from human edits" curator design the paper proposes has no supervision signal; if one does, the reported-but-withheld distribution becomes reportable. Falsifiable on the released 254-edit corpus with a new codebook.
    • SourceDo agents revert from instructed TDD to write-then-test because of instruction decay, a training prior, or both? The three have different fixes and no source separates them.
    • SourceIs there any principled way to set the complexity threshold for an agent reader, or is 4 → 6 → 8 pure practitioner feel?
    • SourceThe structural claim is that no layer is sufficient alone, and the paper measures what none of the three layers catches. The falsifiable form is cheap for any team already running all three: attribute each defect found over a quarter to the layer that caught it — preventive (never generated), executable (broke the build), human (caught in supervision or review), or escaped to production — and report the four shares. Until that exists, "layered supervision" is a description of where effort went, with no evidence that the distribution outperforms the review-centric arrangement it replaced.
    • SourcePreventive guardrails are defined here as guardrails nothing checks, whose effect is conditional on the generator consulting them — and the practitioner answer is to duplicate every relevant rule into an executable check. Does the steering artifact then contribute anything? The discriminating ablation is one repository and three arms: the rule in the steering file only, in the lint rule only, and in both, scored on violation rate at first generation rather than at merge. If the both-arm matches the lint-only arm, the preventive layer is buying trajectory quality rather than conformance, and should be argued for on that basis instead.
    • SourceThe authors refuse "maturity" for a three-dimensional configuration space — governance posture, system criticality and homogeneity, team composition — because five points show no linear progression. Is the space genuinely non-ordered, or is the refusal an artifact of n = 5? The survey instrument that produced two of the three dimensions already exists and is published on Zenodo; administering it at a scale that supports ordering, and testing whether guardrail-layer adoption is monotone in any of the three, separates a real finding from a sample-size artifact.
    • SourceHow does the design_system.html stay in sync as the codebase evolves — re-extract on a cadence, or wire it into CI? And is completeness (does every component the product needs have a slot?) the binding constraint rather than freshness — the Astryx test suggests uncovered components fail silently, which staleness checks wouldn't catch.
    • SourceDoes a rendered, model-readable design system measurably improve on-brand output vs. a plain CSS/token file, or is the win mostly human legibility? Partially answered: astryx agent ready design system test shows structured artifacts (DESIGN.md + layered tokens) flip output from generic to on-brand, so the artifact clearly beats no artifact — but it does not isolate the variable this question actually asks about, since the artifacts used were plain files, not a rendered page. Rendered-vs-plain remains untested.
    • SourceAt what project size does maintaining the artifact cost more than the consistency it buys?
    • SourceThe authors couldn't locate the saturation point because synthesis was manual — how few documents actually suffice, and can a cheaper sample match the 3,100-doc theory?
    • SourceAutomating the codes→theory step failed with a naïve bottom-up prompt; is that a prompt/scaffolding limitation or a genuine ceiling on LLM interpretive synthesis over thousands of codes?
    • SourceThe three-lens design manages coder bias, but the relevance judge and segmenter are single-model — do those upstream gates impose their own systematic slant on what reaches the codebook?
    • SourceDoes agent-authored contribution measurably change merge rate, revert rate or post-merge defect rate in a repository, versus human-authored contribution to the same project? Omarchy's numbers are self-reported and have no control. Partially answered 2026-09-22 by not all agents are equal five coding agents post merge (Kraishan, arXiv 2609.17598, empirical), which builds exactly the control this bullet asks for: 4,027 human PRs kept only from the 810 repositories that also contain agent PRs, restricted to the same December 2024–July 2025 window and capped per repository under a fixed seed. Two of the three named quantities land. Revert rate: answered, and the answer is per-vendor rather than per-authorship — within 90 days of merge, human 11.5% against Codex 6.1% (OR 0.50), Devin 14.5% (OR 1.31), and Copilot / Cursor / Claude Code statistically indistinguishable from humans after BH correction. Post-merge defect rate: proxied, not measured — reverts and size-normalized churn are the proxies, and the authors state that revert detection by commit message misses silent rewrites, so a PR quietly rewritten out of existence counts as surviving. Merge rate: not answered. The corpus table reports a merged share by group (43.0% Copilot to 82.6% Codex against a human 76.4%), but that column pools all 33,596 agent PRs across 2,807 repositories while the human row is the capped 810-repo baseline, and the paper runs no test on it — descriptive, not controlled. Scope to carry with all of it: public repositories above 100 stars, Python/JS/TS only, and the maintainer in this dataset is the median >100-star project, not a maintainer with a weekly-doubling backlog. See Agent-Vendor Heterogeneity for the full treatment. Partially answered again 2026-10-01, on post-merge defect rate, by who finishes the job follow up fixes agent prs (Takerngsaksiri et al., arXiv 2609.26847, empirical). It is the closest the vault has to a direct defect measure: a later PR that repairs the merged change, human- and judge-verified, within 30 days. Merged agent PRs in ≥500-star AIDev repositories draw one at 3.68% against 2.34% for human merges in the same window, with a within-repo MH odds ratio of 1.62 [1.10–2.39] across 218 shared repositories. That is the opposite direction from the revert result above: agent code is reverted less but fixed forward more. For the maintainer this page is about, the cost is partly self-absorbed, because 69.6% of those fixes come from the same agent product. Merge rate is still not answered: this paper studies merged PRs only. See Follow-Up Fixes on Agent PRs.
    • SourceDHH claims rejection is socially cheap because "the clanker won't mind" — does the human who dispatched the agent experience rejection the same way, or does the submitter-side cost simply become invisible to the maintainer?
    • WaitKarpathy's open frontier: can "understanding" itself eventually be automated, or is it definitionally the human residue? His "back in a couple years" hedge leaves it open. Relevant prediction logged 2026-09-21 from after math de toffoli duede tao guest post (practitioner-opinion, no measurement): De Toffoli and Duede, arguing against current systems, nevertheless expect the automation branch — future AI is "likely to produce genuine proofs that are at once formally certified and fully intelligible to mathematicians," i.e. machine-delivered understanding, not merely machine-delivered results. Two philosophers of mathematical practice and Karpathy land on the same hedge from opposite motives, which is worth noting and settles nothing.
    • SourceIf understanding is the bottleneck, is the highest-ROI skill learning how to build understanding fast (knowledge-base hygiene, asking the right projections) — and can that be taught? Partially answered (2026-09-22) by psychological costs ai adoption software engineering (case-study, N = 21 at one regulated-software firm one year into adoption), on the first half only. It is the corpus's first workplace inventory of what practitioners converge on unprompted when understanding becomes the binding constraint, and the list is not knowledge-base hygiene: a daily manual-coding ritual ("if you don't use it, you lose it"), attempting a solution before consulting the model to protect independent reasoning, delegating boilerplate while retaining business-logic-heavy work, prior codebase knowledge treated as a prerequisite for oversight rather than an output of it, and complexity-based delegation triage. Every one preserves the capacity to understand rather than accelerating the act of understanding — so the practitioner answer to "highest-ROI skill" is closer to not losing the capability than to building understanding fast, which is a meaningfully different bet than this bullet assumes. The teachability half is untouched: the authors prescribe curricula in "AI-assisted verification" and name self-then-AI sequencing as a pedagogical candidate, but that is a recommendation section, not a result, and nothing in the study measures whether any of these practices preserves anything. Extended (2026-09-22) by training novices to think or giving them llms rct (empirical, preregistered 2×2 RCT, n=1,053 first-year undergraduates), which supplies the teachability half — randomized, and for a reasoning skill rather than a hygiene practice. A ~6-minute game (12 binary-choice items in four sections, each with a worked example and feedback, against a placebo posing the same items with no scaffolding) raises mechanism identification ~0.55 SD and falsification logic ~0.85 SD in writing produced afterwards, and raises idea diversity ~0.47 SD between solutions. The part that bears directly on this bullet: the trained skill is not crowded out when a capable model is in reach — every interaction of training with ChatGPT access is positive and significant on top of both main effects (mechanisms +0.129, falsifiability +0.203, coherent logic +0.228), and the diversity effect is identical with and without the tool. So a cognitive skill can be taught cheaply and is still exercised rather than delegated. Three limits keep it partial. The skill taught is causal reasoning, not "building understanding fast" — it is the disciplined-thinking half of the bullet, not the knowledge-base-hygiene half. The outcome is reasoning style in the assisted output, measured by an LLM rubric; there is no unaided post-measure, no tool withdrawal, and the authors say outright they cannot tell knowledge acquired from output procured. And on the paper's own graded outcome the training did not pay (−0.111, p<0.01) — because the rubric penalized mechanisms, falsifiability and distance from the modal answer — which is a warning this bullet should carry: teachable and used is not the same as rewarded.
    • WaitDoes the human share of planning decisions fall over time as models improve (the ceiling rising into the planning layer), or is ~70% a stable human floor?
    • Source"Decision attribution" is inferred from transcripts. When Claude proposes a plan and the user assents, is that scored as the user's planning decision or Claude's? The rubber-stamping boundary is exactly where the measure is hardest. Partially answered on the execution half only (2026-08-12): DECODE shows the inference is avoidable below the planning layer — an edit trajectory records what a developer changed in an accepted completion as a byte-level fact, with no transcript reading, and 56% of those edits change functionality rather than naming. The rubber-stamping boundary this question actually asks about is untouched, because assent to a proposed plan leaves no edit trace at all; the residual claim is that the hardest attribution problem is specific to planning, and that the execution share is measurable without a classifier if you instrument the editor rather than the conversation.
    • SourceHeadless/SDK/pipeline usage (excluded here) is where execution autonomy is highest and planning is front-loaded into a single prompt — does the 70/20 split survive there, or collapse toward full delegation?
    • SourceThe completion pool is 2024-to-early-2025 inline autocomplete, median 9 lines. Do the bimodal retention shape, the 15-minute knee and the ~31% removal-edit rate hold at 2026 agentic granularity, where the unit is a multi-file diff the developer never watched being written? The discriminator is the same trajectory extraction run over agent edits rather than completions.
    • SourceThe customize-then-remove path (23.4%, against 12.2% after a functionality edit) is offered as evidence that subtly misaligned completions resist adaptation. A reading-depth explanation predicts the same matrix: customizing requires reading the completion closely, and close reading is when its real flaw surfaces. Discriminating them needs a signal outside the edit stream — time-to-first-edit conditioned on completion length, or an eye-tracking or dwell-time proxy. Which mechanism is right decides whether the fix is better generation or earlier forced inspection.
    • SourceIs completion retention predictable in principle? Fine-tuning lifts classification only to F1 0.45 against a 0.33 random baseline, and generation on the dominant edit class (changing functionality, 56% of snapshots) tops out at 0.49 Levenshtein similarity for every model tried. Either the signal is in context the models were not given (repository, task history, the developer's other files) or retention is a property of intent that no amount of code context contains — and the paper's proposed "detect low-editability generations before showing them" product depends on which.
    • SourceEvery one of the 67 relationships is a hypothesis, not a finding — the paper's explicit call is for causal-estimand studies (controlling for the other constructs) to confirm, reverse, or drop each edge. Which of P1–P17 survive measurement? Partially answered (direction only, no edge confirmed): Security Debt of Agent-Generated Code (empirical, arXiv 2607.12428) supplies outcome-level evidence pointing the way P1 (load → shallower review) and P4 (surface plausibility disarms the reviewer) predict — 81.1% of genuine credentials in agentic PRs drew no reviewer comment before integration, and the reviewer-focus finding it cites (Haider & Zimmermann, arXiv 2601.19287) is that inline comments on AI-authored code address logical and functional correctness rather than security posture. But it isolates no mechanism, controls for none of the other constructs, and has no human-PR baseline, so it corroborates a direction without confirming an edge. First production test, and it is a null (2026-08-12): Tran et al. have the human baseline the security paper lacks and find no correlation between review time or iteration count and the survival of inefficient AI-generated code — so on the one outcome class they measured, the P1 chain does not reach the outcome. It is reported without a statistic or specification, and it tests one defect class in one review-mature monorepo, so it does not reverse P1 either. The usable result is narrower and new: the mechanism map needs a defect-class dimension, because review depth cannot plausibly moderate what review cannot see. The dimension gets its positive pole the same day: Dipongkor et al. (empirical, 4,882 agentic PRs) identify a class where attention does have purchase and say where to point it — error-handling constructs, unexercised 81.0–86.0% of the time in both languages whether or not the agent wrote tests. So the two ends of the dimension are now instantiated rather than merely postulated: a class review cannot catch at any depth, and a class review can catch if told where to look. P8/P9 remain untested (2026-08-12): Cynthia et al. is the largest study yet of automated reviewers but measures the adoption of their comments (71.4% pooled), not the throughput P8 claims or the quality P9 contests — so the two edges nearest to automated review still have no measurement, only a new third quantity beside them. A fourth quantity, and the closest approach yet to P9 (2026-08-12): Greptile measures automated reviewers' recall on ~1,500 labelled high-severity bugs (52–62% by arm) — not adoption, not throughput, and nearer the quality half of P9 than anything before it, but still not a quality outcome, since it scores detection against a self-built label set rather than defects that shipped. It is case-study, not empirical, and it moves the moderator rather than the edge: automated-reviewer capability turns out to depend on which model authored the code under review. The moderator moves a second time, and the same way (2026-09-22): goldman code review comments developers resolve ase 2025 (Melbourne / Atlassian, ASE 2025, empirical) runs one LLM reviewer across two corpora and finds its comment-type mix flips between them — on Atlassian's internal code it files 2.8x the human share of bug comments and 1.4x the maintainability share, while on open-source CuRev the bug advantage collapses (20.1% vs 18.1%) and humans file more design comments than it does (23.6% vs 15.7%). So automated-reviewer capability is not a scalar property of the model on either axis measured so far: it depends on who authored the code (Greptile) and on which codebase it is pointed at (Goldman). It still does not touch P8 or P9 — it measures neither throughput nor any quality outcome — and its adoption figure (39.3% of 4,000 comments, on an exact-line-modified construct) is a third incommensurable quantity rather than a replication of the 71.4% above; see the Agent Review Comment Resolution page for why the two do not average.
    • WaitDoes the no-review convergence hold as agentic authoring crosses from <1% of PRs toward double digits, or does the early-adoption discipline break down under volume the way Faros predicts? (Bears on it but does not settle it, 2026-09-22: not all agents are equal five coding agents post merge reports per-vendor review coverage spanning 5.4% to 51.2% of PRs across the same period. That is AIDev review-table coverage, not a no-review rate — it says how often a review record exists for a PR, not whether a human looked — so it cannot be read against the >50%→~14% trend, and the study has no human arm to converge toward. What it does add is that any single convergence number is an average over a tenfold spread between products. Trigger event unchanged.) Partially answered 2026-09-22 — the named predictor's own follow-up, and the trigger half-lands: Faros's The Speed Trap (vendor-claim) reports agents opening 13–14% of PRs at a leading edge (against "<1%" in the report this question was written against) and the unreviewed-merge metric growing +76.3% period-over-period, up from +31.3% — Faros's prediction vindicated in direction, on Faros's own instrument, with its own reading that review capacity is not scaling and ad hoc skipping is calcifying into a default. It does not retire the question, for three reasons. The two authoring figures have different denominators (pooled panel share vs an unspecified leading-edge slice), so no population is shown to have crossed <1% → double digits. The 76.3% is a growth rate in a count, not a share — the infographic formats it identically to every other growth row and the prose is ambiguous — and a growth rate cannot be read against the >50%→~14% share trend this question is about; a rising count and a falling share coexist under volume growth, which is exactly the non-comparability The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence resolved for the +31.3% figure one window earlier. And there is no agent-vs-human split of the unreviewed PRs, so the report cannot say agent PRs are the ones skipping review, which is the entire content of "discipline breaks down under agentic volume." What would settle it: an agent-vs-human breakdown of the no-review population on a calendar axis in any population, or the gated report's definition of the metric. Trigger event updated — the crossing is now observed at a leading edge, so the remaining requirement is a share-denominated no-review measurement split by authorship at that adoption depth. (Instrument note, 2026-09-22, not an answer: ai to ai code reviews shows the review stream on agent PRs is itself increasingly agentic — 248,641 agent-authored PRs with an AI-attributed review by 2026, up two orders of magnitude across 2025 — and that its "closed loop" does not require the absence of a human. So whichever study eventually supplies the share, it has to separate AI-attributed review events from human ones or it will measure convergence in a mixture. See Closed-Loop AI Review.)
    • SourceThe paper's own question: which decisions, under which conditions, push the system toward the virtuous loop rather than the vicious one? — the system-dynamics leverage-point analysis it gestures at but doesn't run.
    • SourceThe three contested edges (automated review → quality/security, P9; governance → latency, P17; and one more) are contested because their sign is moderator-set — what are the moderator thresholds that flip them?
    • SourceIs the approval drift habituation or calibration? Rising approval is predicted both by lowered scrutiny and by correct learning that an agent's PRs are usually fine. The paper names post-merge defect rates as the clean test and does not have them. Linking these AIDev reviews to verified follow-up fixes (the Follow-Up Fixes on Agent PRs construct, same corpus family) and asking whether late-half approvals are fixed forward more often than early-half ones would separate the two.
    • SourceDoes the per-reviewer embedding-drift signal track anything a team cares about? It is validated only against sequence position, a label chosen to avoid leakage. It has not been tested against missed defects, reverted merges, or a reviewer's own report of reviewing less carefully. A monitor that fires on 60% of reviewers is useless if drift and harm are uncorrelated.
    • SourceDoes the slope persist or flatten beyond 207 days, and does it appear in enterprise review with required approvers? The corpus is seven months of open source, where a reviewer can simply stop reviewing. The linear fit predicts continued drift, and a longer panel or enterprise telemetry would falsify it.
    • SourceDoes mutation-testing or complexity-gate adoption actually track agent adoption, or is Martin's revival idiosyncratic to a practitioner who already owned the tools? Partially answered 2026-09-22 by layered supervision ai assisted software engineering (case-study, five unrelated practitioners, ESEM 2026 Software Engineering in Practice) on the idiosyncrasy half only. Four of the five describe executable enforcement expanding in scope specifically because AI-generated volume outran review — extended checks for architectural violations, code duplication and dependency misuse; the build system as arbiter — and one states the economic stance independently of Martin: "as long as I can do it automatically it costs me nothing . . . I really want this convention to be strictly upheld." So the revival is not one practitioner's habit. The adoption half is untouched and arguably not addressed at all: none of the five names mutation testing, CRAP score or a complexity gate, the checks described are conformance and policy rules rather than the remediation-heavy tools this page is about, and five interviews cannot establish a trend in any case. The question still wants a population measurement — mutation-testing or complexity-gate presence in repositories, stratified by agent-authorship share.
    • SourceWhat is the real ceiling on gate stacking — at what number of must-pass gates does the agent's throughput advantage over a human disappear? Martin says he has not found it.
    • SourceThe case study reports volume and never efficacy. What is the escaped-defect or incident rate of auto-approved PRs versus the human-stamp baseline the Slack channel used to produce? PostHog has both populations in its own history, which makes this a checkable before/after rather than a request for new instrumentation. The same history answers a second question for free (added 2026-08-12): plotting the share of merged PRs that would clear the 500-line/20-file ceiling, quarter by quarter, dates how fast the gate's coverage decays as the ambient PR-size distribution rises. (Bears on it without answering, 2026-10-01: Duckbill's before/after (what is happening with code reviews, practitioner-opinion, relayed first-party figures) is the first before/after on any tiered gate in the corpus, but it reports merge volume and latency only, and the gate change landed together with a guardrail upgrade. The defect half is still unmeasured anywhere.)
    • SourceA keyword deny-list is a proxy for blast radius, and Security Debt of Agent-Generated Code shows the proxy misses where the measured debt concentrates (CI/container plumbing, 87.6%). Does adding CI/IaC paths to the deny-list restore coverage, or does it shrink the auto-approvable set so far that the gate stops paying for itself? A third candidate predicate is now priced but not tested (2026-08-12): Dipongkor et al. (empirical) show that "the existing suite covers this diff" is computable, deterministic, and rarely true — 27.0% of changed lines in Python, nothing at all in 64.8% of PRs. That reframes this bullet as a choice between two extensions with the same failure mode rather than one open question: both a path deny-list and a coverage floor buy precision by shrinking the auto-approvable set, and neither has been measured against the volume it costs. The same before/after PostHog already has in its history would answer both at once.
    • SourceIf the size ceiling becomes a design target — agents instructed to emit stacks under 400 lines — does total risk fall, or does it just redistribute into more PRs each below the gate, with the integration risk moving to the seams between them? The security study's size gradient is measured per-PR and cannot distinguish these.
    • SourceThree tiering keys are now on this page — the change (size ceiling, deny-list), production (σ control bands), and the reversibility of the action (the SOC ladder). Only the first has a deployment with a coverage number attached. Does a reversibility-keyed tier admit a larger auto-approved set at equal escaped-defect rate than a size-keyed one? Checkable in the same PostHog history the first question asks for: partition merged PRs by how cheaply the change could be undone (revert-only versus data migration, config, or a released artifact) instead of by line count, and compare the auto-approvable share and the incident attribution of each partition. The prediction from the SOC ladder is that reversibility dominates size, since every gate that tiers on undo cost stops at exactly the point where undo stops being possible. A qualitative datum, 2026-10-01: when Spotify's fleet auto-merge let through a dependency upgrade that broke production (ai changed how spotify builds quality at higher velocity, case-study), the fix it reports leaned toward cheaper undo (more rollback capacity, owners on hand when changes land) more than toward a stricter gate. That is consistent with the prediction but measures nothing. A second qualitative datum, 2026-10-01: two practitioners in Orosz's survey (practitioner-opinion) key human review on undo cost without naming it. Duckbill routes non-additive schema changes to a human (additive ones may pass), and Sigil's Luo reviews only the schema because "everything, besides data, is fluid and recoverable". Also consistent, also unmeasured.
    • SourceEvery number here rests on a ground truth Greptile built from "sentiment analysis, upvote/downvote ratios, and git archaeology," with no protocol, no agreement statistic, no released artifact, and an LLM judge doing the matching without published validation. Does the crossover survive on a label set someone else constructed — a curated defect corpus with human adjudication, or an injected-bug benchmark where ground truth is golden by construction? Until then the direction is a vendor's finding and the magnitudes are uncheckable.
    • SourceIs the blindness a property of shared training lineage or of stylistic fit? The result is compatible with a much duller explanation: each reviewer happens to be strong on the bug mix the other agent produces, for reasons unrelated to authorship. The published category data does not settle it (composition reproduces only ~7% of the effect, which argues against the dull reading but is not a test of it). The discriminating experiment is cheap and Greptile has the datasets: run a third-family model — Gemini, or an open-weight reviewer — across both corpora. If lineage is the mechanism, the third model has no same-model arm and should score near the cross-model band on both. (Bears on it, does not settle it — 2026-09-22: ai to ai code reviews establishes at population scale that the dull channel is real and large. With CodeRabbit held fixed as reviewer and the author varied, its comment-category mix moves 24.5pp between Claude Code- and Copilot-authored PRs across 35,248 classified comments — a pure author main effect on reviewer output, no same/cross contrast in it, since every one of those pairs is cross-product. That does not touch the crossover, which is what survives both main effects; it does mean "the two corpora simply differ" is no longer a hypothetical alternative. It cannot be the discriminating test this bullet asks for, because the study measures no recall against any ground truth and attributes at product level rather than model level — and its own future-work section proposes exactly the missing experiment: functionally equivalent PRs from known models, source-blinded, scored on defect recall and false positives across same-model, related-model and independent-model reviewers. See Closed-Loop AI Review.)
    • WaitCaridad predicts the effect shrinks as models converge: "a year ago, the performance difference in the opening figure would likely have been larger." Does the same protocol, rerun on the next generation of both families, show a narrower gap? The trigger event is the next paired frontier release measured the same way; note that the prediction is also the one Greptile's own product would least like to be true.
    • SourceThe 38.9% smell rate has no human-authored-PR control over the same high-risk paths, and the 2.7×-vulnerability figure it leans on traces to a vendor blog and a Substack post. Does agentic authorship raise smell density, or merely raise the volume of CI/IaC files touched? A matched human baseline on the same path set would settle it. Partially answered 2026-08-12 by Tran et al. (empirical, Google, arXiv 2608.06640): the corpus now has a matched human baseline — 3.52M production changes with authoring-time provenance and a human-written cohort — and on its taxonomy AI-generated code is below parity on Correctness and Safety (0.94x), API misuse (0.93x) and lifetime/ownership hazards (0.71x), with the excess concentrated in efficiency and interface coupling instead. That is direct evidence against the 2.7x vulnerability figure this page marks as uncorroborated. It is not the same path set: application C++ scored by clang-tidy categories, not GitHub Actions / Dockerfiles / IaC scored for security smells, and the AI cohort there does touch more files per change (3 vs 2 median), which is the volume half of this bullet left unmeasured. So the density-vs-volume question survives, with the prior it was testing now leaning the other way. Further partially answered 2026-09-22, and this time on the same corpus family: not all agents are equal five coding agents post merge (Kraishan, Texas Tech, arXiv 2609.17598, empirical) is the first AIDev study with a human-PR control — 4,027 human PRs kept only from the 810 repositories that also contain agent PRs, same window, capped per repo under a fixed seed. On smell presence it finds pooled agents below humans: 2.9% vs 4.6%, OR 0.63 [0.47, 0.85], χ²(1) = 8.59, p = .003. On smell density — the quantity this bullet actually asks about — every pairwise Cliff's δ against humans is negligible (Codex −.02, Copilot −.02, Cursor −.03, Devin ns, Claude Code +.05), so the authors' own summary is that per-line density is similar across authorship. That is the density half answered for its path set, which is not this one: regex over added Python/JS/TS application lines, against this page's LLM-judged CI / Dockerfile / IaC high-risk paths, so the two differ on instrument, language set and file type at once and the 38.9% and the 2.9% are not comparable numbers. The volume half is untouched — nothing there counts how many CI/IaC files each cohort opens. Two results worth carrying regardless: the pooled agent-human density gap lives entirely in the XL PR bucket (δ = −.08, p = .025, no difference in the four smaller buckets), which is the size gradient this page measures arriving from a second direction; and on the credential class specifically, agents are cleaner than humans — hardcoded credentials in 0.9% of agent PRs vs 2.2% of human PRs (OR 0.39 [0.25, 0.62], p < .001) — which is the same unexpected side this page's own RQ2 landed on by a completely different route. Vendor spread swamps the pooled effect in both directions: 1.6% (Cursor) to 9.5% (Claude Code) on presence, and a fourfold spread on credentials alone. See Agent-Vendor Heterogeneity.
    • SourceDoes "no reviewer comment" mean undetected? 60 of the 74 genuine credentials were removed without a comment, yet the abstract reads the same rows as detection failure. Commit-history analysis of who removed them and when is a checkable discriminator between silent remediation and coincidental churn.
    • SourceReview coverage of agent PRs is converging toward the human baseline while efficacy on credentials sits near zero. Do the two trends move together as teams mature, or independently — i.e. does showing up to review buy any measurable catch-rate improvement? Not answered by the closest-looking evidence (2026-08-12), and the resemblance is a trap. Cynthia et al. (empirical, arXiv 2607.21997) is the largest efficacy-adjacent measurement of the review layer to date — 54,713 agent review comments, 71.4% resolved — but it measures the opposite direction: the agent reviewing a pull request and the human deciding whether to act, rather than a reviewer catching defects in agent-authored code. Its resolution rate is an adoption proxy, and its own construct-validity section concedes that "comments may be resolved without being useful or remain unresolved despite being valuable." What it does supply is the negative half of the diagnosis: developers in its argued sample are engaged, not asleep, so whatever suppresses catch-rate on this page's credentials is unlikely to be plain inattention. The question still needs a study that pairs review presence with a defect ground truth on the same PRs. (Bears on it without settling it, 2026-10-01: Reviewer Habituation on Agent Pull Requests shows that showing up is not a constant. The same reviewer approves agent PRs more often as her exposure grows, so a coverage series that converges can hide a per-reviewer approval series that drifts. It still has no catch-rate.)
    • WaitDoes the cost of change actually stay near zero as a system grows, or does it re-inflate at some size — which would restore the case for up-front design and settle Martin's self-flagged prediction?
    • SourceDoes storage-per-se move any outcome once content and read-back behaviour are held fixed — the same specification supplied as a committed file vs. pasted into the conversation, with the agent required to re-read it before each decision in both arms, scored on rework, defect rate or drift? Nothing in the corpus runs this; Khatri's with/without ablation harness (the correctness null on the agent-context-files page) is the closest existing instrument. A null retires persistence as even a proxy; a measurable effect gives Is Persistence the Line Between Prompting and Spec-Driven Development?'s authority line its first measured competitor.
    • ResolvedIs persistence the right definitional line between prompting and spec-driven development, as Pocock proposes and Martin's practice assumes? Answered: Is Persistence the Line Between Prompting and Spec-Driven Development? — no. Persistence is an observable proxy for authority (is later work judged against the artifact?), and it misclassifies the corpus in both directions: CLAUDE.md, WORKFLOW.md, REVIEW.md, the ticket graph and Martin's own dependency-rule file are persistent and returned-to without being specs, while the grilling transcript, Carey's why-conversation and Martin's deleted plans are ephemeral and do the entire specification job. Both cited positions are also misread here: Pocock's sentence is a conjunction — "are you persisting your specifications… are you returning to those specifications?" — and only the second clause survives (he closes rather than deletes the PRD, revoking authority while keeping the file); Martin persists constraints a checker enforces and discards only plans that nothing enforces. practitioner-opinion throughout, and nothing in the corpus measures persistence-per-se.
    • Is there a non-vendor telemetry dataset large enough to adjudicate the maturity-protection question independently of Faros's commercial framing? Partially answered: CMU's arXiv 2607.07980 supplies exactly this — a non-vendor, 2.5M+-PR GitHub telemetry study — and it (a) finds the agent no-review rate converging toward the human baseline rather than a widening gap, and (b) argues the effect's sign is set by team practice, closer to DORA. The catch: its own headline is that the telemetry is direction-unstable, so it counterweights Faros without cleanly settling the maturity question — the honest verdict is "surface telemetry alone, vendor or not, can't adjudicate this." The worked example of that verdict is The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence: the two telemetry studies' opposite under-review headlines dissolve once metric, population, time axis, and authorship unit are aligned — the datasets never disagreed, only the framings did.
    • SourceDoes anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction? Kalff & Simbeck document respondents excluding predictive ML from the category while counting ChatGPT in it, which would bias reported adoption down for exactly the tool class their headline finding is about. Falsifiable two ways: a split-ballot survey where one arm names the underlying technique ("software that scores turnover risk") instead of the label, or validating firm-level self-reported adoption against vendor-license/spend records for the same firms — the Ramp instrument applied as a criterion rather than a substitute. Related evidence (2026-08-12), not an answer: DX reports a self-reported AI-generated code share of 34% to 52% across Q1-Q2 2026 in 500+ organizations, while Google's authoring-time provenance over the overlapping window reads 28.99% to 68.62%. Different populations (a cross-industry customer panel vs one C++ monorepo), different constructs (code share vs adoption), no matched firms, and DX's question wording sits in a gated PDF the vault does not hold — so this settles nothing. It is nonetheless the corpus's first side-by-side of a self-reported code share against a provenance-measured one, and the direction is the one Kalff & Simbeck's construct-collapse mechanism predicts: self-report reads below the instrument that counts bytes, not above it. The criterion validation this bullet asks for is that same comparison run on the same firms. A second instance of the mechanism, documented but unquantified (2026-09-22): mckinsey state of ai 2026 road to roi publishes the most-cited AI adoption series in business media — 21% in 2017 to 89% in 2026 — with a footnote conceding the question was redefined four times over that span (2017 "a core part of the business or at scale"; 2018-19 embedding at least one AI capability in processes or products; 2020+ adopted AI in at least one function; 2025-26 regular use in at least one function). Each revision loosens the threshold, so an unknown share of the series' slope is the definition moving rather than the behaviour. The same edition discloses a second break: prior years asked about AI's cost impact at the use-case level and rolled up, 2026 asked at the business-function level directly. This is the bullet's mechanism caught in the wild at the largest scale available — and it is still not the answer, because no year is fielded under two definitions, so the effect size the split-ballot design would measure is exactly what remains unmeasured. What it adds is that the question is not hypothetical: a decade-long trend line in wide circulation already contains it. The same mechanism on the telemetry side, which this bullet assumed was immune (2026-09-22): adoption telemetry enterprise ai production signals reports one behavioural dataset (ActivTrak, 120,620 workers) reading 82% adoption under a sustained-usage definition and ~2% under a workflow-embeddedness one — same logs, ~40× spread, definition alone. Secondhand and practitioner-opinion, so not the criterion validation asked for; it does mean the implied remedy (validate self-report against telemetry) must pin the telemetry threshold first, or inherit two definitional degrees of freedom instead of one.
    • ResolvedSurveys and telemetry measure different things (felt productivity vs. system outcomes); is the "contradiction" partly a category error — both true at their own layer — rather than one being wrong? Answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — yes for the measurement halves: surveys capture felt productivity (genuinely real at the individual layer — task completion did rise) while telemetry captures system outcomes that haven't propagated into feeling yet; during a fast transition the layers legitimately diverge, and the AEI's linked telemetry+survey design already treats them as complements. But not entirely: the maturity-protection claim is one proposition about one layer (system outcomes) and remains substantively contested — Faros's "no protection" aligns with its vendor incentive, CMU's moderator theory argues the sign is team-set (closer to DORA) but measures no maturity effect. The category-error dissolution cleans the framing; the one real disagreement stays open (tracked in the non-vendor-dataset question above).
    • SourceDoes the cost premium of iterating on an incoherent codebase actually fall as agent context budgets grow, or does re-derivation cost stay flat because the agent re-reads regardless?
    • SourceDo successive model generations get measurably more tolerant of incoherent code — the capability question Martin's account turns on, and which would separate his mechanism from DHH's?
    • SourceThe greenfield/legacy split (Omarchy 100% vs Basecamp "surprisingly tricky") is unexplained by the token-scarcity account — is the binding variable codebase age, user count, or architectural coherence?
    • SourceThe chain's own indicators are the test of it, and Anthropic has the population to run it: across customers adopting these plays, does rework after build starts (spec.md commits dated after the first plan.md for the same change) actually fall, or does the earlier, cheaper artifact simply move where the churn lands? A falling intent→spec latency with flat rework would mean the chain sped up document production without improving decisions.
    • SourceDoes an artifact chain make review deeper or only earlier? Faros measured review time exploding and Tran et al. measured blocking threads at 1.92× human; the playbook predicts both fall once mechanical evidence arrives attached. Nothing yet measures a cohort with committed plan.md gates against one without on the same quality outcomes.
    • SourceWhich chain position degrades first under corrector fatigue? The design's compounding risk is that spec.md is generated from intent.md and the PR is checked against plan.md, so an unread early artifact becomes the standard later gates measure against — but no source in the corpus measures review attention by artifact type, only by diff. (Unchanged, 2026-10-01: Reviewer Habituation on Agent Pull Requests is the first over-time measurement of corrector drift, and it too is diff-only, covering PR approvals and inline comments.)
    • SourceNg asserts the developer's QA burden fell "significantly." Faros's 2026 telemetry measures the opposite for production orgs. Is the split really 0-to-1-vs-production, or is Ng's self-report subject to the same optimism bias the survey literature keeps finding? Partially answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — both-and, not either/or. The scope split is real and does most of the work (0-to-1 builds lack the review functions that make verification expensive in production: no queue, no incident budget, no future maintainer needing comprehension), so Ng's burden could genuinely fall; simultaneously his evidence is self-reported felt burden, the instrument class shown to lag system reality and skew rosy (The Automation–Optimism Link's no-deficit self-reports vs the measured vanished gains in Contractor & Reyes's randomized study), so "significantly" is a feeling, not a magnitude. The Faros-outranks-both tiebreak for the org case stands. Missing: any measured QA-time series for 0-to-1 builders. The consequence half gets a number, 2026-08-12 — DX's Q2 2026 panel (vendor-claim, 500+ organizations) reports AI users saving an estimated 4-6 hours per week while the innovation ratio (share of time on new features versus maintenance and overhead) stays flat. That is a partial concession to Ng and a rebuttal of what he draws from it: the hours really do come free, and at panel scale they are not landing where his account says they go — on higher-level product decisions. Note what it is not. It is a vendor's self-selected customer panel with the methodology in a gated report, it measures time allocation rather than the QA burden itself, and a flat ratio is consistent with the freed hours being consumed by the review and incident load Faros measures rather than with them never existing. The 0-to-1 scope split survives it untouched, since a solo builder has no innovation ratio to move. The hours doubled and still did not land, 2026-09-22: dx ai accelerates output not innovation (DX, vendor-claim, same customer base) updates the figure to 6.1 h/week saved in Q2 2026 from 3.0 in Q3 2025 and regresses the innovation ratio on 15 workflow metrics: AI output explains 63% of the variance in time saved but carries only a standardized β of 0.16 against the innovation ratio, in a model that explains 13% of it. That sharpens the concession-and-rebuttal above without changing it — the freed time is a larger effect than a quarter earlier, and its conversion into new-feature work is measurably weak — under the same caveats (self-report on both sides, vendor panel, methodology unpublished) and with one addition: the strongest predictor of a lower ratio is information-seeking friction (β −0.19), a context problem rather than a QA problem, so DX's own data does not say the hours went to QA either. Ng's promotion story and the whiplash mechanism are both left standing by a model that explains one-eighth of the outcome. The stated mechanism is the weak link, 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — Ng credits the relief to agents being "much more able to test their own code." On open-source agentic PRs that is mostly not what happens: agent-written tests raise coverage of the agent's own diff in only 35.9% (Java) / 22.5% (Python) of the PRs that include tests, and 50.4% of code-changing PRs carry no test change. So a felt fall in QA time is also consistent with less checking, a third reading beside scope and optimism. It is weak for Ng's setting, because in a 0-to-1 build every line is new and the diff-coverage gap may be smaller. The deciding datum is unchanged and absent from the wiki: a measured QA-time or escaped-defect series for 0-to-1 builders. Retagged to #oq/source. A portfolio that did move, at one production org, 2026-10-01: ai changed how spotify builds quality at higher velocity (Spotify, case-study, self-classified) reports merged changes doubling year over year (~8,100 → ~17,000 in August). Over the same year the quality-and-optimization share of merged PRs rose 27% → 31%, maintenance and configuration fell 31% → 25%, and feature work grew in absolute terms. Unlike DX's flat innovation ratio, the freed capacity landed somewhere measurable. It went to quality work and away from maintenance, not to the higher-level product decisions Ng describes, so it concedes the hours and still rebuts where he says they go. It counts PRs by type, not hours, in one org, so the panel reading stands.
    • SourceThe external loop is the unshortened one. Is that physics (users take time to react) or an unautomated frontier (synthetic users, deployment simulation applied to products rather than models)?
    • WaitIf the human's presence in the middle loop is justified by a context advantage that is closable, the middle loop is a transitional structure. What does a two-loop world look like — and who translates the external loop's signal then?
    • SourceWhere's the boundary of "council of LLM judges" reliability — does it hold for genuinely contested value judgments, or only for quality/coherence? Partially answered (2026-08-04) by Zhou (2026), and it inverts the question's premise. The question assumes the council is safe on the easy end and asks how far up it holds; the measurement says it fails on the easiest end — objective correctness on grade-school math — once anything optimizes against it. Three cross-family judges accepting only unanimously still pass 55% of manufactured wrong answers, and Proposition 2 shows no monotone rule over a shared plausibility signal can do better. Two further findings sharpen where the boundary actually sits. The council is fine as a static rater and fails as a reward: the same judges hold usable discrimination (0.21–0.38) before optimization and collapse to 0.05–0.17 after, so the binding variable is optimization pressure, not the contestedness of the judgment. And reference-free verdicts track prompt framing rather than correctness — with unit-test ground truth held fixed, Llama's gap@16 swings −0.106 under a strict instruction to +0.722 under a lenient one, so on the fuzzy end there may be no stable operating point to have a boundary about. And the council's headroom is small before any of that. Yang et al. (2026) measure juror error correlation on ordinary preference grading with nothing optimizing against the judges — ρ = 0.944–0.972 for repeated samples of one judge, 0.664–0.706 across a stronger family, and family-mixed juries also below independence predictions — so five jurors buy 0.463 → 0.482 on LLMBar. Condorcet's amplification requires independent voters and LLM judges are not that, optimization or no. The council was never carrying the weight the thesis assigns it; optimization pressure only makes the shortfall adversarial. What is not answered: nothing here tests contested value judgments, where there is no anchor to audit against and hence no way to run this measurement at all.
    • SourceThe "labs care" dependency is fragile: capabilities can appear or stagnate based on lab priorities you don't control. How should a product hedge against the data-distribution rug-pull?
    • SourceIs "the first model bottlenecked by my unknowns" a property of Fable or of Thariq? A frontier-lab engineer with deep model fluency hits the human-side ceiling before an average user does — which would make this a leading indicator rather than a current universal. (Bears on it without settling it, 2026-10-01: Houck (DX, practitioner-opinion) offers a third reading, that it is a property of neither. The context was always the constraint, and human colleagues silently repaired it until agents removed the repair loop. That is a second practitioner outside a frontier lab reaching the relocation claim, but it is opinion against opinion, with no measurement of when in model capability the human side starts to bind.)
    • SourceThe quiz gate is self-administered and self-graded (by the model, on the model's own work). What stops a comfortable equilibrium where the quiz gets easier as the reviewer gets lazier? Cf. the maker/checker problem in Verification as the New Bottleneck. (The drift the question worries about is now measured in one register — Security Debt of Agent-Generated Code finds humans committing 67.6% of the genuine leaked credentials inside agent PRs, read by its authors as reduced vigilance. That establishes the direction is real; it says nothing about what would stop it, so the question stands unanswered.) Partially answered from an unexpected direction (2026-08-12), and it moves the defect earlier: Greptile (case-study) measures a model's recall on high-severity bugs in code its own family authored at 6–12 points below its recall on the other family's code. The quiz gate does not need to decay to be weak — a self-authored quiz asks about what the authoring model thinks matters, and the categories it under-weights in review are correlated with the ones it under-weights in authoring, so the blind spots are missing from the question set on day one. That is a different failure from the equilibrium this bullet asks about, and it comes with the obvious mitigation attached (a different model family writes the quiz). The equilibrium question itself — what stops the quiz getting easier as the reviewer gets lazier — is still unanswered; nothing measures a self-graded gate over time. The "reviewer gets lazier" half is now measured on human review, though not on a quiz gate (2026-10-01): Yu, Liu & Zhang (empirical, 400 AIDev reviewers over 207 days) find approval of agent PRs rising gradually and linearly with each reviewer's exposure (30.5% → 36.6%, d = 0.25). Nothing in the review language flags it in advance: four lexical quality metrics stay flat, and the language shift that does exist trails the approval shift. So the drift is real in the adjacent setting, and it has the shape the question fears, a slope with no threshold to watch for. It is also confounded with the agent's code improving, and still no study measures a self-graded gate.
    • SourceElicitation has a cost. Every technique here spends a session's worth of tokens and attention on not building. Nothing in the source bounds when the blindspot pass costs more than the bug it prevents.
    • SourceIf unknown knowns are extractable, are they extractable once? Does a codified blindspot pass become a skill file that permanently narrows the gap, or does each new territory reopen it?
    • SourceFung's own open question: "How far do you push fully automated reviews?" — where's the speed/safety balance, and how do you keep humans confident without re-introducing the review bottleneck? Sharpened by Review as the Control Point: full automation reliably raises review throughput and cuts latency (its P8), but its effect on code quality and security is contested (P9), and the latency effect of a review-governance policy flips sign by calibration — a risk-tiered policy that gates only material changes lowers latency, a blanket policy raises it (P17). So "how far" has no single answer: the safe frontier is set by automated-reviewer capability and process design (two of that page's three moderators), not by a fixed dial. Partially answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — "how far" is a partition, not a dial: automate 100% of mechanical checking (style, lint, spec-drift, tests); the contested zone is automated quality/security judgment (P9, vendor claims unmeasured); and the hard limit is not defect-catching but the functions review performs besides it — reviewer-skill growth (P12/P13), collective ownership (P14), comprehension-debt paydown (P15) — which erode under automation even if the machine catches every bug. Residual: the moderator thresholds are still unmeasured. Floor datum added 2026-07-29 by Security Debt of Agent-Generated Code (empirical): on hard-coded credentials — the one smell class with purpose-built automated detection — seven distinct bots and human reviewers together commented on just 18.9% of the genuine live credentials in 4,022 agentic PRs. So the "automate mechanical checking fully" half of the partition is a prescription, not a description: where it is fully mechanizable, it currently isn't mechanized well. Deployed calibration added 2026-07-29 by Risk-Tiered Auto-Approval (case-study): PostHog's StampHog answers "how far" as as far as the cheap structural checks reach — PR state, a blast-radius deny-list, and a <500-line/<20-file ceiling gate the decision, with an LLM check demoted to a last-position veto that may tighten but never loosen — and that reaches ~1 in 3 PRs merged into their main repo (1.6K in a month). Two qualifiers keep this from being an answer: what it replaced was a Slack stamp-exchange ritual by engineers with "little to no context," so it converts an implicit rubber stamp into an explicit gate-checked one rather than automating substantive review; and the account reports volume with no false-approval or escaped-defect rate, which is the contested half of P9 left unmeasured at scale. A different kind of answer, 2026-08-12 — Tran et al. (empirical, 3.52M production changes with a human control cohort): for at least one defect class, "how far" is the wrong axis, because the human pass contributes nothing measurable to catching it at any depth. They tested review time and iteration count against the survival of inefficient AI-generated C++ and found no correlation, concluding that upstream automated intervention is necessary rather than merely faster. That reframes the partition above: it is not only "mechanical checking vs judgment," it is also which defects are legible to a reader at all — a missing move constructor inside a correct, idiomatic function is invisible to attention and trivially visible to a static category. The caveats are that the null is reported without a statistic or specification, and that a monorepo with mature static analysis has already mechanized much of what review would otherwise catch. The adoption number arrives, 2026-08-12 — Cynthia et al. (empirical, 54,713 agent review comments across 341 repos): where the partition above says "automate mechanical checking fully," this is the first population-scale reading of whether the automated layer's output is taken. It is, roughly seven times in ten (72.9% Copilot / 67.2% Cursor / 54.8% Codex), and the lever that moves it is actionability, not eloquence — an inline code suggestion is the strongest predictor (OR 1.62) while length and sheer explanation count hurt. Two things keep this from extending the partition further. The pooled model's AUC is 0.58, so comment design explains very little of the outcome and the agent-level spread survives controlling for it. And adoption is not efficacy: the study measures whether a human closed the thread, never whether a defect existed — so it fills in "does the machine's output get used" and leaves "does it catch anything" exactly where the 18.9% credential floor left it. How practitioners actually draw the line, 2026-09-22 — layered supervision ai assisted software engineering (case-study, 5 interviews) is the first source here to report the decision procedure rather than a capability estimate, and it is a ratchet: any rule relevant enough to recur in review is promoted to a lint rule permanently, the build system becomes the arbiter, and the human residue is defined negatively as what cannot be encoded (contextual interpretation, architectural tradeoff reasoning, maintainability assessment). That sharpens the partition's mechanism — the boundary moves on evidence of recurrence, never on evidence of the automated checker's accuracy, which is exactly the quantity the 18.9% credential floor says nobody is measuring. It answers nothing on the safety half: five interviews, no defect rate, no escape rate, no review-time measurement. The first outcome-flavoured signal for the automated half, and it carries no number (2026-09-22): Faros's The Speed Trap (vendor-claim) reports that teams with heavy agentic-review adoption see faster first reviews and lower change-failure rates — the first association in the corpus between automating review and a quality outcome rather than an adoption or detection rate. It is worth almost nothing as evidence and is recorded for the direction only: no percentages for either claim, explicitly labelled correlational by Faros itself, measured on a vendor's own customers with no method published, and undercut in the same post by the observation that unreviewed merges keep rising even where agentic-review adoption is high — which is the volume channel this page's partition was built to handle. It does not move the contested quality edge, and at 50–80% of PRs at many companies the automated layer is now large enough that its efficacy being unmeasured is the more striking fact. A "by whom" axis, 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — the partition says how much to automate, not who automates it. On public GitHub the automated layer is mostly the authoring product reviewing itself: same-product AI review outnumbers cross-product about 4.6:1, and 91% of attributable agent PRs draw no detectable AI review at all. That is the configuration the wiki's only recall measurement disfavours (6–12 points of high-severity recall lost on same-family code; different construct, so a tension rather than a finding). A size-keyed partition line is also a silent vendor filter: a 500-line ceiling passes the median Codex/Devin/Copilot/Cursor PR and blocks the median Claude Code PR, while the vendor gap that survives a median-size comparison (Codex vs Devin reverts) is invisible to it. The safety half still has no false-approval or escaped-defect rate for automated review at scale; ICONIQ's 20–30% false-positive "tax" is a cost datum, not an escape rate. Retagged to #oq/source: only a new source can supply that rate.
    • SourceIf CI/build is the hidden jam, does verification infrastructure (test runners, CI capacity) become the actual capex of an AI-native org? Partially answered 2026-07-29 by Agent-Generated Test Quality (empirical): it supplies the mechanism but not the cost. Agent-authored tests in AIDev carry flakiness indicators — unmocked file I/O, random, datetime.now — at a 0.44 rate vs 0.30 for human-authored tests, so the throughput increase arrives with a compounding rerun tax on the runner rather than a one-off load increase. Two gaps keep this short of an answer: the study measures candidate rate (its specified dynamic re-run stage is never reported), and nothing in it prices CI capacity, so the jam is evidenced while the capex claim is not. A CI-spend-per-merged-PR series stratified by agent authorship is what would settle it. Priced, but only by a vendor, 2026-07-29 by circleci q2 pulse 2026 (vendor-claim): CircleCI now supplies the cost half the study omitted — a countable cycles-to-merge metric (median MER 3.9 vs 1.3 for its elite cohort) and a modeled ~$900K/yr delivery cost for a 50-developer team, ~$700K of it claimed recoverable by shifting checks into the inner loop, including a "token reload penalty" for agents idling on CI. That is the first attempt in the vault to put a currency figure on the jam, and it is an argument that verification infrastructure is a real capex line. It does not settle the question: the $900K is a model over CircleCI's own customers with no published inputs, the recoverable figure is the sales case for CircleCI's inner-loop products, and it prices CI time and tokens rather than the runner-capacity build-out the question asks about. The stratified non-vendor spend series is still the thing that would settle it. The jam is measured again and still not priced, 2026-09-22 by Faros's The Speed Trap (vendor-claim): average task time in QA +300.6%, the largest deterioration of any metric in its dataset and a near-tenfold acceleration on the prior window's +33.7%. That is the strongest evidence yet that the constraint is real and worsening at the verification stage — and it is the wrong unit twice over. It measures stage time in a work tracker, not runner capacity or CI spend, so it prices nothing; and it cannot separate a capacity shortfall from a human-QA shortfall from the mechanical fact that the same report's PRs grew 71.8% larger. Two vendor instruments (CircleCI's cost model, Faros's stage time) now agree the verification layer is where the cost lands, and neither has published a capacity or spend figure. The trigger is unchanged: CI spend per merged PR, stratified by agent authorship, from a non-vendor source. A weak side-reading from a different instrument, 2026-09-22 — context, not a partial answer: dx ai accelerates output not innovation (DX, vendor-claim, 500+ customers, survey-plus-telemetry panel) regresses the innovation ratio — share of engineering effort on new capabilities versus maintenance — on 15 workflow metrics, and the drag that surfaces is information-seeking friction (standardized β −0.19, p < 0.01), not anything at the verification stage; deploy frequency carries a small positive coefficient (+0.06). If the verification jam were consuming the freed hours at the level a survey panel can see, a shipping-velocity or QA-stage metric should appear as the negative predictor, and in DX's model none does. Two reasons that is nearly weightless here: the model explains 13% of the outcome's variance, and the metrics DX names are workflow constructs with no CI-capacity or CI-spend measure among them, so the quantity this question asks about was never in the model. What it records is narrower: on the one instrument that asked developers where their effort goes, the correlate of less feature work is finding context, the same diagnosis Faros's successor reaches from the telemetry side (restarts +66.7%, read as agents lacking context) — and neither prices CI.
    • SourceHow should slice granularity be tuned? Too thin = many merge conflicts; too thick = back to horizontal. Partially answered 2026-10-01: Checkpoint-Gated Convergence — Shopify's Helix gives a practitioner rule without measurement. Start at the thinnest reviewable unit (a skeleton, then one small section) and grow slice size only once earlier slices pass review. Size is bounded by what a reviewer can judge "at a glance" and what fits a small context window. The merge-conflict half is untouched: Helix runs checkpoints serially within a screen and parallelizes only across screens, so its slices never compete for the same files.
    • ResolvedCan the planner agent be trusted to slice vertically once told to, or does it need a verifier that flags horizontal slices? Pocock's experience: it needs the verifier, at least through 4.7. Answered: Verifying Without a Compiler: Cowork's Harness vs Claude Code's, and Why the Slice Verifier Stays — it needs the verifier by design, not just empirically: slice shape is a mechanically checkable invariant (does the ticket touch schema + service + UI?), and checkable invariants belong in the deterministic checker regardless of model trust. "Slice vertically" in a prompt is a behavior request — unreliable against the training prior on the way up, compounding-prone once the behavior goes native — while the verifier is a constraint: it doesn't compound, costs ~nothing, and catches drift in either direction. Correct trajectory as models improve: prune the prompt line when ablation shows it's native; keep the checker, the way tests outlive the model learning to write correct code.
    • SourceKarpathy hints at "one domain that's very [valuable]" for founders but won't say which (didn't want to "vague-post on stage"). What verifiable RL-environment domain is he gesturing at?
    • WaitIf the mediocre/AI-native spread keeps widening, what does that do to team composition — a few extreme outliers plus agents, vs. broad mid-level staffing?