Sources#
- AI accelerates output, not innovation
- AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity
- Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think
- Thread by @AndrewYNg
Summary#
Andrew Ng's June 2026 letter in The Batch opens by observing that "loop engineering" had become a buzzphrase off the back of Boris Cherny and Peter Steinberger — the same two practitioners the wiki's Loop Engineering page is named on. Ng's contribution is to point out that they are all talking about one loop, and that building a product runs on three, nested, at different timescales. "These loops guide not just how I build software, but also how I decide what software to build." (practitioner-opinion — a taxonomy plus a personal build account, not measurement.)
The taxonomy's value to this wiki is corrective. Nearly everything in the corpus about loops — Agent Loop Pattern, Loop Engineering, /goal, Ralph loops, maker/checker sub-agents — optimizes the innermost loop, the one the agent closes by itself. Ng's outer two loops are where the human still is, and they are the loops that decide whether the thing being built is worth building.
The three loops#
| Loop | Who closes it | Cadence | What it consumes | What it produces |
|---|---|---|---|---|
| Agentic coding loop | the agent, alone | "every few minutes" | a spec, optionally evals | code that passes its own tests |
| Developer feedback loop | the human | "tens of minutes to hours" | a working build | a revised spec / steer |
| External feedback loop | the market | "hours… days or even weeks" | a shipped thing | a revised vision |
The nesting is strict: the external loop informs the developer's vision, the vision drives the spec, the spec drives the coding agent. Feedback flows back up the same chain. Each outer loop runs ~1–2 orders of magnitude slower than the one it contains, which is what makes the inner loop's speedup so consequential and so limited — you can make the innermost loop instant and the product still moves at the speed of the outermost one.
1. The agentic coding loop#
"Given a product specification and optionally a set of evals… have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification." Ng dates the closing of this loop to "around the end of last year" and calls it "a game changer in enabling coding agents to work longer productively without human intervention." His own datum: building a typing-practice app for his daughter, "my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me." (An anecdote, not a measurement; compare Task Time-Horizon Scaling for the curve this is one point on.)
"This is an active area of invention!" — which is Loop Engineering and Agent Loop Pattern in their entirety.
2. The developer feedback loop#
The loop that changed the most, and the reason to read the piece. Ng's account of what the human used to do:
"Last year, a lot of developers (including me) were acting as the QA function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions."
The human didn't get removed from the loop; they got promoted out of QA. Notice this cuts against the wiki's dominant framing. Verification as the New Bottleneck holds that once coding is cheap, verification becomes the scarce resource; Ng reports the opposite motion — self-testing agents drained the human verification burden and freed attention upward. The two are reconcilable (Ng builds 0-to-1 personal products, where a wrong build is cheap; Fung's claim is about production orgs, where it isn't), but the disagreement is real and worth keeping visible rather than averaging away. Faros's telemetry — median time-in-PR-review up 441.5% — is the higher-evidence source and sides against Ng for the org case.
Two mechanisms Ng names inside this loop:
- Spec translation is the work. "When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec." Vision → spec is lossy and iterative; seeing the build is how you find out what the spec should have said. This is the unknowns problem stated as a loop rather than as a technique.
- Evals are what you build when the loop fails the same way twice. "If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful." A discipline of provocation, not prophylaxis — the opposite ordering to "ten great evals" written as the spec. Ng's version is cheaper and lazier; Cat Wu's is what you do once the feature is ambiguous enough that "it failed" isn't self-evident.
3. The external feedback loop#
"Asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing." Slow — "rarely taking less than hours and sometimes taking days or even weeks." This loop is the only one that updates the vision rather than the spec, and it is the one the AI-native tooling has done least to accelerate. Ng notes AI-native teams increasingly automate its inputs (usage-data analysis, feedback summarization, competitive analysis) without shortening the loop itself.
Why the human is in the middle loop#
Ng's account of what the human contributes is his most-quoted line, and it gets its own page: not taste but a context advantage. "For pretty much all the products I'm involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in." The stopping condition follows directly: "So long as the human knows something the AI does not, human-in-the-loop is needed to inject that knowledge into the system."
Read against the taxonomy, this says something specific: the human sits in the developer feedback loop because that is where knowledge from the external loop (which only the human has run) gets injected into the coding loop (which the agent runs alone). The human is the transmission between the loop that knows the users and the loop that writes the code.
The role consequence#
"With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!"
This is Engineer PM Convergence from a third independent vantage (after Cat Wu and Boris Cherny), and Ng identifies the specific failure it produces: engineers newly holding the middle loop over-invest in it. Building is the loop they know how to run, so the external loop — slow, unpleasant, unautomatable — gets skipped. Ng closes symmetrically: "engineers are playing an expanded role (just as product managers and designers now do more engineering)."
Tension: is the loop already obsolete?#
Two days before Ng's letter, Andrew Ambrosino — who leads the Codex desktop app at OpenAI — told Lenny's Podcast that "loops are so last week," arguing that orchestrated loops are a transitional harness that autonomous, long-horizon models are already outgrowing (see Vibe Coding vs. Agentic Engineering). Ng is publishing a loop taxonomy in the same week.
They may both be right, because they mean different loops. Ambrosino's "loops" are the agentic coding loop — the scaffolding that pokes an agent until it converges — and his claim is that model capability absorbs it (Harness Shrinkage as Models Improve). Ng's outer two loops are not harness; they are the structure of product development itself, and no model capability dissolves the fact that shipping to users takes days. Which is the useful reading of the taxonomy: the inner loop is a harness and will shrink; the outer loops are physics and won't.
The loops, two months on (August 2026)#
Ng's general-audience restatement (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion) compresses the taxonomy into a slogan — "learn AI, build fast, and talk to customers" — and confirms the reading above that the outer loop is physics. The inner loop is now a weekend habit ("I find myself building things… every week, every weekend because I or someone on our team… have some problem and I have some idea for building some AI thing"; last weekend it was a frontier model analyzing his business metrics because "I didn't have time to go find a data scientist"). The outer loop is what he says still takes "months… maybe years": "deep customer insight… talking to people reading facial expressions surveys doing that over and over until we figure out what to build." His name for the middle is now explicit — "the product management bottleneck" — and its content is the context advantage again, now called "a long-term advantage." Two smaller updates. The typing app that anchored the June letter's one-hour agentic run turns out to have a purpose that bears on the learning debate: he built it because he "did not like any of the… free online learning to type types of things," to have his daughter learn a skill herself — cheap implementation used to make a non-offloading tool for a child, in the same interview where he calls LLMs terrible for learning (Experimental Learning Impact of Generative AI). And the KPI question gets the outer-loop answer: "the business outcome of AI is more a function of the business than a function of the AI," so the only KPIs are business KPIs — there is no inner-loop metric for whether the thing was worth building.
Connections#
- Loop Engineering — the discipline Ng is placing: Osmani's five primitives all live inside the innermost of these three loops; this taxonomy is the map that shows what loop engineering leaves untouched
- Agent Loop Pattern — the primitive that closes the agentic coding loop
- Context Advantage, Not Taste — Ng's reframing of the human contribution, and the reason the human occupies the middle loop specifically
- Unknowns as the Agentic Bottleneck — vision→spec lossiness is the unknowns problem; the developer feedback loop is where unknowns surface after implementation
- Engineer PM Convergence — third independent report of the same convergence, with a named failure mode: engineers over-run the loop they enjoy
- Evals as Product Spec — the productive disagreement: evals as a reaction to repeated failure (Ng) vs evals as the spec authored up front (Cat Wu)
- Vibe Coding vs. Agentic Engineering — Andrew Ambrosino's "loops are so last week," published the same week; the tension resolves once you separate harness-loops from product-loops
- Agent-Generated Test Quality — the first
empiricallook at the artifact Ng's claim rests on. Agent-authored tests really are broader than human ones (edge-case variety 0.62 vs 0.32), which supports the self-testing half of his account — but they carry unmocked file I/O and non-determinism at ~1.4× the human rate, so the QA burden they drain from the developer is partly relocated onto the CI runner rather than eliminated - Verification as the New Bottleneck — the direct disagreement: Ng reports self-testing agents reduced the human QA burden, where Fiona Fung holds verification became the scarce resource; scope (0-to-1 personal builds vs production orgs) is the likely reconciler, and Faros's telemetry outranks both
- Task Time-Horizon Scaling — the ~1-hour unattended run is one anecdotal point on this curve
- Harness Shrinkage as Models Improve — why the inner loop shrinks and the outer two don't
- AI Native Product Cadence — the org-level cadence the outer loops set; the external loop is the floor no tooling has lifted
- Andrew Ng — author
- Boris Cherny / Peter Steinberger — the two practitioners Ng credits with making "loop engineering" a buzzphrase
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — dissolves the Ng-vs-Faros either/or: the scope split explains why the QA burden could truly fall in 0-to-1 builds, the documented optimism bias explains why the self-report can't quantify it
- Implementation Abundance Inverts Product Work — the same bottleneck named from the lab side; Ng's "product management bottleneck" is the outer two loops, and his limit on the inversion is that they still take months
- What the Agent-PR Oversight Numbers Can and Cannot Say — Ng's QA-relief claim tested against its own stated mechanism: diff-coverage data says agents mostly do not test their own changes, so felt relief may be less checking rather than less need
Open Questions#
- Ng asserts the developer's QA burden fell "significantly." Faros's 2026 telemetry measures the opposite for production orgs. Is the split really 0-to-1-vs-production, or is Ng's self-report subject to the same optimism bias the survey literature keeps finding? Partially answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — both-and, not either/or. The scope split is real and does most of the work (0-to-1 builds lack the review functions that make verification expensive in production: no queue, no incident budget, no future maintainer needing comprehension), so Ng's burden could genuinely fall; simultaneously his evidence is self-reported felt burden, the instrument class shown to lag system reality and skew rosy (The Automation–Optimism Link's no-deficit self-reports vs the measured vanished gains in Contractor & Reyes's randomized study), so "significantly" is a feeling, not a magnitude. The Faros-outranks-both tiebreak for the org case stands. Missing: any measured QA-time series for 0-to-1 builders. The consequence half gets a number, 2026-08-12 — DX's Q2 2026 panel (
vendor-claim, 500+ organizations) reports AI users saving an estimated 4-6 hours per week while the innovation ratio (share of time on new features versus maintenance and overhead) stays flat. That is a partial concession to Ng and a rebuttal of what he draws from it: the hours really do come free, and at panel scale they are not landing where his account says they go — on higher-level product decisions. Note what it is not. It is a vendor's self-selected customer panel with the methodology in a gated report, it measures time allocation rather than the QA burden itself, and a flat ratio is consistent with the freed hours being consumed by the review and incident load Faros measures rather than with them never existing. The 0-to-1 scope split survives it untouched, since a solo builder has no innovation ratio to move. The hours doubled and still did not land, 2026-09-22: AI accelerates output, not innovation (DX,vendor-claim, same customer base) updates the figure to 6.1 h/week saved in Q2 2026 from 3.0 in Q3 2025 and regresses the innovation ratio on 15 workflow metrics: AI output explains 63% of the variance in time saved but carries only a standardized β of 0.16 against the innovation ratio, in a model that explains 13% of it. That sharpens the concession-and-rebuttal above without changing it — the freed time is a larger effect than a quarter earlier, and its conversion into new-feature work is measurably weak — under the same caveats (self-report on both sides, vendor panel, methodology unpublished) and with one addition: the strongest predictor of a lower ratio is information-seeking friction (β −0.19), a context problem rather than a QA problem, so DX's own data does not say the hours went to QA either. Ng's promotion story and the whiplash mechanism are both left standing by a model that explains one-eighth of the outcome. The stated mechanism is the weak link, 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — Ng credits the relief to agents being "much more able to test their own code." On open-source agentic PRs that is mostly not what happens: agent-written tests raise coverage of the agent's own diff in only 35.9% (Java) / 22.5% (Python) of the PRs that include tests, and 50.4% of code-changing PRs carry no test change. So a felt fall in QA time is also consistent with less checking, a third reading beside scope and optimism. It is weak for Ng's setting, because in a 0-to-1 build every line is new and the diff-coverage gap may be smaller. The deciding datum is unchanged and absent from the wiki: a measured QA-time or escaped-defect series for 0-to-1 builders. Retagged to#oq/source. A portfolio that did move, at one production org, 2026-10-01: AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity (Spotify,case-study, self-classified) reports merged changes doubling year over year (~8,100 → ~17,000 in August). Over the same year the quality-and-optimization share of merged PRs rose 27% → 31%, maintenance and configuration fell 31% → 25%, and feature work grew in absolute terms. Unlike DX's flat innovation ratio, the freed capacity landed somewhere measurable. It went to quality work and away from maintenance, not to the higher-level product decisions Ng describes, so it concedes the hours and still rebuts where he says they go. It counts PRs by type, not hours, in one org, so the panel reading stands. - The external loop is the unshortened one. Is that physics (users take time to react) or an unautomated frontier (synthetic users, deployment simulation applied to products rather than models)?
- If the human's presence in the middle loop is justified by a context advantage that is closable, the middle loop is a transitional structure. What does a two-loop world look like — and who translates the external loop's signal then?
Sources#
- Thread by @AndrewYNg — Andrew Ng, The Batch, published 2026-06-30, clipped from X (
practitioner-opinion). The three-loop diagram is an image hosted on X and not transcribed. - AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity — Tyson Singer, Spotify Engineering, 2026-09-16,
case-study: the merged-PR work-type mix only (cited in the first open question). Full treatment at Acceleration Whiplash - Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28,
practitioner-opinion): learn-AI/build-fast/talk-to-customers, the weekend-build habit, the months-to-years outer loop, the typing app's purpose, and the business-not-AI KPI answer - AI accelerates output, not innovation — Grace Fu, AI accelerates output, not innovation (DX newsletter, 2026-09-09;
vendor-claim). Cited only in the first open question: time saved 3.0 → 6.1 h/week, AI output β 0.16 against the innovation ratio, information-seeking β −0.19, model R² 0.13. Paraphrased-digest raw; numbers preserved, prose not quoted
Cited by 23
- Acceleration Whiplash×3
Three Loops Of Ai Native Building — the telemetry that outranks Andrew Ng's self-report: he claims…
- Andrew Ng×3
Three Loops Of Ai Native Building — the agentic coding loop (agent-closed, minutes), the developer…
- Open Questions Backlog×3
Three Loops Of Ai Native Building (88d) — The external loop is the unshortened one. Is that physics…
- What the Agent-PR Oversight Numbers Can and Cannot Say×2
Three Loops Of Ai Native Building: is Ng's "QA burden fell significantly" a 0-to-1-vs-production…
- Context Advantage, Not Taste×2
In a June 2026 letter otherwise devoted to a loop taxonomy, Andrew Ng makes an aside that quietly…
- Engineer PM Convergence×2
Ng adds what the Anthropic accounts don't: a named failure mode. Engineers newly holding the…
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?×2
Three Loops Of Ai Native Building — is the Ng-vs-Faros QA-burden split really 0-to-1-vs-production,…
- Implementation Abundance Inverts Product Work×2
Andrew Ng gives the bottleneck a name from outside the labs — "the product management bottleneck" —…
- Loop Engineering×2
Two weeks after Osmani's essay, Andrew Ng responded to loop engineering "becoming a hot buzzphrase…
- Agent-Generated Test Quality
Acceleration Whiplash — the authoring-quality thesis measured on the test suite: assertion drift…
- Agent Loop Pattern
Three Loops Of Ai Native Building — Andrew Ng's taxonomy places this primitive: it closes the…
- Andrew Ambrosino
Three Loops Of Ai Native Building — Andrew Ng published a three-loop taxonomy the same week…
- The Automation–Optimism Link
Three Loops Of Ai Native Building — a candidate instance of the gap: Andrew Ng self-reports that…
- Boris Cherny
Three Loops Of Ai Native Building — Andrew Ng credits him (with Peter Steinberger) for making "loop…
- Deployment Simulation
Three Loops Of Ai Native Building — the open frontier the taxonomy exposes: the external feedback…
- Evals as Product Spec
Three Loops Of Ai Native Building — the productive disagreement on when to write evals: Andrew Ng…
- Experimental Learning Impact of Generative AI
Three Loops Of Ai Native Building — the builder's side of the same author's learning claim: Ng's…
- AI Coding Practice
Three Loops Of Ai Native Building — Andrew Ng's nested-loop taxonomy for 0-to-1 products: the…
- Peter Steinberger
Three Loops Of Ai Native Building — Andrew Ng credits him (with Boris Cherny) for the buzzphrase,…
- Task Time-Horizon Scaling
Three Loops Of Ai Native Building — one anecdotal point on the curve: Andrew Ng's coding agent…
- Unknowns as the Agentic Bottleneck
Three Loops Of Ai Native Building — vision→spec lossiness is this problem stated as a loop: Andrew…
- Verification as the New Bottleneck
Three Loops Of Ai Native Building — the direct dissent. Andrew Ng reports the opposite motion:…
- Vibe Coding vs. Agentic Engineering
Three Loops Of Ai Native Building — Andrew Ng published a loop taxonomy the same week Ambrosino…
Related articles
- Andrew Ng
Founder of DeepLearning.AI and AI Fund, founding lead of Google Brain, co-founder of Coursera; writes The Batch, where…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Outsource Your Thinking, Not Your Understanding
"You can outsource your thinking but not your understanding"; understanding as the non-delegable human bottleneck; know…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Boris Cherny
Creator of Claude Code at Anthropic; phone-driven workflow with hundreds of agents; primary advocate of `/loop` primiti…
