H
Howardism
Plate IIAI Coding Practice中文HOWARDISM

The Three Loops of AI-Native Building

Andrew Ng's nested-loop taxonomy for 0-to-1 products: the agentic coding loop (minutes, agent-closed), the developer feedback loop (tens of minutes to hours, human-closed), and the external feedback loop (hours to weeks, market-closed); loop engineering has been optimizing only the innermost one, and the human's remaining job is a context transfer that lives in the outer two

Article metadata
Publication details
Published:July 9, 2026
Filed:Concept
Domain:AI Coding Practice
Reading:19 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for The Three Loops of AI-Native Building

Sources#

Summary#

Andrew Ng's June 2026 letter in The Batch opens by observing that "loop engineering" had become a buzzphrase off the back of Boris Cherny and Peter Steinberger — the same two practitioners the wiki's Loop Engineering page is named on. Ng's contribution is to point out that they are all talking about one loop, and that building a product runs on three, nested, at different timescales. "These loops guide not just how I build software, but also how I decide what software to build." (practitioner-opinion — a taxonomy plus a personal build account, not measurement.)

The taxonomy's value to this wiki is corrective. Nearly everything in the corpus about loops — Agent Loop Pattern, Loop Engineering, /goal, Ralph loops, maker/checker sub-agents — optimizes the innermost loop, the one the agent closes by itself. Ng's outer two loops are where the human still is, and they are the loops that decide whether the thing being built is worth building.

The three loops#

LoopWho closes itCadenceWhat it consumesWhat it produces
Agentic coding loopthe agent, alone"every few minutes"a spec, optionally evalscode that passes its own tests
Developer feedback loopthe human"tens of minutes to hours"a working builda revised spec / steer
External feedback loopthe market"hours… days or even weeks"a shipped thinga revised vision

The nesting is strict: the external loop informs the developer's vision, the vision drives the spec, the spec drives the coding agent. Feedback flows back up the same chain. Each outer loop runs ~1–2 orders of magnitude slower than the one it contains, which is what makes the inner loop's speedup so consequential and so limited — you can make the innermost loop instant and the product still moves at the speed of the outermost one.

1. The agentic coding loop#

"Given a product specification and optionally a set of evals… have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification." Ng dates the closing of this loop to "around the end of last year" and calls it "a game changer in enabling coding agents to work longer productively without human intervention." His own datum: building a typing-practice app for his daughter, "my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me." (An anecdote, not a measurement; compare Task Time-Horizon Scaling for the curve this is one point on.)

"This is an active area of invention!" — which is Loop Engineering and Agent Loop Pattern in their entirety.

2. The developer feedback loop#

The loop that changed the most, and the reason to read the piece. Ng's account of what the human used to do:

"Last year, a lot of developers (including me) were acting as the QA function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions."

The human didn't get removed from the loop; they got promoted out of QA. Notice this cuts against the wiki's dominant framing. Verification as the New Bottleneck holds that once coding is cheap, verification becomes the scarce resource; Ng reports the opposite motion — self-testing agents drained the human verification burden and freed attention upward. The two are reconcilable (Ng builds 0-to-1 personal products, where a wrong build is cheap; Fung's claim is about production orgs, where it isn't), but the disagreement is real and worth keeping visible rather than averaging away. Faros's telemetry — median time-in-PR-review up 441.5% — is the higher-evidence source and sides against Ng for the org case.

Two mechanisms Ng names inside this loop:

  • Spec translation is the work. "When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec." Vision → spec is lossy and iterative; seeing the build is how you find out what the spec should have said. This is the unknowns problem stated as a loop rather than as a technique.
  • Evals are what you build when the loop fails the same way twice. "If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful." A discipline of provocation, not prophylaxis — the opposite ordering to "ten great evals" written as the spec. Ng's version is cheaper and lazier; Cat Wu's is what you do once the feature is ambiguous enough that "it failed" isn't self-evident.

3. The external feedback loop#

"Asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing." Slow — "rarely taking less than hours and sometimes taking days or even weeks." This loop is the only one that updates the vision rather than the spec, and it is the one the AI-native tooling has done least to accelerate. Ng notes AI-native teams increasingly automate its inputs (usage-data analysis, feedback summarization, competitive analysis) without shortening the loop itself.

Why the human is in the middle loop#

Ng's account of what the human contributes is his most-quoted line, and it gets its own page: not taste but a context advantage. "For pretty much all the products I'm involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in." The stopping condition follows directly: "So long as the human knows something the AI does not, human-in-the-loop is needed to inject that knowledge into the system."

Read against the taxonomy, this says something specific: the human sits in the developer feedback loop because that is where knowledge from the external loop (which only the human has run) gets injected into the coding loop (which the agent runs alone). The human is the transmission between the loop that knows the users and the loop that writes the code.

The role consequence#

"With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!"

This is Engineer PM Convergence from a third independent vantage (after Cat Wu and Boris Cherny), and Ng identifies the specific failure it produces: engineers newly holding the middle loop over-invest in it. Building is the loop they know how to run, so the external loop — slow, unpleasant, unautomatable — gets skipped. Ng closes symmetrically: "engineers are playing an expanded role (just as product managers and designers now do more engineering)."

Tension: is the loop already obsolete?#

Two days before Ng's letter, Andrew Ambrosino — who leads the Codex desktop app at OpenAI — told Lenny's Podcast that "loops are so last week," arguing that orchestrated loops are a transitional harness that autonomous, long-horizon models are already outgrowing (see Vibe Coding vs. Agentic Engineering). Ng is publishing a loop taxonomy in the same week.

They may both be right, because they mean different loops. Ambrosino's "loops" are the agentic coding loop — the scaffolding that pokes an agent until it converges — and his claim is that model capability absorbs it (Harness Shrinkage as Models Improve). Ng's outer two loops are not harness; they are the structure of product development itself, and no model capability dissolves the fact that shipping to users takes days. Which is the useful reading of the taxonomy: the inner loop is a harness and will shrink; the outer loops are physics and won't.

The loops, two months on (August 2026)#

Ng's general-audience restatement (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion) compresses the taxonomy into a slogan — "learn AI, build fast, and talk to customers" — and confirms the reading above that the outer loop is physics. The inner loop is now a weekend habit ("I find myself building things… every week, every weekend because I or someone on our team… have some problem and I have some idea for building some AI thing"; last weekend it was a frontier model analyzing his business metrics because "I didn't have time to go find a data scientist"). The outer loop is what he says still takes "months… maybe years": "deep customer insight… talking to people reading facial expressions surveys doing that over and over until we figure out what to build." His name for the middle is now explicit — "the product management bottleneck" — and its content is the context advantage again, now called "a long-term advantage." Two smaller updates. The typing app that anchored the June letter's one-hour agentic run turns out to have a purpose that bears on the learning debate: he built it because he "did not like any of the… free online learning to type types of things," to have his daughter learn a skill herself — cheap implementation used to make a non-offloading tool for a child, in the same interview where he calls LLMs terrible for learning (Experimental Learning Impact of Generative AI). And the KPI question gets the outer-loop answer: "the business outcome of AI is more a function of the business than a function of the AI," so the only KPIs are business KPIs — there is no inner-loop metric for whether the thing was worth building.

Connections#

Open Questions#

  • Ng asserts the developer's QA burden fell "significantly." Faros's 2026 telemetry measures the opposite for production orgs. Is the split really 0-to-1-vs-production, or is Ng's self-report subject to the same optimism bias the survey literature keeps finding? Partially answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — both-and, not either/or. The scope split is real and does most of the work (0-to-1 builds lack the review functions that make verification expensive in production: no queue, no incident budget, no future maintainer needing comprehension), so Ng's burden could genuinely fall; simultaneously his evidence is self-reported felt burden, the instrument class shown to lag system reality and skew rosy (The Automation–Optimism Link's no-deficit self-reports vs the measured vanished gains in Contractor & Reyes's randomized study), so "significantly" is a feeling, not a magnitude. The Faros-outranks-both tiebreak for the org case stands. Missing: any measured QA-time series for 0-to-1 builders. The consequence half gets a number, 2026-08-12 — DX's Q2 2026 panel (vendor-claim, 500+ organizations) reports AI users saving an estimated 4-6 hours per week while the innovation ratio (share of time on new features versus maintenance and overhead) stays flat. That is a partial concession to Ng and a rebuttal of what he draws from it: the hours really do come free, and at panel scale they are not landing where his account says they go — on higher-level product decisions. Note what it is not. It is a vendor's self-selected customer panel with the methodology in a gated report, it measures time allocation rather than the QA burden itself, and a flat ratio is consistent with the freed hours being consumed by the review and incident load Faros measures rather than with them never existing. The 0-to-1 scope split survives it untouched, since a solo builder has no innovation ratio to move. The hours doubled and still did not land, 2026-09-22: AI accelerates output, not innovation (DX, vendor-claim, same customer base) updates the figure to 6.1 h/week saved in Q2 2026 from 3.0 in Q3 2025 and regresses the innovation ratio on 15 workflow metrics: AI output explains 63% of the variance in time saved but carries only a standardized β of 0.16 against the innovation ratio, in a model that explains 13% of it. That sharpens the concession-and-rebuttal above without changing it — the freed time is a larger effect than a quarter earlier, and its conversion into new-feature work is measurably weak — under the same caveats (self-report on both sides, vendor panel, methodology unpublished) and with one addition: the strongest predictor of a lower ratio is information-seeking friction (β −0.19), a context problem rather than a QA problem, so DX's own data does not say the hours went to QA either. Ng's promotion story and the whiplash mechanism are both left standing by a model that explains one-eighth of the outcome. The stated mechanism is the weak link, 2026-10-01: What the Agent-PR Oversight Numbers Can and Cannot Say — Ng credits the relief to agents being "much more able to test their own code." On open-source agentic PRs that is mostly not what happens: agent-written tests raise coverage of the agent's own diff in only 35.9% (Java) / 22.5% (Python) of the PRs that include tests, and 50.4% of code-changing PRs carry no test change. So a felt fall in QA time is also consistent with less checking, a third reading beside scope and optimism. It is weak for Ng's setting, because in a 0-to-1 build every line is new and the diff-coverage gap may be smaller. The deciding datum is unchanged and absent from the wiki: a measured QA-time or escaped-defect series for 0-to-1 builders. Retagged to #oq/source. A portfolio that did move, at one production org, 2026-10-01: AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity (Spotify, case-study, self-classified) reports merged changes doubling year over year (~8,100 → ~17,000 in August). Over the same year the quality-and-optimization share of merged PRs rose 27% → 31%, maintenance and configuration fell 31% → 25%, and feature work grew in absolute terms. Unlike DX's flat innovation ratio, the freed capacity landed somewhere measurable. It went to quality work and away from maintenance, not to the higher-level product decisions Ng describes, so it concedes the hours and still rebuts where he says they go. It counts PRs by type, not hours, in one org, so the panel reading stands.
  • The external loop is the unshortened one. Is that physics (users take time to react) or an unautomated frontier (synthetic users, deployment simulation applied to products rather than models)?
  • If the human's presence in the middle loop is justified by a context advantage that is closable, the middle loop is a transitional structure. What does a two-loop world look like — and who translates the external loop's signal then?

Sources#

§ end
Cited by 23
Related articles
  • Andrew Ng

    Founder of DeepLearning.AI and AI Fund, founding lead of Google Brain, co-founder of Coursera; writes The Batch, where…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Outsource Your Thinking, Not Your Understanding

    "You can outsource your thinking but not your understanding"; understanding as the non-delegable human bottleneck; know…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Boris Cherny

    Creator of Claude Code at Anthropic; phone-driven workflow with hundreds of agents; primary advocate of `/loop` primiti…