Sources#
- Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think
- Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Thread by @AndrewYNg
Summary#
In a June 2026 letter otherwise devoted to a loop taxonomy, Andrew Ng makes an aside that quietly reframes the wiki's most-asked question:
"Many people describe this human contribution as 'taste,' but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can't be automated: so long as the human knows something the AI does not, human-in-the-loop is needed to inject that knowledge into the system."
The wiki has spent a whole page asking whether taste is a genuine ceiling or "just another AI capability that AI systems fail at for a time, then get good at." Ng's answer is: neither — it's a category error. What people call taste is not a capability the human has and the model lacks; it is knowledge the human has and the model hasn't been given. "We know a lot more than the AI system about the users and the context the product has to operate in."
The reframe is not cosmetic. It changes what kind of thing the human role is, and therefore what would end it. (practitioner-opinion — Ng offers no measurement, and states the preference as a preference.)
Why the reframe is load-bearing#
Three consequences follow immediately, and each one flips the sign of a claim elsewhere in this wiki.
1. It makes the human role falsifiable. "Taste" is unfalsifiable by construction — you cannot check whether a model has it, only whether experts like its output. A context advantage has a truth condition: does the human hold information the model doesn't? You can enumerate that information, and you can watch the list shrink.
2. It makes the role a gap, not a moat. Ng's own justification for preferring the frame is instrumental: it "gives us a clearer path to helping AI systems get better." A moat is something you defend; a gap is something you close. Read literally, the sentence "so long as the human knows something the AI does not, human-in-the-loop is needed" states the human's necessity and its expiry condition in the same breath. It is a strictly weaker claim about human durability than the taste framing it replaces, and Ng seems to intend it that way.
3. It relocates the work from cultivation to transfer. If the residue is taste, the response is to develop taste. If the residue is a context asymmetry, the response is elicitation and transfer — get the knowledge out of the human's head and into the system. Which is precisely what Thariq Shihipar's field guide, published four days later and citing nothing of Ng's, is a manual for.
The convergence: unknown knowns are the context advantage#
Thariq's four-quadrant breakdown names a cell he calls unknown knowns: "what's so obvious I'd never write it down, but would recognize it if I saw it."
That is the context advantage seen from inside the human's head. It is invisible to its holder for the same reason it is invisible to the model — nobody writes down the obvious. Ng says the human's value is knowing things the model doesn't; Thariq says the hard part is that the human doesn't know which things those are. Put together:
- Ng supplies the criterion. The human is needed exactly as long as the asymmetry exists.
- Thariq supplies the protocol. Blindspot passes, interviews, prototypes, and references are all instruments for pumping information across the asymmetry — before, during, and after the work.
Two practitioners, one week apart, no cross-citation, describing the same object from opposite ends. This is the strongest cross-source convergence in the corpus on what the human is actually for.
The uncomfortable implication#
If the human's necessity is an information asymmetry, then every artifact that transfers context narrows it. [[agent-context-files|CLAUDE.md and AGENTS.md]], skills, memory files, codebase indexes, production-sourced evals — the entire externalized-context stack exists to move what the human knows into where the model can read it. Under the taste framing these tools amplify the human. Under Ng's framing they spend the human's advantage, deliberately, one file at a time.
This is not an argument against writing them. It is an observation that the wiki's two most-recommended practices — externalize your context, and treat human judgment as the durable residue — are in tension, and Ng's frame is what makes the tension visible. The practices are consistent only if the asymmetry regenerates faster than it is transferred: new users, new markets, new products, new constraints. Whether it does is an empirical question nobody in the corpus has posed, let alone answered.
Where the reframe is too clean#
Ng's frame explains the deployment asymmetry well and the generative one badly.
- Deployment asymmetry (Ng is right). "We know a lot more about the users and the context the product has to operate in." This is private information. It is transferable in principle, and it is exactly what Anthropic's 400K-session study measures: domain understanding of the problem — not coding skill — predicts who gets quality work out of an agent. Anthropic reads a decrease in the returns to expertise over time as evidence that "models are starting to supply the essential judgment users currently bring." Under Ng's frame, that is not a mystery: it is the context gap closing, and the returns-to-expertise curve is the instrument for watching it close.
- Generative asymmetry (Ng doesn't address it). DeepMind's argument is that AI may be bounded by human conceptual frameworks — not lacking a fact, but unable to originate the concept in which the fact would be stated. Ditto Boden level-3 creativity: creating a new conceptual space is not an information deficit that better context transfer repairs. If a human's contribution is that they can invent a frame the model cannot, "knows something the AI does not" is technically true and completely misleading.
So the honest statement is that Ng dissolves the product-work version of the taste question and leaves the research-frontier version standing. Which is consistent with the domain he's writing from: 0-to-1 consumer products, where the missing knowledge really is "what my daughter wants from a typing app," not a new conceptual space.
A third arrival, from the side that calls it taste#
The convergence above is Ng-and-Thariq, both writing about the mechanism. A month later Akshay Nathan (OpenAI, productivity engineering) supplies the same diagnosis from the other camp — someone who uses the word "taste" for the bottleneck and then, unprompted, explains it as an information problem (Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI, practitioner-opinion):
"The one automation that I would love to work and it doesn't work is bring me new ideas. Somehow LLMs are just not it. One interesting part about ideas is, like, they're not, like, in a vacuum… they usually come from somewhere and, like, in product development, they're coming from talking to users or reacting to friction that you're seeing or feedback, building on some foundation that you already had planned out before."
This is a negative capability observation with a positive context explanation attached: the automation fails, and the reason offered is not that the model lacks an idea-generating faculty but that it lacks the grounding stream ideas condense out of. That is Ng's criterion applied to the one task the taste framing would call irreducibly human. It is also a clean instance of the third open question below — the failure Nathan describes is one that better context plumbing (continuous user-conversation and friction telemetry into the model) would attack, not a better model. Note the limit: this is idea generation in product work, which is the deployment asymmetry, not the generative one — nothing here touches The Abstraction Barrier.
The author, two months later: "a long-term advantage"#
Ng's own restatement, to a general audience in August 2026 (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion), keeps the mechanism and changes its expiry. The mechanism is intact and gets its best worked example: "for AI as data scientists or AI brainstorming partner it often comes up with… one or two good ideas two or three mediocre ones and… four atrocious ones and sometimes you wonder how could my AI have thought… that could even be a plausible idea" — and the atrocious ideas are ones "incredibly obvious to you" were awful, from years of "we talked to customer we saw the funny facial expression… our manager said hey… I really care about this." He is explicit that this is what taste is: "the technical thing that underlies why humans have better judgment and better taste than AI is this context advantage."
What changed is the durability. In June the frame was preferred because it "gives us a clearer path to helping AI systems get better" — the gap reading this page is built on. In August: "almost all humans… just know a lot of stuff that the plumbing does not exist and I don't think exists for the foreseeable future for AI to get"; "this is a long-term advantage like no one's going to solve this… in a few years"; and the frame is wired directly into his no-job-apocalypse argument — "one of the reasons why AI will not replace our jobs… anytime soon is because humans have a massive context advantage." Same author, same tier, no new evidence, so nothing here is superseded; it is a within-author drift and the page carries both readings. Read charitably the two are consistent: the gap is closable in principle (June) and nobody is building the plumbing in practice (August) — which is the third open question below, answered by assertion. Read less charitably, "long-term advantage" is the moat reading this page argued the frame ruled out, now held by the frame's author. The tell is the word plumbing: he has moved from a claim about information to a claim about infrastructure, and the infrastructure is exactly what Agent Context Files and Production-Sourced Evaluation are.
Against it: taste as something other than information#
The strongest counter is that some of what "taste" names is not knowledge at all but discrimination under uncertainty — a reliable preference ordering over options none of which you can articulate a criterion for. Thariq's color-grading episode is the crisp case: he had the information (Claude could generate variations) and still could not proceed, because he couldn't grade the options. His fix was to acquire the criterion — which is Ng's frame, one level up: the missing thing was information about what good looks like, and it was obtainable.
That the hardest example in the corpus resolves in Ng's favor is a point for the reframe. Design is where it will be tested: if design taste is a context gap, it closes; if it is discrimination without a statable criterion, it doesn't.
Connections#
- GDPval Benchmark — the asymmetry run as an experiment on paid work, and since 2026-09-10 with the primary paper's numbers rather than a lecturer's. GDPval writes all the context a professional averaging 14 years in the occupation carries in their head into the prompt; remove it again — prompts rewritten to omit where the data lives, how to approach the problem and what the deliverable should look like, down to 42% of the original token length — and GPT-5-high's win-or-tie rate moves 47.7% → 44.3% (wins-only 43.3% → 39.8%). A 3.4-point metric move, and the paper is explicit that the qualitative failure is larger than the metric: "the models struggled to figure out context." Two cautions for anyone quoting it: the arm ran on an earlier version of the gold subset, so its 47.7% baseline is not the 38.8% in the main text and the two must never be differenced; and a small drop is genuinely ambiguous evidence here — it is consistent with the asymmetry mattering little, and with the models failing at a task the metric barely scores. The lecturer's gloss is still this page's thesis in occupational form: "humans are basically architecting what is the set of problems and then the model can go solve it"
- Task Time-Horizon Scaling — the same asymmetry priced in duration: on internal pull requests, contractors with no codebase context run 5–18× slower than maintainers, and model performance tracks the contractor times because the model has never seen the codebase either. What the horizon measures is a smart newcomer, and the newcomer's handicap is exactly the information gap this page names
- Research Taste as the Human Bottleneck — the page this reframes. Its central open question ("is taste a genuine ceiling or the next jagged valley?") presupposes taste is a capability; Ng argues it is an asymmetry, which makes it neither ceiling nor valley but a gap that closes as context transfers
- Unknowns as the Agentic Bottleneck — the extraction protocol for the asymmetry; Thariq's unknown knowns are this concept viewed from inside the human's head
- The Three Loops of AI-Native Building — the context advantage is why the human occupies the middle loop: they are the transmission between the loop that knows the users and the loop that writes the code
- Returns to Expertise in Agentic Coding — the instrument. If the human role is a closable context gap, the measured decline in expertise premium is the gap closing, tracked in usage data
- Outsource Your Thinking, Not Your Understanding — Karpathy's residue is understanding, which is closer to Ng's "context" than to "taste"; both locate the human in what they know rather than in what they can appreciate
- Agent Context Files — the transfer mechanism, and the uncomfortable implication: every skill file you write spends a little of your own advantage
- The Abstraction Barrier — the limit case Ng's frame doesn't reach: being bounded by human conceptual frameworks is not a missing fact
- Transformative Creativity — Boden level-3 as the generative asymmetry that context transfer cannot repair
- Why AI Lags at Design — the test case: design taste is either a context gap (closes) or criterion-less discrimination (doesn't)
- Jagged Intelligence (Ghosts, Not Animals) — the framing Ng displaces: "a capability AI fails at then masters" assumes taste is a capability at all
- Implementation Abundance Inverts Product Work — Ambrosino names curation/taste as the new bottleneck without asking what taste is; Ng supplies the answer that makes the bottleneck perishable
- Engineer PM Convergence — the practical consequence: if the convergent skill is proximity to the user rather than taste, "hire for taste" is a perishable strategy
- Production-Sourced Evaluation — context transfer as infrastructure: sourcing evals from real usage moves the human's private knowledge of users into where the model can read it
- Andrew Ng — author of the reframe
- AI-Assisted Error Analysis — the sharpest disagreement with this page's reframe, and worth holding open. Shankar agrees the human contribution is information the AI lacks ("what good means for your product is living in your head. It's not in the traces") but argues the gap is structurally unclosable rather than an engineering gap: if a tool could fully find and fix your product's failures, it could do the same for every competitor, so nothing would differentiate the product. Ng's asymmetry is a closable lead; Shankar's is a moat — same observation, opposite prediction about whether it survives
- Printing Press Software Democratization — Cherny's "domain knowledge displaces coding skill" is this page's reframe stated as a diffusion: the accountant who writes accounting software is bringing context, not taste, and Returns to Expertise in Agentic Coding measured that context as the amplifier
Open Questions#
- Does the asymmetry regenerate faster than it transfers? The whole human role, under this frame, rests on the answer. Nobody in the corpus has posed it.
- Ng prefers the frame because it "gives us a clearer path to helping AI systems get better." That is a reason to adopt the frame, not evidence that it's true. What would distinguish a context asymmetry from a capability gap empirically? (Returns to Expertise in Agentic Coding is the closest thing to an instrument.)
- If the human's contribution is context injection, is the human replaceable by better context plumbing — memory, retrieval, continuous production telemetry — rather than by a better model? That would put the expiry of human-in-the-loop on the infrastructure roadmap, not the scaling curve. Ng's August 2026 answer, by assertion: the plumbing "does not exist and I don't think exists for the foreseeable future" — an opinion about the roadmap, not evidence about it.
- Ng writes from 0-to-1 consumer products. Does the frame survive contact with domains where the missing thing is a concept rather than a fact? Partially answered (2026-09-22) by Training novices to think, or giving them LLMs? Evidence from an RCT (
empirical, preregistered 2×2 RCT, n=1,053) — the first randomized test in this corpus in which a concept, not a fact, is experimentally supplied to the human side. Half the sample gets a ~6-minute causal-reasoning training (mechanisms, falsifiable boundary conditions, explicit cause→effect chains); half gets ChatGPT Edu; the cells cross. The concept takes: mechanism identification rises ~0.55 SD, falsification logic ~0.85 SD, idea diversity ~0.47 SD between solutions, and none of it is crowded out when a model is in reach. The concept does not pay: on the graded output its coefficient is negative (−0.111, p<0.01) while LLM access is worth +0.862 on a control mean of 2.09. So the frame survives, but only by narrowing to a conditional — what the human supplies is decisive only where the evaluator demands it. Here it was not: the rubric loaded positively on the model's coherence and idea count and negatively on falsifiability (−0.109), mechanisms (−0.166) and distance from the modal answer (−7.676, ≈ −0.29 points per SD of novelty). Two limits on how far to take this. The subjects are novices with no domain context at all, which is the opposite end of the population Ng writes about, and the task was chosen precisely because the model is already well-trained on it — so it tests the frame where the context gap is smallest by construction. And the study measures assisted output, never withdrawing the tool, so it says nothing about whether the concept was internalized.
Sources#
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks — Patwardhan et al. (19 authors, OpenAI), arXiv 2510.04374 v1, 2025-10-05, 29pp (
empirical). Cited here for Appendix A.2.7 and Figure 15 (the under-contextualized arm: prompts at 42% of original length, GPT-5-high 47.7% → 44.3% wins-and-ties, with the earlier-gold-set caveat the paper attaches) and §5's own limitation that GDPval tasks are precisely-specified and one-shot by construction. Full treatment and COI on GDPval Benchmark - Thread by @AndrewYNg — Andrew Ng, The Batch, published 2026-06-30 (
practitioner-opinion); the context-advantage passage sits inside the developer-feedback-loop section - A Field Guide to Fable: Finding Your Unknowns — Thariq Shihipar, 2026-07-04 (
practitioner-opinion); unknown knowns as the same object, and the color-grading episode as the hard case - Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI — Latent Space, 2026-07-28 (
practitioner-opinion); Akshay Nathan's failed "bring me new ideas" automation, explained as a grounding-stream deficit rather than a capability one - Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28,
practitioner-opinion): the brainstorming-ideas example, "the technical thing that underlies… taste," and the "long-term advantage" / "plumbing does not exist" restatement that softens the June gap reading. Auto-caption transcript
Cited by 25
- Research Taste as the Human Bottleneck×5
Under this reading the residue is not a faculty but an information asymmetry — the human knows the…
- Andrew Ng×3
Context Advantage Over Taste — his reframing of the residual human role as an information asymmetry…
- Implementation Abundance Inverts Product Work×3
That is the context-advantage explanation, not the capability explanation — and it bears directly…
- The Three Loops of AI-Native Building×3
Ng's account of what the human contributes is his most-quoted line, and it gets its own page: not…
- Aakanksha Chowdhery×2
Context Advantage Over Taste — the thesis her lecture-8 synthesis lands on, from three benchmarks…
- Engineer PM Convergence×2
Ng also declines the word this page leans on. What Cat Wu calls taste, he calls a context advantage…
- GDPval Benchmark×2
Context Advantage Over Taste — the underspecification ablation is this thesis as an experiment, now…
- Open Questions Backlog×2
Context Advantage Over Taste: Ng writes from 0-to-1 consumer products. Does the frame survive…
- Oversight When the Signals Give Out: the Activation Fallback and the Taste Reward×2
The observed failure mode already exists in miniature. Design By Selection: left undirected,…
- Task Time-Horizon Scaling×2
Context Advantage Over Taste — what the horizon is a horizon of: models track low-context…
- Unknowns as the Agentic Bottleneck×2
Unknown knowns are the same object Andrew Ng calls the human's "context advantage" — knowledge the…
- The Abstraction Barrier
Context Advantage Over Taste — the generative asymmetry this page defends, stated as the boundary…
- Agent Context Files
Context Advantage Over Taste — the uncomfortable reading: if the human's necessity is an…
- AI-Assisted Error Analysis
Context Advantage Over Taste — the direct disagreement worth holding open. Andrew Ng
- Andrew Ambrosino
Context Advantage Over Taste — supplies what Ambrosino's "taste is the bottleneck" leaves…
- CS329A: Self-Improving AI Agents (Stanford)
The synthesis is the lecture's contribution rather than any one benchmark. All three failures…
- Jagged Intelligence (Ghosts, Not Animals)
Context Advantage Over Taste — Andrew Ng displaces the framing this page supplies for taste: "a…
- AI Economics & Labor
Context Advantage Over Taste — Andrew Ng's reframing of the residual human contribution: not…
- Outsource Your Thinking, Not Your Understanding
Context Advantage Over Taste — Andrew Ng locates the residual human role in what they know rather…
- Printing Press Software Democratization
Implementation Abundance Inverts Product Work — the two answers to "what is scarce once anyone can…
- Production-Sourced Evaluation
Context Advantage Over Taste — production telemetry as context transfer: sourcing evals from real…
- Returns to Expertise in Agentic Coding
Context Advantage Over Taste — the instrument. If the human's role is a closable information…
- Thariq Shihipar
Context Advantage Over Taste — his unknown knowns are Andrew Ng's context advantage seen from…
- Transformative Creativity
Context Advantage Over Taste — the limit Andrew Ng's frame doesn't reach: Boden level-3 creation of…
- Why AI Lags at Design
Context Advantage Over Taste — the test case for Andrew Ng's reframe: if design taste is a closable…
Related articles
- Andrew Ng
Founder of DeepLearning.AI and AI Fund, founding lead of Google Brain, co-founder of Coursera; writes The Batch, where…
- Research Taste as the Human Bottleneck
The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Jagged Intelligence (Ghosts, Not Animals)
"Ghosts not animals": jagged statistical circuits, no intrinsic motivation; car-wash/strawberry failures; stay in the l…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
