Howardism · Vol. 03Plate II · No. 02
Product & Org, in order.
Notes24DomainProduct & OrgOpen Qs54Newest22 Sept 2026Oldest6 May 2026
Product cadence, org design, and the AI-native team.
Map of Content for the product-org domain — 20 concepts. Curated entry point; see Home for all domains.
- AI-Native Organization — Garry Tan's org-design mapping: skill files = employees, resolver tables = org charts, filing rules = process, trigger evals = performance reviews — a company whose operations are encoded as markdown that agents execute, with engineers hired to maintain the skills; claimed record revenue-per-head (Emergent ~$15M ARR at 15 people, Retell $60M at ~40) — and ICONIQ's four-year functional headcount split (S&M/R&D/G&A flat to within a few points) says the top-level org chart has not actually changed yet, whatever the intent surveys report
- AI Native Product Cadence (hub) — Cat Wu's 6mo→1mo→1day cadence at Anthropic: research-preview branding, mission-as-tiebreaker, evergreen launch room, lighter PRDs, weekly metrics readouts
- Build Instead of Buy Under Agentic Coding — McKinsey's 2026 state-of-AI survey supplies the first population-scale reading of agentic coding reaching the procurement decision: 32% of 1,719 respondents report their organization decided against buying at least one software product or feature because it could be built in-house with agentic coding tools (Exhibit 4) — 41% in technology, 19% in insurance, 17% in the public sector, and nearly half among the 6% AI high performers against 31% of everyone else. The enabling condition and the bound arrive with it: about two in ten organizations are scaling software coding agents (31% at ≥$1B revenue, 17% below), and the orgs furthest along report AI operating costs constraining coding-agent use three times as often as others. It is a decision that was reported, not a build that was verified — self-report, no dollar value, no completion, no maintenance horizon.
- Community Smells Under AI Adoption — PLS-SEM on 152 software professionals: AI adoption is associated with fewer socio-technical anti-patterns, by two different mechanisms — indirectly in specialization work (AI → more peer consultation → less knowledge fragmentation) and directly in coordination work (AI → better communication quality, with interaction frequency unchanged) — while a vocal minority of the same respondents report in free text that AI replaced their teammates
- Compounding Loop Optimization — Dan Carey's discipline of instrumenting and automating every recurring step of the build loop — because when internal tooling is an-afternoon-cheap, each optimization pays back ×(50–100 iterations per project)
- Dogfooding as Product Discipline — Product sense is built by relentless first-hand use ("ant food"); Mr. Peanut catch; cross-source (Cat Wu vibe-checks, Glasgow founder-led sales)
- Engineer PM Convergence — Generalists across disciplines; product taste as bottleneck skill; Anthropic Claude Code team as case study; "just do things" cultural substrate
- Evals as Product Spec — Cat Wu's framing of evals as the emerging core PM skill: ten great evals beats a hundred mediocre; encode what done looks like for ambiguous AI features; companion to introspection (hypothesis) and vibe-check (direction). Shopify raises the stakes — the rubric becomes the RL reward, so a mis-specified spec trains a bad model daily rather than shipping one bad feature — and supplies the corpus's only falsifiability test for a spec itself: two experts, 25 random samples, Cohen's κ, rewrite below ~0.2
- Excellence as an Operating System — Elizabeth Stone's account of Netflix culture: talent density, agency, and accountability are not values but a mechanism for excellence — resist process even when things go wrong (blameless retros + individual responsibility instead), run the keeper test in both directions; Lenny's observation that top AI labs now converge on the early Netflix culture deck
- Implementation Abundance Inverts Product Work — Andrew Ambrosino's inversion thesis: when talking to a frontier model can stand up any feature from scratch, implementation stops being the expensive step you derisk up front — so the process runs backwards and the costly work becomes curating the 90 uncoordinated builds people already produced; taste is the new bottleneck
- Managers as ICs — Every Claude Code manager starts as an IC; flat org; agentic coding collapsed the onboarding cost that pushed managers out of the codebase
- Model Introspection Feedback — Cat Wu's underrated technique: ask the model why it failed; treat answer as harness-debugging signal not model criticism; caveats around model self-report fidelity
- Pilot-to-Production Gap — Anthropic × Accenture's account of why enterprise AI pilots don't predict production: the pilot's success conditions — curated data, handpicked AI-native teams, protected budgets, narrow scope, hidden manual intervention, no downstream stakeholders — are exactly the complements production won't supply, so a pilot measures a system nobody will run. The prescription is a seven-decision blueprint front-loaded before the pilot (four pre-pilot, one preproduction, one in-production, one at scale), each with a named owner; the supporting figures are vendor-published surveys (23% sustained enterprise-wide impact, 64% past pilots vs 7% data-ready) and unattributed customer anecdotes.
- Polish No Longer Signals Readiness — Andrew Ambrosino's observation that the medium used to encode process-stage — a production-looking artifact meant late-stage, derisked, design-and-business-approved — but cheap implementation divorces polish from maturity: a 90-person exploration can look ready-to-ship while being early design work, and over-anchoring on it ('can we release this now?') is the trap
- Prototype Fidelity After Cheap Polish — Hundhausen's argument that GenAI decoupled polish from effort, invalidating the empirical basis of the low-fidelity-first playbook: the classic finding was that polish suppresses feedback because it signals sunk effort, and that signal is now false while the psychological barrier likely persists — plus the revival of Boehm's evolutionary prototyping and three unanswered research questions
- Prototype Over PRD — Dan Carey's prototype-replaces-PRD method: record a why-not-what conversation, transcribe it, hand the transcript to Claude, ask for a few prototype variations; the prototype is the spec, not a downstream artifact
- Psychological Costs of AI Adoption — Alami, Paja & Tiwari's one-year case study of a 1,200-person regulated-software firm (N=21 interviews): five costs practitioners carry that adoption strategies never price — uncertainty distress, accountability anxiety, cognitive load intensification, craft identity disruption, meaning and satisfaction erosion — plus two constructs, agency displacement (cost tracks how much of a role's generative work AI takes) and the verification tax (retained accountability converted into line-by-line review labour)
- Role Averaging, Not Role Elimination — Andrew Ambrosino's nuanced OpenAI-side take on role collapse: your role is 'the average of what you spend your time on' and tool-gatekeeping is eroding — but eliminating roles dangerously eliminates specialties with knowable best practices ('getting rid of the product role is a terrible idea'), and 'zone defense' coverage plus managers remain necessary because not everyone can work on everything in both breadth and depth
- Standardize the Infrastructure, Not the Tools — Shopify's inversion of the one-tool-per-job norm for AI: route every coding agent through a central LLM proxy so leadership gets cost control, per-team usage analytics, and model portability, while engineers keep free tool choice — buying optionality under uncertainty about which model or workflow wins, with MCP servers extending the same governs-access-not-engineers principle to internal systems; Accenture's Tokenomics figures price what the meter is for (42% of orgs have no single AI-cost owner; formal chargeback ties 32¢ of every token dollar to an outcome, 6× no allocation); and the first population-scale record of the model-mix decision actually being exercised — frontier-model token share 53%→45% in five weeks as firms impose company-wide defaults on cost-effectiveness grounds (Ramp, September 2026), outcome observed, mechanism still not
- Systems Thinking Over Specialization — Elizabeth Stone's Netflix hiring thesis: in an agent-heavy org the scarce profile is the systems thinker who abstracts across business domains into paved paths, design systems, and source-of-truth data — narrow specialists shrink to a few irreplaceable niches; AI fluency becomes a cross-level career-ladder overlay, and the trainable move is 'step out one click'
Derived#
- AI-Native Product Org Bottlenecks — AI-native product-org bottleneck is accountable taste at speed: dogfooding trains taste, evals encode it, and accountability owns the consequences as output volume rises
- Playbook Boundary Conditions: the Devil's-Advocate Substrate and the Prototype's Edge — Joint answer to two #oq/now items about where AI-native playbook prescriptions stop being general. (1) The founder's devil's-advocate prescription interacts with character training complementarily, not conflictingly: the prompted moves are framing-compliance tasks that work on any instruction-follower (asking for the competitor's best case never requires disagreeing with the founder), while character training supplies the unprompted honesty the prompts can't manufacture — so the technique is model-portable but its safety net is Claude-specific, and the residual risk (framing bias within the assigned adversarial task) is exactly the part neither layer covers. (2) Prototype-over-PRD's breakdown boundary is not backend-vs-frontend but observable-surface-vs-invariant: the corpus already holds a domain-matched artifact for each spec job (three PRs, tracer-bullet slice, ten evals, design_system.html), so what breaks at the backend is the clickable prototype, not the artifact-over-document principle; the PRD survives where no cheap artifact's surface covers the risk — cross-cutting invariants and cross-team coordination. Both answers have the same shape: every prescription has a substrate; know the substrate, know the boundary
- The PRD-Replacement Spectrum at AI-Native Speed — Four positions (grill-then-PRD → lighter-PRD → build-to-decide → prototype-is-spec) are one spectrum once you decompose the PRD into three jobs: AI-native speed dissolves specification, relocates alignment, and orphans rationale
- Where Does the Why Live? — Rationale (the 'why') is well-homed at authoring time — it's the recorded why-not-what conversation and the grilling session — but orphaned for future readers: AI-native methods delete the PRD, bury discussion in PRs, and the prototype shows what not why; code explicitly can't hold it, context files hold policy not product-rationale, and only the richer-artifact axis partly answers it
Open questions 54 open
- SourceDoes the cadence scale beyond ~100 people? Anthropic itself is bigger (~30-40 PMs alone), but the Claude Code team that visibly drives cadence is small.
- SourceWhat's the equivalent of research-preview branding for B2B enterprise launches where customers expect stability? Cat doesn't address.
- SourceHow much of the cadence is structural (process choices) vs cultural (talent density)? Probably both, ratio unclear.
- AI-Native Organization3 open
- SourceTan's revenue-per-head figures (Emergent ~$15M ARR at 15 people, Retell $60M at ~40) are stated from stage without sourcing. Do third-party data (Carta/Standard Metrics cohorts, press-verified ARR) corroborate record revenue-per-head at AI-native YC companies, or do these examples regress toward the AI Investment Story, Not Efficiency Story mean on inspection? (Partially answered: TechCrunch, 2026-07-15 corroborates Emergent as a real, fast-growing $1.5B unicorn — $120M company-reported ARR, 200K+ paying customers — so the direction holds. But the record per-head claim compresses: at scale it is ~$600K/head (200 employees), below the $100M+ top-decile AI-company RPE of $960K AI Investment Story, Not Efficiency Story and about half the ~$1M/head of Tan's own 15-people/$15M snapshot — the per-head extreme is a low-headcount-phase artifact that regresses as the company staffs up. Caveats keeping this open: TechCrunch's figures are themselves company-reported
vendor-claim, not Carta-audited, and the Retell half ($60M at ~40 ≈ $1.5M/head) remains unverified. AWS's June-2026 founder survey adds a population reading (55% of AI-natives self-report $400K+/head) but it is self-report, not the cap-table/press verification this question asks for. See Emergent.) Sharpened by a peer-group benchmark, and it points against the claim (2026-09-22): iconiq pacesetter index 2026 (empirical, quarterly operating financials 2024 – Q2 2026, not a survey) is the closest thing yet to the third-party cohort data this bullet asks for. It benchmarks revenue per FTE for companies that are AI-forward and top-quartile three-year growers — Tan's exhibits' own peer group — at $655K median / $890K top quartile for $100M+, and $75K / $115K / $225K medians in the three smaller bands. Two consequences. Emergent's ~$600K/head lands marginally below the median of its peer group, not at a record. And the sub-$10M median of $75K/FTE says the fast-growing AI cohort is, at small scale, less productive per head than conventional software — so Tan's "15 people at $15M ARR" (~$1M/head) is an extreme outlier against its own cohort rather than a new normal. Still not a full answer: ICONIQ selects its cohort with no control group and is benchmarking its own portfolio in part, the Retell half remains unverified, and this is a cohort distribution rather than the per-company verification the bullet requests. - SourceThe org mapping predicts a testable staffing signature: AI-native companies should hire engineers to maintain skills rather than function-specific staff. Does job-posting data show a "skill maintainer / agent ops" role emerging as a distinct hiring category? (Partially answered: ICONIQ, Q2 2026 confirms the composition shift the signature predicts — 45% of ~305 AI-builders plan a "different mix of roles (fewer operational, more AI-fluent talent)," function-level headcount reallocates toward R&D/Product/Sales and away from Customer Support/G&A, and G&A operators are "removing finance-ops and order-management roles… redirecting budget to strategic and AI-specific functions." It also names the concrete new engineering hiring categories: forward-deployed engineers (~50% scaling as a permanent motion) and AI safety / trust & reliability engineers. What it does not supply is the specific "skill maintainer / agent ops" title from job-posting data — ICONIQ measures function-level headcount intent and two named roles, not an occupational taxonomy. The direct test (a "skill maintainer / agent ops" posting category) still needs job-posting/occupational-emergence data, e.g. the un-ingested arXiv 2606.22769 "Agent Systems Engineer" signal from the 2026-07-21 research pass. See the restructuring section above and AI Product Economics Maturation for the FDE detail.) Checked against the largest available population instrument and still not answered (2026-09-22): mckinsey state of ai 2026 road to roi (
empirical, self-reported, 1,719 respondents in 97 nations) publishes no occupational or role taxonomy at all — its workforce section forecasts headcount levels by function (39% expect an AI-related decline in the coming year) and never asks what roles are being created. The nearest row in its practice chart runs the other way: "completed a strategic workforce planning exercise" is reported by only ~24% of AI high performers and ~13% of everyone else (Exhibit 11, gridline-read, approximate), the second-narrowest gap in the chart. So at population scale the exercise that would produce a named skill-maintainer role is the practice organizations are least likely to have done, including the ones attributing EBIT to AI. The direct test still needs job-posting or occupational-emergence data. Three more names, and a negative result on the coarser instrument (2026-09-22): iconiq 2026 state of scaling names AI Business Analyst, AI Integration Engineer and GTM AI leader as emerging titles in ICONIQ's portfolio — closer in spirit to the skill-maintainer signature than anything previously logged, since all three sit between a function and the AI systems it runs. But they arrive as an investor's uncounted qualitative observation, with no share of companies, no posting volume and no definition, so they are anecdote, not the occupational data the bullet specifies. The same report supplies the negative result at the level it can measure: functional headcount distribution is flat across four years (S&M/R&D/G&A ≈ 50/41/9 under $100M, 47/39/14 above), with ICONIQ concluding the workforce model 'has yet to meaningfully shift.' That is a coarse instrument — an intra-functional role substitution is invisible to a three-bucket chart — so it bounds rather than answers: whatever role emergence is happening is not yet large enough to move the buckets. The direct test is unchanged. - SourceIs the encoded-role form of the employee metaphor actually accountability-preserving, as the synthesis above suggests, or do Kropp-style framing effects attach to skill-files-as-employees too once teams talk about them that way? No study has tested framing effects on artifact-level anthropomorphism.
- SourceTan's revenue-per-head figures (Emergent ~$15M ARR at 15 people, Retell $60M at ~40) are stated from stage without sourcing. Do third-party data (Carta/Standard Metrics cohorts, press-verified ARR) corroborate record revenue-per-head at AI-native YC companies, or do these examples regress toward the AI Investment Story, Not Efficiency Story mean on inspection? (Partially answered: TechCrunch, 2026-07-15 corroborates Emergent as a real, fast-growing $1.5B unicorn — $120M company-reported ARR, 200K+ paying customers — so the direction holds. But the record per-head claim compresses: at scale it is ~$600K/head (200 employees), below the $100M+ top-decile AI-company RPE of $960K AI Investment Story, Not Efficiency Story and about half the ~$1M/head of Tan's own 15-people/$15M snapshot — the per-head extreme is a low-headcount-phase artifact that regresses as the company staffs up. Caveats keeping this open: TechCrunch's figures are themselves company-reported
- SourceDoes the forgone purchase stay forgone? No source records whether an in-house replacement survives a maintenance cycle, or whether the vendor is re-engaged at renewal + 24 months. The decision is captured at decision time only, and the re-buy is the falsifying event. A follow-up wave asking the same respondents about the specific product they declined would settle it.
- SourceIs the 32% moving real money? The instrument has no dollar value and no size floor, so the figure is compatible with a large budget reallocation and with thirteen declined add-ons. A payment-rail instrument applied to software-vendor spend — the Ramp aperture pointed at SaaS line items rather than AI vendors — would separate the two, since a genuine displacement shows up as decelerating net-new software vendor spend in the industries reporting the highest rates.
- SourceDoes the build option get rationed by tokens? High performers report build-instead-of-buy at ~48% and cost-constrained coding-agent use at 18% (3× everyone else), but the survey publishes no cross-tab, so it cannot say whether the same organizations hold both. If they do, the procurement effect has a price ceiling and will plateau rather than compound.
- SourceThe design cannot separate "AI adoption improves team social health" from "healthier teams adopt AI better", and the authors say so. The discriminating study is the one they name: longitudinal or quasi-experimental tracking of AI adoption and team dynamics over time, controlling for communication culture, org maturity, leadership practice and seniority. Until then every coefficient here is an association. Still unfilled as of 2026-09-22 — psychological costs ai adoption software engineering (
case-study) is the corpus's closest approach and is explicitly not it: a holistic single-case design at one firm, one year in, cross-sectional by the authors' own statement ("our snapshot cannot establish which relationships persist, weaken, or change over time; these temporal boundaries require longitudinal investigation"), with no comparison organization and no control for any of the four confounders named here. It substitutes depth for identification, so it can describe mechanisms this survey can only infer and cannot break the reverse-causation tie. - SourceThe Information Sharing path — AI adoption directly worsening documentation and information governance — is the only harm signal in five models and lands at p = .069, below the β ≈ .20 the sample can reliably detect. Does it survive at n ≈ 400, and does it strengthen in teams without a documentation discipline? This is the falsifiable half of the paper's "governance-dependent" conclusion.
- SourceThe study measures peer-interaction frequency, not what the interaction contains or what expertise is retained. If AI raises the count of specialization-oriented exchanges while lowering their depth (the comprehension register), frequency would rise exactly as measured while the underlying transactive memory thins. Nothing here distinguishes the two. Partially answered (2026-09-22) by psychological costs ai adoption software engineering (
case-study, N = 21 interviews at one regulated-software firm one year into adoption): it looks at exactly the layer this question says is unmeasured — what circulates — and finds it filtered rather than thinned. The firm's entire adoption strategy was "experiment-and-share," and sharing held "where roles were similar, psychological safety was sufficient, and permissibility was clear," failing on three named channels: experiments shared enterprise-wide "weren't really relevant for my role" (testers got nothing usable from a developer-dominated stream); practitioners withheld experiences "for fear of judgment," perceiving a high barrier to exposing their AI practices to colleagues; and governance ambiguity suppressed disclosure at source ("I am very hesitant at sharing this"). That is content-level evidence that AI-era exchanges are selected on who is safe to tell and what is safe to say — a distinct thinning mechanism from the depth erosion this bullet hypothesizes, and one a frequency count would also miss. Two reasons it is not an answer: it never measures interaction frequency, so it cannot say whether the count rose while content narrowed, and it studies one firm whose compliance regime plausibly drives the third channel. The within-subject design this question needs — exchange counts and exchange content on the same teams — remains unbuilt.
- SourceThe design cannot separate "AI adoption improves team social health" from "healthier teams adopt AI better", and the authors say so. The discriminating study is the one they name: longitudinal or quasi-experimental tracking of AI adoption and team dynamics over time, controlling for communication culture, org maturity, leadership practice and seniority. Until then every coefficient here is an association. Still unfilled as of 2026-09-22 — psychological costs ai adoption software engineering (
- SourceThe loop assumes the team is (close to) the user. How much of the compounding advantage survives when the user is unlike the builder and "talk to users" can't be same-room?
- SourceWhere is the line between worthwhile internal tooling and yak-shaving? Carey's "afternoon" bar is the heuristic, but Cat Wu warns that over-customizing setups "becomes distraction."
- SourceDoes Claude-as-first-pass-on-all-feedback ever filter out the rare signal that doesn't cluster? Automating triage optimizes the common case; the tail is where surprising bets come from.
- SourceDogfooding works when the team is the user (Claude Code) or near it (Cat Wu, Boris). How do you build product sense for users very unlike you — does "talk to customers" fully substitute, as Glasgow/Fung's small-business work suggests?
- ResolvedCan dogfooding scale, or does it implicitly cap how large an AI-native product org can stay taste-driven before it reverts to dashboards? Answered: The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales — false binary: dogfooding itself never scales (first-hand use is per-person and breaks when the team stops being the user), but the taste it produces scales through three named mechanisms — encoding into runnable artifacts (Evals as Product Spec), concentrating the rare-trusted-evaluator role plus vibe-check rituals rather than diffusing taste with headcount (Claude Character as Product), and AI-extended contact surface (Carey's Claude-first-pass on every user conversation). The cap variable is not org size but team-user distance plus encoding discipline: an org reverts to dashboards when it stops converting felt use into evals and rituals, at any size.
- Engineer PM Convergence3 open
- WaitDoes this scale beyond ~50-person Claude Code-style teams? Boris hedges: "I think this is going to be a question for years."
- WaitWhat happens to formal PM career ladders in companies where engineers do PM work? Open at Anthropic per Cat.
- SourceCross-disciplinary generalist is a hiring bar — where does the supply come from? Career changers, or new-grad bias toward AI-native education?
- Evals as Product Spec4 open
- How do you write an eval for taste-driven features like character? Amanda's role is canonical for being eval-resistant; Cat names her as someone who is good at evals here, but doesn't describe the technique. Partially answered: How Do You Write Evals for Taste? Character as the Limit Case — the technique is a pipeline (conviction → dogfood-sourced failure modes → MSM-style variant A/B measurement → ~10 interpretable evals); proven on the safety/values core but still tacit on the warm/witty aesthetic surface.
- SourceThe 10-vs-100 number is given without justification. Is there a Goldilocks zone, or does it depend on feature surface area? Client-Side Agent Optimization's framing of combos suggests evals also have a combinatorial explosion problem. Partially answered (2026-08-13), and it reframes the axis: shopify sidekick continual learning loop (
case-study) prescribes "keep each judge small and targeted rather than cramming all of your product's behavior into one… focused judges make these tests easier to interpret and the resulting metrics easier to trust. You can always add more." So the quantity that must stay bounded is behaviours per instrument, not instruments per product — the count scales with surface area and the interpretability of any one instrument does not. Their operational reason is sharper than "maintenance cost": a blended judge cannot pass a per-criterion degradation test, so decomposition is what makes an eval falsifiable at the criterion level rather than merely readable. Still open on the number itself — no source measures where the returns to splitting stop, and this is one unreplicated first-party account. - SourceHow do evals interact with Harness Shrinkage as Models Improve? When a harness asset shrinks because the model now handles it natively, the evals built around the old harness may become artifacts rather than guardrails. Does Anthropic retire evals or repurpose them? Partially answered: Boris Cherny (YC interview, 2026-07-27,
practitioner-opinion) — retire: evals live "one, two, three model generations," then saturate and get thrown away and rebuilt from observed struggle; what persists is the authoring practice, not the artifact. Still open: whether any eval class (safety, character) is exempt from the saturation cycle. - SourceIs there a single non-Anthropic example of a PM-as-eval-writer to cite, or is this currently a Cat-Wu-singular framing? The Matt Pocock workshop reaches the same place from a different vocabulary, but no third source has been ingested yet. Partially answered (with a twist): Google's Agent Quality Flywheel is a third-party arrival at eval-as-the-quality-surface — but its answer is to have the coding agent author the eval, compressing the human role to stating the worry and approving the plan. A second non-Anthropic arrival (2026-07-03) pushes back on exactly that compression: Shankar & Husain teach eval-as-the-core-discipline to practitioners and argue the authoring step is where automation stops paying. Still not a PM-as-eval-writer instance — Shankar is a CS professor and Husain an ML engineer, so the role-relocation half of Cat's claim remains Anthropic-singular.
- SourceIs the AI-lab convergence on early-Netflix operating norms (agency, density, top-of-market pay) causal inheritance (the culture deck as a founding document for lab founders) or convergent evolution under the same constraint (scarce elite talent)? A history of lab founding cultures could settle it.
- WaitDoes talent-density-plus-paved-paths actually substitute for process at agent-scale throughput, or does Netflix eventually show the Acceleration Whiplash quality signature (incident rates, review latency) like Faros's high-maturity cohort? Trigger: future Netflix engineering telemetry or tech-blog disclosures.
- SourceCuration of 90 uncoordinated builds is itself expensive and doesn't obviously scale — is there a point where the cost of curating parallel exploration exceeds the cost it replaced? ("zone defense" is Ambrosino's partial answer.)
- WaitIf taste is the bottleneck and taste is "just another capability" AI eventually masters, does the inversion invert again — does curation migrate into the model?
- SourceThe 90-uncoordinated-builds picture assumes abundant tokens and an agentic culture; how much of the inversion survives outside a frontier lab that gives everyone "unlimited tokens"? Partially answered (2026-09-22) by mckinsey state of ai 2026 road to roi (
empirical, self-reported; 1,719 respondents in 97 nations, fielded May 4 - June 8 2026) — the first population-scale reading of the premise, and it splits in three. The build capability survives the trip out of the lab: 32% of respondents report deciding against a software purchase because the functionality could be built in-house with agentic coding tools, and the industry spread is narrow (41% technology, 38% energy and materials, 38% professional services, 28% engineering and construction, 17% public sector) — so this is not a tech-sector artifact, let alone a frontier-lab one. About two in ten report scaling software coding agents (31% at $1B+ revenue, 17% below). The unlimited-token premise does not survive: ~20% report AI operating costs including tokens constraining their use, and the bite is concentrated where the mechanism runs — the high-performer cohort reports cost-constrained coding-agent use at 18% against 6% of all others, the only tool where the leaders report more constraint than the rest, with Chui stating the reason (consumption growing faster than price falls). Outside the lab the abundance is metered, and the meter tightens as practice approaches the frontier. The inversion itself is untouched: the survey counts purchase decisions and scaling phases, never parallel builds, curation cost or process shape, so the specific claim that the expensive step migrates to curation remains unmeasured anywhere, inside a lab or outside one. What changed is that the precondition is now a population fact with a price attached rather than a lab affordance.
- Managers as ICs2 open
- WaitFung's own open question: "Do you still need separate iOS and Android orgs?" — if engineers flex across platforms via Claude, the traditional platform-split org may dissolve too. How far does flattening go?
- WaitDoes manager-as-IC scale past a certain org size, or only work while Claude Code is small and the codebase is Claude-legible?
- How reliable are 4.7-class introspective reports? Anthropic's interpretability research suggests partial fidelity but not full. Empirically, Cat reports it's good enough to drive harness fixes — but unclear at what model scale this technique becomes load-bearing. Partially answered: Self-Report as a Safety Signal — reliability is context-dependent. In a benign debugging setting the report is good enough to drive harness fixes; in an adversarial safety setting, open-weight models (3B–70B) fail to recognize their own compromised outputs 27.3% of the time, and the recognition that exists is the refusal circuit firing late rather than genuine own-output introspection. So the channel Cat relies on is a weak safety signal even where it is a useful debugging one. Also partially answered: Introspective Coupling puts a floor under the scale question — an untrained 8B model manages only 14–18% exact match at predicting its own counterfactual behavior, so at that scale the channel carries essentially nothing without explanation training.
- Does adversarial introspection ("why did you fail?") yield different signal than neutral ("walk me through your reasoning")? Worth probing. Partially answered: Self-Report as a Safety Signal finds self-attribution is heavily framing-dependent — an "intention" probe and a "tampering" probe elicit qualitatively different answers on the same models (some families deny tampering ~100% of the time regardless), so the phrasing of the introspective question materially changes the signal.
- SourceCould a meta-agent run introspection automatically against logged failures? Sounds tractable but no public implementation.
- Pilot-to-Production Gap4 open
- SourceGartner's February 2025 prediction — organizations abandon 60% of AI projects unsupported by AI-ready data through 2026 — reaches its window this year, and this document cites it without checking it. Did the abandonment rate land near 60%, and was data readiness the separating variable it was predicted to be? A resolved forecast would convert the document's central data-readiness claim from a vendor survey into something gradeable. Partially answered (2026-09-22) by mckinsey state of ai 2026 road to roi (McKinsey/QuantumBlack,
empiricalbut self-reported, 1,719 respondents in 97 nations fielded May 4 - June 8 2026 — squarely inside the prediction's window, and the strongest in-window population datum the vault holds). It settles the direction and leaves the rate untouched, and the two halves of the question come apart. On abandonment: nothing in the aggregate looks like a 60% walk-away. The share of AI-using organizations stuck in experimenting fell 32% to 22% year over year while piloting rose 30% to 34% and at-least-scaling rose 38% to 44%; 89% report regular use in at least one function, 56% in three or more. But the survey never asks about abandoned, cancelled or paused projects, so no abandonment rate of any kind exists in it — and the unit is wrong for the forecast either way: the phase question is asked of the organization, so a firm scaling in two functions can have killed five projects in others with the instrument seeing none of it. Gartner's number is a rate over projects; this is a distribution over organizations, and no arithmetic converts one into the other. On data readiness as the separating variable: the only data-layer practice published is whether the organization has a semantic layer or knowledge graph, ~26% of high performers against ~13% of all others (Exhibit 11, gridline-read, approximate). A 2x gap, and one of the smallest in that chart — against ~72% vs ~25% on workflow redesign, ~62% vs ~15% on transformative ambition and ~65% vs ~32% on senior-leader role modeling. So on this instrument data readiness separates, but three organizational variables separate harder, and the prediction's implied causal privilege for the data layer is not what the population shows. Both discounts stand: the high-performer cohort is defined by its outcome, and the whole survey is self-report from a panel McKinsey fields and whose AI transformation work McKinsey sells. What would still close this bullet is an instrument that counts projects — a portfolio census or a vendor-side cancellation series — and none exists in the corpus. One more datum, in the right unit but the wrong weight (2026-09-22): techcrunch startup arr less secure madrona reports Madrona's survey of 150 enterprise IT professionals finding fewer than half of their AI pilots ever reach full production — a project-level rate, which is Gartner's unit and not McKinsey's, but at n=150 from a VC survey reaching the wiki through a news article (practitioner-opinion, primary unread), and measuring non-promotion rather than the abandonment Gartner forecast; a pilot that never ships is not necessarily a project that was killed. It is directionally compatible with a 60% attrition rate and cannot distinguish it from a 40% one. A third unit, secondhand (2026-09-22): adoption telemetry enterprise ai production signals cites S&P Global Market Intelligence for the share of companies abandoning most of their AI initiatives rising 17% → 42% in a single year. That is abandonment, unlike McKinsey's phase distribution, but it counts companies-abandoning-most rather than projects-abandoned, carries no data-readiness split, and reaches the vault third-hand inside apractitioner-opinionpaper that itself cites it as "directional evidence of the failure consensus" and nothing more. The project census this bullet needs still does not exist in the corpus. - SourceThe blueprint's load-bearing claim is an ordering claim: decisions made pre-pilot cost less than the same decisions made post-pilot. Nothing in the document measures this, and the counterfactual is available in principle — two programs, same use case, differing only in whether ownership, data audit, infrastructure and security were settled before the pilot. Does front-loading actually shorten time-to-production, or does it move the same calendar time earlier and add executive-attention cost the document never prices? Extended 2026-09-22 by mckinsey state of ai 2026 road to roi, which bears on this only negatively and is worth recording so the next reader does not re-derive it. Exhibit 11 measures the presence of ten practices in a cross-section; it never measures when any of them was adopted. Two of its widest gaps are this blueprint's pre-pilot decisions restated — transformative ambition (a strategy decision, ~62% vs ~15%) and workflow redesign (~72% vs ~25%) — which is consistent with the ordering claim and exactly as consistent with its converse, that organizations which got value went on to acquire the practices. A survey fielded once cannot break that tie, and this one publishes no adoption dates, no time-to-production, and no cost of the decision process. The counterfactual the bullet asks for remains unavailable.
- SourceConsideration 05 prescribes converting operators into overseers ("the claims processor who used to review every invoice now audits the system that reviews them") while The Tragedy of the Cognitive Commons argues the judgment that role needs is regenerated by the execution work being removed. Is there a deployment where an oversight workforce was staffed without an execution cohort beneath it, and did its error-catching hold over more than a year? This is the falsifiable core of the disagreement and neither source supplies it. Partially answered (2026-09-22) by psychological costs ai adoption software engineering (
case-study, N = 21 interviews at a 1,200-person regulated-software firm, exactly one year after adoption launched): it supplies the deployment and one half of the answer. Engineers there were converted into overseers of AI-authored code while their accountability for it was left unchanged, and the conversion held — the overseers do catch things, by reading "the result line by line" and by applying heightened scrutiny specifically to AI-authored contributions in peer review. But the conditions are the ones that make the finding not transfer to the question's premise. The oversight cohort is entirely ex-executors: 19 of 21 participants have a decade or more of industry experience, median 20 years, so the judgment being exercised was accumulated before the execution work was removed, not regenerated after. The prediction survives untested, and the cohort itself names the mechanism unprompted — fears of "skill atrophy" that might leave practitioners "unable even to critically review AI-generated solutions," countered privately with deliberate manual-coding rituals nobody asked them to perform. Two further gaps: the study measures no error-catching rate of any kind (it is interview evidence about experience, not an accuracy instrument), and one year is inside, not past, the window the question asks about. What it does settle is that the conversion is not costless even when it works — the overseers pay a verification tax the organization neither measured nor budgeted. A second axis added 2026-09-22 by layered supervision ai assisted software engineering (case-study, five practitioner interviews, ESEM 2026 Software Engineering in Practice), which does not answer the question but changes what it is asking. Consideration 05 converts the operator into an overseer of the same artifact — the invoice auditor now audits the invoice-auditing system. These five teams moved oversight somewhere else: off the artifact entirely and up to architectural reasoning, contextual appropriateness and long-term maintainability, with complete line-by-line comprehension explicitly abandoned and replaced by operational explainability, the requirement that a generated system stay diagnosable and reconstructable on demand through abstraction layers, visualizations, logs and higher-level interpretive tooling. That distinction bears directly on the cognitive-commons disagreement this bullet is about: an overseer who audits the artifact needs the judgment the execution work used to regenerate, while an overseer who audits the architecture needs a judgment the execution work never supplied in the first place — so the two positions may not even be disagreeing about the same workforce. The same source supplies the staffing prediction the blueprint omits, and it points the pessimistic way: teams that adopt AI tools widely "without proportionate senior supervisory capacity may accumulate risk not visible from within the team", and the participants who used executable constraints most heavily were the ones with the strongest architectural expertise. The gaps are the same ones as before and one worse: five interviews, no error-catching rate of any kind, no longitudinal arm, and a cohort that is again entirely senior (architect/CTO, staff manager, team lead, two senior engineers), so the regeneration prediction is untested for a third time. A third axis added 2026-09-22 by mckinsey state of ai 2026 road to roi (empirical, self-reported, n=1,719): the first population-scale, role-level reading of the cost, on a panel three orders of magnitude larger than either case study and with the executive layer included as a comparison group. Exhibit 5 cuts perceived AI impacts by organizational level, and the strains sit below the executive line — 47% of midlevel managers and individual contributors report at least one negative effect, against 31% of executives and senior managers. The item closest to this bullet's mechanism is "hurts my ability to think critically": 20% of midlevel managers and 16% of individual contributors, against 9% of C-level and 8% of executive/senior managers. Career anxiety splits the same way (19% / 18% vs 6% / 10%), as does mental fatigue (14% / 7% vs 7% / 7%). Three things this is not: an error-catching rate, a longitudinal measure, or an observation of anyone's judgment — it is one-shot self-perception, and the survey has no oversight-workforce question at all. What it adds is that the cost the 21-person interview cohort named privately is distributed across 1,719 respondents in 97 nations, and that it is systematically not felt by the executives who commission the conversion — which is a structural reason the blueprint prices the capability and not the transition. One counter-row in the same exhibit keeps this honest: midlevel managers also report "think more clearly" at the highest rate of any level (54% vs 39 / 33 / 37), so the cut is a divergence in experience, not uniform pessimism below the C-suite. A third position on the same disagreement, with no deployment behind it (2026-09-22): when does ai augment work workflow framework (CIVIC-AI workshop whitepaper,practitioner-opinion, no measurement) sides against Consideration 05 as written and supplies the lever the disagreement lacked: reviewer competence "does not persist automatically. It requires organisations to deliberately assign workers enough substantive review work and enough exposure to AI failure to keep the verification skill current." On that account an oversight workforce is staffable conditionally, which narrows and cheapens the falsifiable core — not "was an oversight workforce ever staffed without an execution cohort beneath it," but "was its substantive-review share deliberately maintained, and did its error-catching hold." Its own worked example applies the rule (juniors must perform the initial protocol design and coding themselves, justified by anchoring bias and the risk of "vacuous verification" of an AI's generation). No deployment, no year, no error-catching measurement: unmoved on evidence, better specified on design. - SourceNANTE's practical claim is that stall types are different kinds of problem: a Navigate stall with a shallow-plateau flag needs workflow redesign, a Navigate stall with a low-task-success flag needs enablement, and applying either to the other is waste. Nothing tests this. Does a cohort diagnosed with one stall type respond differentially to its mapped intervention class versus the alternatives? adoption telemetry enterprise ai production signals names this as the strongest of its own proposed validations (§8.1) and runs none of it — its evaluation is synthetic populations whose pathologies and thresholds were co-developed, which establishes computability and not correctness. Falsifiable with a real deployment: classify cohorts prospectively, assign interventions across rather than only along the mapping, and compare depth conversion afterwards.
- SourceGartner's February 2025 prediction — organizations abandon 60% of AI projects unsupported by AI-ready data through 2026 — reaches its window this year, and this document cites it without checking it. Did the abandonment rate land near 60%, and was data readiness the separating variable it was predicted to be? A resolved forecast would convert the document's central data-readiness claim from a vendor survey into something gradeable. Partially answered (2026-09-22) by mckinsey state of ai 2026 road to roi (McKinsey/QuantumBlack,
- SourceIf the medium no longer signals stage, what does — is explicit human labeling ("this is exploration") the only mechanism, or can tooling re-attach the signal (e.g. a visible "exploration / preview / prod" marker on every build)?
- WaitDoes over-anchoring get worse as builds get more polished, or does everyone eventually recalibrate and learn to discount fidelity entirely?
- SourceDoes the Schumann effect survive the loss of its mechanism — do users still soften feedback on polished artifacts once told the artifact took an hour? The article's Question 1, and the field's key unknown.
- SourceWhich feedback is lost to polish: strategic (workflow, information architecture) or tactical (visual polish)? The distinction determines whether the low-fi-first playbook mattered for the reasons its advocates claimed.
- SourceDoes GenAI-generated prototype code actually evolve into production, or rebuild? Hundhausen poses it as open; this corpus's debt evidence suggests rebuild, but no source measures prototype-to-production survival directly.
- Prototype Over PRD1 open
- NoteThe prototype-as-spec must not become the prototype-as-validation trap Problem-Solution Fit Discipline warns about: a fast prototype proves the build was solvable, not that the problem is real.
- ResolvedIf there is no PRD, where does the rationale ("why we chose variation B") live for future readers? Same rationale-capture gap flagged in Building Is Cheap, Arguing Is Expensive. Answered: Rationale as a Dated Record: Where the Why Lives for the Next Reader — in a dated decision record written after the choice and committed beside the prototype: for this method, the recorded why-not-what transcript plus one line naming the winning variation and why. It does not reintroduce the PRD, because what made the PRD a PRD was authority (later work judged against it), and a post-hoc record binds nothing;
intent.mddoes not qualify, since the spec is generated from it. Earlier partial answer: Where Does the Why Live? (well-homed at authoring time, orphaned at read time). - ResolvedWhere does prototype-over-PRD break down? Carey's domain is a visual design tool where a prototype is the product surface; for backend/infra/data work the prototype may not capture the spec (cf. AI Native Product Cadence's "full PRD for heavy-infra features"). Answered: Playbook Boundary Conditions: the Devil's-Advocate Substrate and the Prototype's Edge — the boundary is observable-surface-vs-invariant, not backend-vs-frontend: every domain has its own tracer artifact (three PRs, vertical slice, ten evals, design_system.html), so what breaks at the backend is the clickable prototype, not artifact-over-document; the PRD survives where no artifact's surface covers the risk — cross-cutting invariants and cross-team coordination.
- WaitAbsorption is a response category defined by the loss being appraised as irreversible, and the paper cannot say what it turns into. Do practitioners who reported absorbing craft-identity and meaning losses at one year show elevated attrition, burnout or disengagement at three — or does absorption stabilize into a new baseline once the comparison to pre-AI work stops being available? The design that settles it is the authors' own: track this cohort, not a new one.
- SourceThe pathway figures give craft identity disruption and meaning erosion no organizational amplifier, which is a strong claim derived from one firm's coding — it implies no adoption strategy can touch the two costs practitioners called irreversible. Does a firm that deliberately preserved generative work for engineers (protected non-delegated tasks, bounded AI scope by role) show lower identity and meaning costs at comparable adoption depth, or do the two costs track raw exposure regardless of work design? The intervention gets a second, independent prescription — and still nobody has run it (2026-09-22): when does ai augment work workflow framework (CIVIC-AI workshop whitepaper,
practitioner-opinion, no measurement) makes job purpose one of six conditions for genuine augmentation and states this bullet's concern as that condition's failure mode almost verbatim — "AI does desirable tasks, while humans are relegated to undesirable ones" — while its worked example prescribes exactly the protected-generative-work arm the bullet asks for: juniors perform the initial protocol design and coding themselves rather than reviewing an agent's, on anti-anchoring and anti-vacuous-verification grounds. The useful addition is the argument for running the comparison, which is not a morale argument: the framework puts learning and purpose upstream of meaningful human control, making protected generative work a prerequisite for the firm's own oversight integrity rather than an equity concession. No firm, no measurement, no matched adoption depth — the bullet stands as posed. - SourceEvery magnitude in the role analysis rests on a sample whose industry experience has a median of 20 years. Do practitioners whose professional identity formed after generative AI report craft identity disruption and meaning erosion at all, or does the cost profile collapse to the accountability/verification half? A replication stratified by career-entry cohort would separate a transition cost from a permanent one, and the paper's 10/10 engineer rows cannot.
- SourceWhere is the equilibrium between fluidity and specialty — how much role-averaging before a company loses the accumulated best practices Ambrosino warns about?
- SourceZone defense assumes enough high-taste people to cover the whole company; does it degrade in orgs without OpenAI's talent density, collapsing back to top-down planning?
- SourceDoes "your role is the average of what you spend time on" survive performance review and career ladders, or does it fragment them the way Cat Wu flags ("we're sacrificing product consistency")? Partially answered: Netflix's move (Systems Thinking Over Specialization) is to leave per-level criteria untouched and add a cross-level AI-fluency overlay, explicitly because the tech shifts too fast to encode per level — one large-org existence proof that ladders survive by absorbing fluidity as an overlay rather than rewriting levels. Extended 2026-09-22 by psychological costs ai adoption software engineering (
case-study, one 1,200-person regulated-software firm, N = 21): a second large org built the overlay — a five-level AI maturity model, novice to agentic, that practitioners self-assess against — and it landed as a pressure instrument rather than a neutral rubric, which is the failure mode the Netflix case does not exhibit. It was never formally attached to performance review; it did not need to be. Leadership read it as evidence of progress ("a lot of them has moved to step 3 and 4"), practitioners read it as evaluation ("so they defined maturity levels ... So they're pushing us to evaluate us. They are pushing us towards a certain level in the department") against a background of visible layoff news ("all the measuring of people's AI maturity and that it's all over the news that this is causing people to be fired ... everybody is also motivated to use AI by fear"). So the answer sharpens: an overlay can be absorbed without rewriting levels, but an overlay that scores individuals on tool adoption is read as a ladder whether or not it is wired to one — and here it was wired instead to an unmet enterprise target ("achieving 50% adoption rate among software engineers"), which put the pressure on without putting any criteria in writing. The still-open half is the same one: neither case measures what happens at promotion time.
- SourceDoes a central LLM gateway actually change model-mix decisions, or only report on them? The claimed benefit is portability; no source in the corpus records an org exercising it. Partially answered (2026-09-22): ramp ai index september 2026 records the outcome across ~70,000 US businesses — frontier-model token share falling 52.5% to 44.7% over five weeks while standard-tier share rose 25.7% to 35.8%, with buyers telling Ramp they are "imposing company-wide defaults that reduce usage of frontier models" because standard models are "still highly performant and also more cost effective." Organizations demonstrably do exercise a model-mix switch, deliberately and company-wide, and the switch is large enough to move a market aggregate. The gateway half is untouched: a payment-rail instrument cannot see whether a proxy, a shared-harness config, a procurement rule or a vendor-side default enforces the decision, and the letter names no gateway anywhere. What is left of the question is precisely the mechanism — is a portability layer what makes the switch cheap, or do firms reach the same place without one? A confound also arrived with the answer: effective token prices fell 41% over the same window, so the mix could be moving on price alone. A third mechanism, added 2026-09-22 — and it is a genuinely new axis, because it is neither a gateway nor a price. techcrunch startup arr less secure madrona (
practitioner-opinion, press reporting of an unread VC survey) reports Madrona finding 77% of 150 enterprise IT professionals re-evaluate their AI vendors every six months or on a rolling basis, with Madrona's own gloss that "switching costs are lower and the re-evaluation cadence is relentless." That is a procurement calendar, not a portability layer: it makes switching cheap by making the decision routine and scheduled rather than by making the migration technically easy, and it reaches a vendor-level switch with no infrastructure at all. So the bullet's remaining question — is a portability layer what makes the switch cheap, or do firms reach the same place without one? — now has a live candidate for the without-one branch, and this source cannot adjudicate it either: it names no infrastructure, never asks what a re-evaluation costs to act on, and reports a stated cadence rather than an observed switch. Note also the unit mismatch with the Ramp reading — Madrona counts AI vendors (applications), Ramp counts model tiers, and nothing establishes that the same cadence governs both. - SourceWhat does per-team AI usage analytics get used for once it exists — cost containment, capacity planning, or performance evaluation of engineers? The third would collide with everything Telemetry vs. Survey Measurement establishes about what instrumented output data can and cannot support. Partially answered (2026-09-22): the first of the three is now observed, at firm rather than team granularity. ramp ai index september 2026 reports businesses acting on AI-spend visibility by imposing company-wide model defaults on explicitly cost-effectiveness grounds — cost containment, and the only one of the three uses any source in the corpus has recorded. Two limits keep it open: the granularity is wrong (vendor-level firm spend on a card rail, not per-team usage), and the publisher sells the spend-visibility product, so the framing is marketing surface as well as measurement. Capacity planning and performance evaluation remain entirely unobserved. Extended 2026-09-22 by mckinsey state of ai 2026 road to roi (
empirical, self-reported, n=1,719), which adds the population denominator the Ramp letter cannot: ~20% of respondents report AI operating costs, including tokens, actually constraining their organization's AI use (Exhibit 8), spread 12% to 25% by industry and "broadly consistent" across company sizes. That is the first estimate of how many organizations have reached the point where a cost reading binds rather than merely informs, and it is the cost-containment use observed at a third instrument. Two qualifications keep the bullet open in the same place. The constraint is not a beginners' problem — AI high performers report cost-constrained coding-agent use at 18% against 6% of others (Exhibit 14) — and, more to the point of this bullet, the survey asks about outcomes of cost pressure and never about the instrument: no question mentions a gateway, a proxy, chargeback or per-team analytics, so what an organization looked at before constraining itself is still unobserved. Capacity planning and performance evaluation of engineers remain unobserved at every instrument.
- SourceDoes a central LLM gateway actually change model-mix decisions, or only report on them? The claimed benefit is portability; no source in the corpus records an org exercising it. Partially answered (2026-09-22): ramp ai index september 2026 records the outcome across ~70,000 US businesses — frontier-model token share falling 52.5% to 44.7% over five weeks while standard-tier share rose 25.7% to 35.8%, with buyers telling Ramp they are "imposing company-wide defaults that reduce usage of frontier models" because standard models are "still highly performant and also more cost effective." Organizations demonstrably do exercise a model-mix switch, deliberately and company-wide, and the switch is large enough to move a market aggregate. The gateway half is untouched: a payment-rail instrument cannot see whether a proxy, a shared-harness config, a procurement rule or a vendor-side default enforces the decision, and the letter names no gateway anywhere. What is left of the question is precisely the mechanism — is a portability layer what makes the switch cheap, or do firms reach the same place without one? A confound also arrived with the answer: effective token prices fell 41% over the same window, so the mix could be moving on price alone. A third mechanism, added 2026-09-22 — and it is a genuinely new axis, because it is neither a gateway nor a price. techcrunch startup arr less secure madrona (
- WaitDoes agent-era recentralization (common paved paths, solve-once infrastructure) hold up against the local-team autonomy that Stone credits for Netflix's historical speed — i.e., will local teams accept the paved path when their problem doesn't fit it, or does shadow infrastructure reappear?
- WaitStone keeps AI fluency as a deliberately vague overlay because the tech "evolves by the quarter." Does it ever crystallize into per-level ladder criteria (as conventional competencies did), or is permanent-overlay the stable state? Trigger: Netflix's next ladder revision.
- SourceStone claims specialists can now broaden "quickly" with AI tools. Does the wiki's evidence support cheap breadth acquisition — the concave novice→intermediate curve in Returns to Expertise in Agentic Coding suggests yes for working grasp, but is there evidence on speed of cross-domain ramp for experienced specialists? Partially answered: Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale — split the claim: tool-in-hand performance breadth is measurably cheap (the concave curve makes the needed increment small; the expertise meta-skills — framing precision, verify-specification, who-corrects-whom — transfer across domains, per the management edge, so an experienced specialist enters above the novice floor; AI-assisted onboarding compresses ramp further). Retained-capability breadth is unproven and the only randomized evidence cuts against it: automation-mode gains vanish when the tool is removed while self-report hides the deficit — the augmentation/automation usage split decides which good you get. Ramp speed itself is measured nowhere; the mechanism argument stands in for it. Retagged 2026-10-05: ramp speed for experienced specialists is measured nowhere in the vault; needs a time-to-proficiency study of cross-domain work under AI tools.