Sources#
- Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals
- Deploying AI from pilot to production: A practical blueprint for CIOs and technical leaders
- Startup ARR is less secure than ever, new research shows
- The state of AI in 2026: On the road to ROI
- When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering
Summary#
Deploying AI from pilot to production (Anthropic × Accenture, 2026-09-11, 38pp, vendor-claim) argues that the enterprise AI pilot is a structurally misleading instrument: it succeeds because of conditions that production will not reproduce, so its success is not evidence about the system anyone will actually run.
"That insulation makes pilots successful, but it also makes them an unreliable signal for how enterprise AI will perform in production at scale."
The claim is not that pilots are run badly. It is that the pilot's insulation is the thing being measured, and the fix is to move the production decisions before the pilot rather than after it — which reframes the pilot from a capability demonstration into a readiness exercise. This is the practitioner statement of the thesis Organizational Complements to AI makes from the economics side: the pilot artificially supplies the complements, so the complement gap is invisible until the complements are withdrawn.
The insulation catalog#
The document's most transferable contribution is the itemization — what specifically makes a pilot unrepresentative. Six conditions, stated as the reasons pilots succeed:
| Pilot condition | What production supplies instead |
|---|---|
| Curated datasets, prepared for the exercise | The real data estate: distributed, inconsistently formatted, governed by rules not designed for AI, owned by teams on other timelines |
| Handpicked, AI-native engineering teams | Standard delivery teams, "not representative of the broader enterprise workforce" |
| Protected budgets, insulated from normal org dynamics | Steady-state operational economics |
| Narrow, defined scope with clear timelines | Fragmented enterprise systems and bidirectional integration |
| A high level of manual intervention | Whatever the system does unattended |
| Limited exposure to downstream stakeholders | IT, Legal, Finance, HR and Compliance, each with a different definition of "working" |
The fifth is the one that generalizes furthest and is named explicitly as a design rule — "minimizing hidden manual intervention". A pilot that quietly depends on a human repairing inputs measures a hybrid system whose human half is not in the production budget. It is the enterprise-deployment instance of the failure class Failures That Look Like Success names: the demo reads as working, and the part that made it work is not in the artifact.
The prescription is to build the production conditions into the pilot design: define production success criteria upfront, test against representative enterprise conditions, minimize hidden manual intervention, validate delivery scalability beyond AI-native specialists, embed governance and observability early, and measure long-term operational economics rather than short-term model performance.
Front-load the decisions#
The structural argument is about ordering, not content — every decision deferred past the pilot is claimed to compound:
"Every decision deferred past the pilot creates engineering debt: parallel systems, integration patterns that never standardize, and a platform foundation that's always being renegotiated."
Its sharpest illustration is the infrastructure hedge: an organization that cannot settle build-versus-buy runs the pilot on a managed API while a platform team builds in-house "for when we need more control", and six months later maintains two systems with every new use case reopening the debate. The document's claim is that most infrastructure problems come from never committing, not from committing wrongly — which is the same optionality argument Standardize the Infrastructure, Not the Tools resolves the other way, by making the substrate the committed layer and leaving tool choice deliberately open.
The deployment blueprint#
Seven decisions across four lifecycle stages, each with a named owner. Four of the seven land before the pilot begins, and the pilot stage itself introduces no new owners or decisions — it only tests the decisions already made.
| Stage | # | Decision | Scope | Owner |
|---|---|---|---|---|
| Pre-pilot — decide before the pilot begins | 01 | Strategy and ownership | Success criteria, ROI thresholds, go/no-go gates | Executive sponsor |
| 02 | Data and integration | Audit the data estate, assign source and pipeline owners | Data and IT change owners | |
| 03 | Infrastructure | Align platform choices to use-case requirements | CIO / platform lead | |
| 04 | Security and trust | Classify data; clear security, regulatory and access requirements | Security and compliance lead | |
| Pilot — validate, don't defer | — | No new owners or decisions | The pilot tests the decisions already made | — |
| Preproduction — prepare the organization | 05 | Org readiness | Sequence rollout, fund reskilling, name champions | Exec sponsor / change lead |
| Production — govern by risk | 06 | Governance and risk | Risk taxonomy, monitoring thresholds, review authority | Risk and monitoring owner |
| Scaled deployment — operate as a capability | 07 | Scale and evolution | Start simple, scale deliberately; understand failure before scaling | AI program owner |
(Source PDF p.35. This page renders as page graphics, so both docling and pdftotext recover nothing from it — the table above is a manual transcription made at ingest from the rendered page; see the parse note in Sources.)
Each consideration closes with a two-column question set the document splits by decision type: "Work out" questions need cross-functional input and have multiple owners; "Assign" questions are ownership decisions a CIO or business leader must make alone. The split is the document's operational core — it is a claim about which deployment questions are deliberative and which are executive, and treating an Assign question as a Work-out question is how ownership diffuses.
Ownership decomposes into three non-substitutable parts#
The named owner needs decision rights (make calls without convening a meeting), escalation authority (blockers reach senior attention in days, not months), and executive backing (signals institutional weight, which determines whether other functions engage as partners or observers).
"Decision rights, escalation authority, and executive backing aren't interchangeable, and the absence of any one of them is enough to stall a program."
The failure mode is measured, after a fashion: Accenture's September 2026 Tokenomics research reports that 42% of organizations rely on shared IT and finance accountability with no single owner responsible for AI costs and outcomes. Contrast AI Employee Framing, which finds the accountability leak coming from the framing of the agent rather than from the org chart (−9pp personal accountability, +44% escalation when an agent is called an "employee") — two independent routes to the same unowned outcome, one structural and one linguistic.
Deployment is not adoption#
The document separates shipping the system from the workforce using it, and is blunt that leadership pressure alone does not close the gap:
"A portfolio manager showing a compliance specialist how she summarized a 200-page filing in three minutes converts more skeptics than any structured rollout."
The prescription is to engineer the demonstration moments deliberately rather than wait for them — identify champions, clear time to experiment, recognize early adopters publicly — which is the same mechanism Standardize the Infrastructure, Not the Tools reports from Shopify (adoption by demonstration, not mandate), arrived at independently.
As adoption spreads, the claim is that work shifts from executing tasks to overseeing a system: identifying incorrect outputs, deciding what escalates, and detecting drift before it becomes a production issue. That shift is asserted, not measured, and it sits directly against the mechanism The Tragedy of the Cognitive Commons describes — if the entry-level execution work is what regenerates the judgment the oversight role requires, an org that converts everyone to overseers has removed the training path for the skill it now depends on. The document does not raise this.
The copilot throughput ceiling#
The scale-stage argument for end-to-end automation over assistance is a clean statement of where assistive AI stops paying:
"The throughput ceiling on a copilot is still the person using it. A tool that helps an analyst write reports faster is still bounded by how many reports that analyst can review and approve."
Only when the AI completes the work — invoices processed end-to-end with exceptions escalated, tier-one tickets from open to resolution — does "throughput scale with volume, not headcount". This is the enterprise-workflow statement of the shift Conversation-to-Delegation Shift measures in developer tooling (99.8% / 63.3% / 16.5% delegated-output share across three populations), and the tension it inherits is the same one Verification as the New Bottleneck names: removing the human from the loop moves the bound from production to verification rather than removing it.
Paired with a counterweight the document states plainly — start simple. A well-designed prompt is fast to test with predictable failure modes; a multi-step agentic system is more capable but harder to debug, more expensive to maintain, and fails in less anticipable ways. Add complexity only when the simpler approach has demonstrably hit its limits. (The document cites Anthropic's own Building effective agents for the same argument.)
And a readiness test that is genuinely sharper than an accuracy number: know the error mode distribution, not just the error rate. "A system that's 95% accurate at limited volume sounds production-ready. But the remaining 5% matters" — a formatting slip a reviewer catches in seconds may be safe to automate; an incorrect financial calculation or a missed compliance flag makes every additional unit of volume add risk. The operative question is what each type of error costs at full production volume. That is the consequence-tiering discipline AI-Assisted Error Analysis applies to eval investment, applied instead to the automation decision.
The figures, and how far they carry#
Every quantity here is vendor-published. The surveys are at least named and dated; the deployment anecdotes are not verifiable at all.
Attributed surveys (Accenture, self-published, methodology not included):
| Figure | Source |
|---|---|
| 23% of C-suite leaders report sustained, enterprise-wide AI impact | Pulse of Change, July 2026 |
| 64% have moved past pilots into production across multiple functions or begun enterprise-wide efforts — but only 7% have the data readiness to scale advanced AI | AI-Ready Data for Advanced AI, May 2026 |
| "Data reinventors" realize EBIT margin uplift of up to 1.6× over industry peers | AI-Ready Data for Advanced AI, May 2026 |
| 42% rely on shared IT/finance accountability with no single AI cost-and-outcome owner | Tokenomics, September 2026 |
| Organizations with formal chargeback accountability link 32¢ of every dollar of AI token spend to a quantified business outcome — 6× those with no allocation | Tokenomics, September 2026 |
Third-party: Gartner predicted (February 2025) that organizations would abandon 60% of AI projects unsupported by AI-ready data through 2026 — a prediction whose window closes this year and which the document cites without checking against outcome.
The 64% vs 7% pair is the load-bearing one, because it is the only figure that separates being in production from being able to scale, and it is the arithmetic behind the whole document. It should be quoted with the caveat that both halves come from one vendor's own survey of its own prospective buyers, and that "data readiness" is a threshold that vendor defines.
Unattributed anecdotes — no company, no methodology, no measurement protocol: a global insurer compressing underwriting review "more than 5×" while improving data accuracy "75 to 90%"; a pharmaceutical company cutting clinical study report production "from ten weeks to under ten minutes"; a telecom deploying "more than 13,000 custom AI solutions" in under a year. The named customer cases are thinner still as evidence, being Anthropic case-study links with a quoted executive: StubHub (30% support-cost reduction, response times from over 20 minutes to near-instant, after side-by-side model A/B testing); Novo Nordisk (the ten-weeks-to-ten-minutes claim, attributed); TELUS (Fuel iX multi-model platform, 13,000 custom solutions, 500,000+ hours saved, 100 billion tokens monthly, "Claude became the overwhelming choice"); Palo Alto Networks (20–30% increase in feature development velocity); NBIM (600+ active users within two months, an AI Ambassador Network of 50 specialists, internal surveys showing 20%+ weekly time saved).
Read these as existence proofs, not effect sizes. The selection is the vendor's, the counterfactual is absent everywhere, and in the two cases where the same number appears both as an anonymous statistic on p.3 and as a named case later (Novo Nordisk, TELUS), the anonymous framing implies a breadth of evidence the named cases do not supply.
What this document does not do#
Three things a reader should not expect from it:
- No failure analysis. Every case is a success; the ~77% of programs the document's own headline statistic says have not reached sustained impact are characterized only by which considerations they skipped, which is an assertion of the thesis rather than evidence for it.
- No priors on the decisions themselves. The blueprint says who decides and when, never what to decide. Build-versus-buy, single-versus-multi-model, cloud-versus-on-prem are all resolved into "commit deliberately" — which makes the document a process artifact, not a technical one. The one exception is the retrieval-architecture note below.
- No cost of the process. Front-loading seven decisions before a pilot has a price in calendar time and executive attention, and it is not estimated anywhere — which matters, because the prescription's own premise is that "AI is compressing that timeline into months or even weeks".
And one thing the co-author has since done with the gap rather than about it: three months later Accenture and Google Cloud announced a 1,000-person forward-deployed-engineer workforce for Gemini Enterprise (Accenture and Google Cloud Deepen Partnership with Formation of New Accenture Gemini Enterprise Business Group, 2026-09-08, vendor-claim) — the same gap sold as staffed delivery capacity instead of a decision discipline, which is the commercial alternative to this document's prescription and is treated at Forward-Deployed Engineering as a Delivery Layer.
The one architectural claim#
Consideration 03 carries the document's only substantive technical recommendation, and it is a subtractive one: audit whether the retrieval layer is solving a current problem or an expired one.
"Many existing pipelines are solving for a context window constraint that no longer applies."
Chunking, pre-send summarization and staged retrieval were workarounds for small context windows; modern models process whole contracts, codebases or research reports in one pass. The document allows that retrieval still earns its place where data freshness or access control are the drivers — which is the same residual Document Parsing as the Retrieval Bottleneck identifies from the opposite direction (long context did not kill RAG; cost, governance and audit keep it), and the same expiry logic Context Window Smart Zone complicates, since usable context is measured well below the advertised window.
The in-window population datum (McKinsey, August 2026)#
McKinsey's 2026 state-of-AI survey (1,719 respondents in 97 nations, fielded May 4 - June 8 2026, GDP-weighted, empirical but wholly self-reported) is the only population-scale reading this vault holds from inside the window Gartner's abandonment prediction covers, and the nearest thing to an outside check on the 64%/7% pair above. Phase of AI use among organizations using AI (Exhibit 1):
| Phase | 2025 | 2026 |
|---|---|---|
| Experimenting | 32 | 22 |
| Piloting | 30 | 34 |
| At least scaling | 38 | 44 |
with 89% reporting regular AI use in at least one function and 56% in three or more (up from 51%). The movement is forward across the whole distribution: the experimenting share fell ten points and both later phases grew. Whatever else 2026 was, it was not the year the cohort walked away.
The survey also restates this document's central problem in its own numbers, and more starkly. 80% of respondents say AI improved their individual productivity; 37% say AI contributed positively to their organization's EBIT, essentially unchanged from 2025 - while the scaling share rose six points over the same year. Deployment breadth grew and financial attribution did not. The gap the blueprint exists to close is, at population scale, not closing, which is the strongest available argument that the bottleneck is the organizational change this document prescribes rather than the deployment itself.
The separating variables it publishes are the practice gaps between AI high performers - the 6% of respondents attributing at least 5% of EBIT to AI and calling the impact significant, a share unchanged from 2025 - and everyone else (Exhibit 11, a dumbbell chart carrying no printed numeric labels; every value below is read off gridlines and is approximate):
| Practice | All others (~) | High performers (~) |
|---|---|---|
| Fundamentally redesigned workflows because of AI | 25 | 72 |
| Expects to use AI for enterprise-wide transformative change | 15 | 62 |
| Senior leaders demonstrate true ownership and commitment | 32 | 65 |
| Spends >15% of IT budget on AI | 16 | 35 |
| Determined how and when outputs need human validation | 25 | 35 |
| Actively manages AI solution costs (tokens, compute, storage) | 25 | 34 |
| Has a semantic layer / knowledge graph for company data | 13 | 26 |
| Defined processes to quantify AI initiative impact | 17 | 25 |
| Completed a strategic workforce-planning exercise | 13 | 24 |
| Transformation office with authority to remove bottlenecks | 20 | 22 |
Two readings, in order. First, the cohort is defined by its outcome, so every row is an association inside a group selected on the dependent variable - this is a ranking of what travels with reported success, not a causal blueprint, and one respondent per organization self-reports both the practice and the result. Second, the ranking's shape is the usable part and it lines up with this document's own ordering: the two widest gaps are workflow redesign and transformative ambition (the blueprint's considerations 01-02), leadership role-modeling is third, and the data-layer row is among the narrowest. Exhibit 10 says the same thing from the objectives side: high performers pursue growth (74% vs 47%) and innovation (65% vs 49%) in addition to efficiency, where the efficiency objective itself is flat across cohorts (78% vs 82%). Nearly three-quarters of high performers report fundamentally redesigning workflows, up from 55% a year ago.
A stage model for the gap, and the boundary of its evidence (NANTE, August 2026)#
Adoption Telemetry (Damon A. Young, PolyWise Partners; arXiv 2608.23617, 2026-08-22, 23pp, practitioner-opinion) is the first source in the vault that takes this page's diagnosis as settled and asks what instrument would locate the failure. Its framing: the cause of enterprise AI failure — organizational, not technical — "is no longer contested," and "a substantial vocabulary has grown up around the phenomenon: pilot purgatory, pilot fatigue, the GenAI divide, AI theater… What has not followed the diagnosis is an instrument."
Provenance before content, because this one matters. Single author, a consultancy (PolyWise Partners) that would deploy the instrument, closing with a design-partner invitation. No figure in the paper comes from a real organization. Every result is generated by the author's own simulator from a fixed seed, and the paper says so at length — see the circularity admission below. The failure statistics it opens with are all secondhand: MIT NANDA's 95%-no-P&L-impact (which the paper itself flags as a v0.1 preliminary report "whose methodology counts vary across secondary accounts," cited as directional only), S&P Global's rise from 17% to 42% in the share of companies abandoning most of their AI initiatives in a single year, and Gartner's prediction that over 40% of agentic projects will be cancelled by end-2027 (the paper notes Gartner publishes no derivation for the 40%). None was checked against its underlying release here.
NANTE — from the Twi nante, "walk" — is a five-stage model, each stage a computable predicate over an event stream of invocations with session ids, per-session turn counts and task outcomes, joined to a provisioned-user roster. The roster is the structural move: "a store containing only events renders Notice invisible, because the users who never arrived generate no rows — the population that most needs counting is exactly the one that leaves no trace."
| Stage | Question | Proposed threshold |
|---|---|---|
| Notice | does the user know it exists and applies to them? | provisioned, no qualifying activity |
| Attempt | have they tried it at all? | first invocation |
| Navigate | has trial become recurring use? | activity in ≥ 3 distinct weeks |
| Transform | has the shape of the work changed? | ≥ 6 active weeks and ≥ 60% of sessions multi-step and ≥ 55% task success |
| Embed | is withdrawal now disruptive? | activity in ≥ 90% of weeks since first use |
Three design choices carry the model. Notice/Attempt/Navigate are breadth; Transform/Embed are depth, and the published evidence puts the characteristic enterprise failure exactly at that boundary. Transform and Embed partition one qualifying pool rather than stacking as filters — an Embed classification implies Transform's criteria, and continuity is what splits them, "so that historical volume cannot substitute for present integration." And "multi-step" is proxied by sessions of ≥ 12 turns, which the paper calls "an honest proxy for workflow depth in an event schema that carries no task-type field" and then lists as a known failure mode: "a user deeply integrating a tool and a user struggling with it both produce many turns." The ≥ 55% success gate is a deliberate addition, because without it "a cohort of frequent, multi-step, but mostly failing users would classify as transformed."
The honest-ceiling design is the most interesting judgment in the paper. Stalls are found by walking the boundaries in order for the first graduation failure. At the breadth boundaries, ≥ 90% graduation reads healthy and < 80% failing. At the two depth boundaries the cutoffs are set deliberately outside the achievable range: the healthy floor is placed above 1.0 — a fraction no cohort can reach, so none ever reads "healthy" there — and the at-risk floor at 0.01, below which a cohort reads failing for falling under the ~2% workflow-integration baseline the published evidence reports. The reasoning is that progression past Navigate "is rare everywhere… an industry-wide cliff, not a cohort-specific defect," and an alarm calibrated to an aspiration nobody meets is an alarm that is always ringing. This is a measurement instrument conceding at design time that the gap this page describes is the field's normal condition.
Four failure signatures are computed as flags beyond the stall point, each with a not-this intervention contrast — the paper's rationale being that "misdirected investment is the modal enterprise failure":
- Shallow plateau (divergence index > 0.3 — retention high, session depth flat): a workflow-fit failure. Redesign the workflow so the tool sits inside the task sequence — not more licenses. "This is the signature of the cohort every usage dashboard reports as successful."
- Low task success (≥ 75% of a cohort's Navigate-classified, tenure-eligible users failing the success gate while ≤ 90% fail the multi-step gate, computed only at a Navigate stall with ≥ 20 such users): skill-building and enablement — not workflow redesign. "The population has already brought the tool into deep work; the work is failing, not absent."
- Champion dependency (Gini > 0.5 over per-user activity among active users): deliberate de-concentration — not celebrating the power users, whose concentration is the risk.
- Usage regression (late-window invocation rate < 40% of an established early-window rate): leadership sponsorship and re-integration — not re-onboarding; "the population once knew how; something stopped rewarding the behavior."
Where a stall and a signature disagree, the signature wins: "the stall locates the population's current position, but a signature identifies the process that produced it." A Notice stall maps to visibility and access, not training — with the honest caveat that "a zero-activity row is behaviorally ambiguous between a user who never learned the capability exists and one who knows and has declined it."
The composite score exists to be distrusted. Stage weights 0/25/50/75/100 reduce the distribution to 0–100, "reported alongside — never instead of — the distribution, the stall point, and the flags," and the paper's own Table 1 is the argument: the shallow-plateau and ability-gap cohorts share an identical stage distribution and an identical score of 49.9 with opposite diagnoses (shallow-but-succeeding vs deep-and-failing), while awareness-gap and reinforcement-decay score within 0.6 points (41.0 and 41.6) despite being a population that never arrived and one that arrived and left. Below 60 post-rollout days or 50 users the instrument reports unavailable rather than computing.
What the evaluation is, and what the paper says it is not. Six synthetic profiles — one healthy, five pathological — each a reference cohort of 167 provisioned users observed over 149 days, a single fixed seed, the default thresholds, run as the repository's CI test suite (agent-adoption-kit, Apache 2.0). The headline pair is the one the framework exists to separate: 98%+ of both rosters progress into recurring use, and they diverge only in depth. The healthy cohort distributes upward (73.7% Navigate, 7.2% Transform, 18.0% Embed; depth conversion 25.1%, score 60.5, no stall); the shallow-plateau cohort puts 99.4% at Navigate with none beyond (depth 0.0%, score 49.9, Navigate stall, divergence index 0.67 against the 0.30 threshold). "A usage dashboard reporting active users would rate the second cohort marginally higher than the first." The discriminating contrast that makes the taxonomy non-trivial: the ability-gap cohort has the same 99.4%-at-Navigate distribution but 34.5% task success against the healthy cohort's 69.5%, while the shallow-plateau cohort's success is 69.6% — "shallow stalling is not failing, and failing is not shallowness," and the two flags never co-fire. (Table 1 reconciled cell-for-cell against pdftotext -layout, exact; the healthy row's 25.1% depth is computed from unrounded fractions and differs by 0.1 from its displayed 7.2 + 18.0, per the paper's own footnote.)
Then the disclaimer, which is unusually complete and should be quoted rather than paraphrased: "the synthetic pathologies and the proposed thresholds were developed together, so the evaluation shows that the instrument separates the failure modes it was designed to separate — a necessary condition and a genuine engineering result, but not evidence that real cohorts classified by these thresholds are correctly diagnosed." One profile is explicitly non-representative by construction: the healthy cohort's 25.1% depth conversion "sits roughly an order of magnitude above the ~2% workflow-integration rate the published evidence reports, and is a constructed upper-contrast reference rather than a field-typical population." So the vault's reading: NANTE is a specification and a taxonomy, not a finding. Nothing on this page is superseded by it, and no number in it is usable as a claim about any real organization.
Two places it bites this page's own arguments. First, on the throughput-ceiling argument for end-to-end automation: the instrument inverts exactly where that prescription leads. For autonomously operating agents "the framework's central signal inverts: deepening adoption reduces human-initiated invocation, so an organization succeeding at delegation and one abandoning the tool produce the same declining curve, and NANTE as specified would diagnose the first as reinforcement decay and prescribe reengagement." The paper treats this as a genuine boundary rather than a patchable defect — a fully delegated workflow is "an organization transferring ownership of a process," not a population progressing through a change. An org that follows this page's automation prescription therefore loses the adoption instrument at the same moment. Second, §8.5 specifies the disposition extension — five states per invocation (completed autonomously, reviewed and accepted, overridden, escalated, abandoned) — and names four failure signatures the current schema cannot see, including rubber-stamping ("high completion, few escalations, no overrides"), which is "the one the present instrument would score most favorably." That is Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?'s question restated as a telemetry schema, and the paper declines to simulate it: no published base rate exists for agentic disposition, so parameterizing one "would mean inventing the parameters and then verifying that our instrument recovers our invention." An honest refusal, and a real gap.
One regularity worth carrying forward, also unvalidated: §8.4 observes that Transform-grade telemetry exists only where organizations built genuine agents rather than distributing seats — seat-assistant admin exports carry breadth and continuity but no task outcomes, leaving Transform's gates uncomputable — "so the availability of depth measurement is itself a coarse indicator of deployment maturity."
Connections#
- The Enterprise AI Adoption Gradient — this page's thesis with a coefficient behind it. OpenAI's enterprise study closes on "adoption is only the beginning of deployment," and decomposes the gap: among ChatGPT Enterprise adopters, larger firms show sharply lower usage per employee (WAU/emp −0.032***, tokens/emp −0.667***) while messages per active user is flat (−0.002, n.s.). The deployment shortfall is entirely breadth — a smaller share of the workforce ever becomes active — not weaker use by those who do
- Telemetry vs. Survey Measurement — where the instrument argument lives. That page carries the same source's claim that the four measurement traditions (observability, usage dashboards, product analytics, change management) each instrument something adjacent to adoption, the ActivTrak 82%-vs-~2% definitional swing inside a single telemetry dataset, and the neutrality argument for why no platform vendor builds a portfolio-neutral adoption instrument. The split of labour: the stage model and its evidence boundary are here because they are claims about deployment failure; the claim that a production log can carry a change-management construct at all is a claim about instruments and lives there
- Usage-Telemetry Classifier Validation — the measurement cost of NANTE's design, on the page that prices the alternative. The stage predicates above are deterministic counts with no LLM anywhere in the measurement path, so there is no classifier accuracy to publish — but the semantic judgment does not vanish, it moves into the 12-turn multi-step proxy and the deployment-reported success label, neither of which anyone validates. A published-but-unvalidated threshold against an unpublished-but-validated classifier: two different failures, and the genre has no instrument that is both
- Layered Supervision — a different destination for consideration 05's conversion. This document turns the operator into an overseer of the same artifact (the invoice auditor now audits the invoice-auditing system); that source reports oversight re-scoped upward instead, off the artifact entirely and onto architectural reasoning, with line-by-line comprehension abandoned in favour of keeping the system diagnosable on demand. It also names the staffing failure this blueprint does not: adopting widely without proportionate senior supervisory capacity accumulates risk invisible from inside the team
- Organizational Complements to AI — the economics of the same claim. That page's thesis is that AI productivity gains depend on complementary workflow, skill and org-design changes; this document is the practitioner-side account of why the dependency stays hidden, since a pilot supplies the complements artificially (curated data, AI-native staff, protected budget) and therefore cannot measure the gap it will hit. The 64%-past-pilots-vs-7%-data-ready split is a complement gap stated as a survey statistic — with the caveat that it is one vendor's survey, where that page's core evidence (the Codex three-population natural experiment) is not
- Risk-Tiered Auto-Approval — consideration 06 is a fourth tiering key, keyed to the consequence of the output rather than the change, the anomaly, or the reversibility of the response: a four-tier oversight ladder (automated / sampled / reviewed / advisory) with review cadences attached, plus the checkpoint-audit discipline and the "structural drag" account of why review layers accumulate. Detailed on that page
- Standardize the Infrastructure, Not the Tools — the same cost-governance layer, with the chargeback numbers this document supplies (32¢ of every AI-token dollar tied to an outcome under formal chargeback, 6× those with none; 42% with no single owner) and the same adoption-by-demonstration mechanism. But the two resolve the commitment question oppositely: this document treats an uncommitted build-versus-buy as the primary source of engineering debt, where Shopify deliberately buys optionality by standardizing only the substrate
- Failures That Look Like Success — "minimizing hidden manual intervention" is this failure class located in the deployment pipeline rather than in an agent trajectory: the pilot reads as working, and the human repair that made it work is neither in the artifact nor in the production budget
- Conversation-to-Delegation Shift — the copilot throughput ceiling is this shift argued prospectively for enterprise workflows, where that page measures it retrospectively in developer tooling. The asymmetry is worth noting: this document asserts that throughput scales with volume once the human leaves the loop and supplies no instance where it was measured, while that page has the delegated-output shares and no claim about the ceiling
- AI Employee Framing — the second route to unowned AI outcomes. This document locates the accountability leak in org structure (42% with no single owner) and prescribes a named owner with three non-substitutable powers; Kropp et al. locate it in language, and find that calling the agent an "employee" costs −9pp of personal accountability with no adoption gain. Both failures survive the other's fix
- The Tragedy of the Cognitive Commons — the unraised tension. The document's consideration-05 prescription is that work shifts from executing tasks to overseeing a system, requiring "judgment-based skills"; Lovett's argument is that those skills regenerate through exactly the entry-level execution work being automated away. The document treats the oversight capability as a training investment; that page treats it as a commons with a removed regeneration mechanism. Neither cites the other, and this is the sharpest thing the vault can say against the blueprint
- Document Parsing as the Retrieval Bottleneck — the same verdict on retrieval's remaining job from the opposite approach: this document argues subtractively (audit whether the pipeline solves an expired context-window constraint), that page argues from the 2024→2026 retrospective that long context did not kill RAG because cost, governance and audit outlive the constraint. They agree on the residual — freshness and access control — and disagree on nothing
- Verification as the New Bottleneck — where the end-to-end automation argument runs out. The document's error-mode-distribution test (what does each type of error cost at full volume, not what is the error rate) is a deployment-side statement of that page's problem, and its escalate-only-the-exceptions design assumes the exception detector is the cheap part
- AI-Assisted Error Analysis — the same consequence-first reasoning applied to a different decision: that page sets eval investment by reasoning through worst-case user outcomes up front, this one sets the automation boundary the same way. Both reject a uniform bar across features
- Anthropic — publisher, with Accenture. The document's "Getting started" section is a reading list of Anthropic product and engineering material (Claude Enterprise Administrator Guide, Scaling agentic coding across your organization, Claude Cowork enterprise controls, Building effective agents, Effective context engineering, MCP), which is the clearest single indicator of what the artifact is for
- Psychological Costs of AI Adoption — what consideration 05's conversion costs the converted. That case study is the vault's only deployment where an oversight workforce has been running long enough to describe the job: engineers who now read AI-authored code they remain accountable for, one year in. The conversion is not a training problem and it is not free — retained accountability without retained authorship generates a verification bill ("if we weren't responsible for the code it produced, it would be a lot faster"), and for the roles whose craft was the displaced execution work it also generates identity and meaning losses that oversight training does not touch. The blueprint prices the capability and not the transition
- Build Instead of Buy Under Agentic Coding — the one decision this document resolves into "commit deliberately" and that the population is now making without it: 32% of a 1,719-respondent panel report declining at least one software purchase because agentic coding could supply the feature in-house. The caution runs both ways — a decision not to buy is exactly the sort of commitment a pilot's insulated success conditions can produce and production cannot honor
- Forward-Deployed Engineering as a Delivery Layer — the same gap monetized rather than designed away: an embedded-engineer workforce is what a buyer purchases when it declines the front-loaded-decision discipline, and the corpus's largest instance is staffed by this document's own co-author
- Claude Code — named as the worked example for the leading-vs-lagging indicator argument: PR cycle time is a leading indicator of value from a developer using Claude Code, but business value is the lagging indicator (revenue, net revenue retention, churn) and must be attributed to what shipped
Open Questions#
-
Gartner's February 2025 prediction — organizations abandon 60% of AI projects unsupported by AI-ready data through 2026 — reaches its window this year, and this document cites it without checking it. Did the abandonment rate land near 60%, and was data readiness the separating variable it was predicted to be? A resolved forecast would convert the document's central data-readiness claim from a vendor survey into something gradeable. Partially answered (2026-09-22) by The state of AI in 2026: On the road to ROI (McKinsey/QuantumBlack,
empiricalbut self-reported, 1,719 respondents in 97 nations fielded May 4 - June 8 2026 — squarely inside the prediction's window, and the strongest in-window population datum the vault holds). It settles the direction and leaves the rate untouched, and the two halves of the question come apart. On abandonment: nothing in the aggregate looks like a 60% walk-away. The share of AI-using organizations stuck in experimenting fell 32% to 22% year over year while piloting rose 30% to 34% and at-least-scaling rose 38% to 44%; 89% report regular use in at least one function, 56% in three or more. But the survey never asks about abandoned, cancelled or paused projects, so no abandonment rate of any kind exists in it — and the unit is wrong for the forecast either way: the phase question is asked of the organization, so a firm scaling in two functions can have killed five projects in others with the instrument seeing none of it. Gartner's number is a rate over projects; this is a distribution over organizations, and no arithmetic converts one into the other. On data readiness as the separating variable: the only data-layer practice published is whether the organization has a semantic layer or knowledge graph, ~26% of high performers against ~13% of all others (Exhibit 11, gridline-read, approximate). A 2x gap, and one of the smallest in that chart — against ~72% vs ~25% on workflow redesign, ~62% vs ~15% on transformative ambition and ~65% vs ~32% on senior-leader role modeling. So on this instrument data readiness separates, but three organizational variables separate harder, and the prediction's implied causal privilege for the data layer is not what the population shows. Both discounts stand: the high-performer cohort is defined by its outcome, and the whole survey is self-report from a panel McKinsey fields and whose AI transformation work McKinsey sells. What would still close this bullet is an instrument that counts projects — a portfolio census or a vendor-side cancellation series — and none exists in the corpus. One more datum, in the right unit but the wrong weight (2026-09-22): Startup ARR is less secure than ever, new research shows reports Madrona's survey of 150 enterprise IT professionals finding fewer than half of their AI pilots ever reach full production — a project-level rate, which is Gartner's unit and not McKinsey's, but at n=150 from a VC survey reaching the wiki through a news article (practitioner-opinion, primary unread), and measuring non-promotion rather than the abandonment Gartner forecast; a pilot that never ships is not necessarily a project that was killed. It is directionally compatible with a 60% attrition rate and cannot distinguish it from a 40% one. A third unit, secondhand (2026-09-22): Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals cites S&P Global Market Intelligence for the share of companies abandoning most of their AI initiatives rising 17% → 42% in a single year. That is abandonment, unlike McKinsey's phase distribution, but it counts companies-abandoning-most rather than projects-abandoned, carries no data-readiness split, and reaches the vault third-hand inside apractitioner-opinionpaper that itself cites it as "directional evidence of the failure consensus" and nothing more. The project census this bullet needs still does not exist in the corpus. -
The blueprint's load-bearing claim is an ordering claim: decisions made pre-pilot cost less than the same decisions made post-pilot. Nothing in the document measures this, and the counterfactual is available in principle — two programs, same use case, differing only in whether ownership, data audit, infrastructure and security were settled before the pilot. Does front-loading actually shorten time-to-production, or does it move the same calendar time earlier and add executive-attention cost the document never prices? Extended 2026-09-22 by The state of AI in 2026: On the road to ROI, which bears on this only negatively and is worth recording so the next reader does not re-derive it. Exhibit 11 measures the presence of ten practices in a cross-section; it never measures when any of them was adopted. Two of its widest gaps are this blueprint's pre-pilot decisions restated — transformative ambition (a strategy decision, ~62% vs ~15%) and workflow redesign (~72% vs ~25%) — which is consistent with the ordering claim and exactly as consistent with its converse, that organizations which got value went on to acquire the practices. A survey fielded once cannot break that tie, and this one publishes no adoption dates, no time-to-production, and no cost of the decision process. The counterfactual the bullet asks for remains unavailable.
-
Consideration 05 prescribes converting operators into overseers ("the claims processor who used to review every invoice now audits the system that reviews them") while The Tragedy of the Cognitive Commons argues the judgment that role needs is regenerated by the execution work being removed. Is there a deployment where an oversight workforce was staffed without an execution cohort beneath it, and did its error-catching hold over more than a year? This is the falsifiable core of the disagreement and neither source supplies it. Partially answered (2026-09-22) by The Psychological Costs of Artificial Intelligence Adoption in Software Engineering (
case-study, N = 21 interviews at a 1,200-person regulated-software firm, exactly one year after adoption launched): it supplies the deployment and one half of the answer. Engineers there were converted into overseers of AI-authored code while their accountability for it was left unchanged, and the conversion held — the overseers do catch things, by reading "the result line by line" and by applying heightened scrutiny specifically to AI-authored contributions in peer review. But the conditions are the ones that make the finding not transfer to the question's premise. The oversight cohort is entirely ex-executors: 19 of 21 participants have a decade or more of industry experience, median 20 years, so the judgment being exercised was accumulated before the execution work was removed, not regenerated after. The prediction survives untested, and the cohort itself names the mechanism unprompted — fears of "skill atrophy" that might leave practitioners "unable even to critically review AI-generated solutions," countered privately with deliberate manual-coding rituals nobody asked them to perform. Two further gaps: the study measures no error-catching rate of any kind (it is interview evidence about experience, not an accuracy instrument), and one year is inside, not past, the window the question asks about. What it does settle is that the conversion is not costless even when it works — the overseers pay a verification tax the organization neither measured nor budgeted. A second axis added 2026-09-22 by When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering (case-study, five practitioner interviews, ESEM 2026 Software Engineering in Practice), which does not answer the question but changes what it is asking. Consideration 05 converts the operator into an overseer of the same artifact — the invoice auditor now audits the invoice-auditing system. These five teams moved oversight somewhere else: off the artifact entirely and up to architectural reasoning, contextual appropriateness and long-term maintainability, with complete line-by-line comprehension explicitly abandoned and replaced by operational explainability, the requirement that a generated system stay diagnosable and reconstructable on demand through abstraction layers, visualizations, logs and higher-level interpretive tooling. That distinction bears directly on the cognitive-commons disagreement this bullet is about: an overseer who audits the artifact needs the judgment the execution work used to regenerate, while an overseer who audits the architecture needs a judgment the execution work never supplied in the first place — so the two positions may not even be disagreeing about the same workforce. The same source supplies the staffing prediction the blueprint omits, and it points the pessimistic way: teams that adopt AI tools widely "without proportionate senior supervisory capacity may accumulate risk not visible from within the team", and the participants who used executable constraints most heavily were the ones with the strongest architectural expertise. The gaps are the same ones as before and one worse: five interviews, no error-catching rate of any kind, no longitudinal arm, and a cohort that is again entirely senior (architect/CTO, staff manager, team lead, two senior engineers), so the regeneration prediction is untested for a third time. A third axis added 2026-09-22 by The state of AI in 2026: On the road to ROI (empirical, self-reported, n=1,719): the first population-scale, role-level reading of the cost, on a panel three orders of magnitude larger than either case study and with the executive layer included as a comparison group. Exhibit 5 cuts perceived AI impacts by organizational level, and the strains sit below the executive line — 47% of midlevel managers and individual contributors report at least one negative effect, against 31% of executives and senior managers. The item closest to this bullet's mechanism is "hurts my ability to think critically": 20% of midlevel managers and 16% of individual contributors, against 9% of C-level and 8% of executive/senior managers. Career anxiety splits the same way (19% / 18% vs 6% / 10%), as does mental fatigue (14% / 7% vs 7% / 7%). Three things this is not: an error-catching rate, a longitudinal measure, or an observation of anyone's judgment — it is one-shot self-perception, and the survey has no oversight-workforce question at all. What it adds is that the cost the 21-person interview cohort named privately is distributed across 1,719 respondents in 97 nations, and that it is systematically not felt by the executives who commission the conversion — which is a structural reason the blueprint prices the capability and not the transition. One counter-row in the same exhibit keeps this honest: midlevel managers also report "think more clearly" at the highest rate of any level (54% vs 39 / 33 / 37), so the cut is a divergence in experience, not uniform pessimism below the C-suite. #oq/source A third position on the same disagreement, with no deployment behind it (2026-09-22): When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration (CIVIC-AI workshop whitepaper,practitioner-opinion, no measurement) sides against Consideration 05 as written and supplies the lever the disagreement lacked: reviewer competence "does not persist automatically. It requires organisations to deliberately assign workers enough substantive review work and enough exposure to AI failure to keep the verification skill current." On that account an oversight workforce is staffable conditionally, which narrows and cheapens the falsifiable core — not "was an oversight workforce ever staffed without an execution cohort beneath it," but "was its substantive-review share deliberately maintained, and did its error-catching hold." Its own worked example applies the rule (juniors must perform the initial protocol design and coding themselves, justified by anchoring bias and the risk of "vacuous verification" of an AI's generation). No deployment, no year, no error-catching measurement: unmoved on evidence, better specified on design. -
NANTE's practical claim is that stall types are different kinds of problem: a Navigate stall with a shallow-plateau flag needs workflow redesign, a Navigate stall with a low-task-success flag needs enablement, and applying either to the other is waste. Nothing tests this. Does a cohort diagnosed with one stall type respond differentially to its mapped intervention class versus the alternatives? Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals names this as the strongest of its own proposed validations (§8.1) and runs none of it — its evaluation is synthetic populations whose pathologies and thresholds were co-developed, which establishes computability and not correctness. Falsifiable with a real deployment: classify cohorts prospectively, assign interventions across rather than only along the mapping, and compare depth conversion afterwards.
Sources#
-
The state of AI in 2026: On the road to ROI — Dan Tinkoff, Lieven Van der Veken & Michael Chui with Tara Balakrishnan, The state of AI in 2026: On the road to ROI (McKinsey / QuantumBlack, 2026-08-25,
empirical, self-reported; online survey, 1,719 participants in 97 nations, fielded May 4 - June 8 2026, GDP-weighted). Cited here for Exhibit 1 (phase distribution 2025-26), the 37%-EBIT / 80%-individual-productivity pair, Exhibit 10 (objectives by cohort), Exhibit 11 (the high-performer practice gaps) and Exhibit 5 (perceived impacts by organizational level). Exhibit 11 carries no printed numeric labels — every value quoted from it here was read off the chart's gridlines at ingest and is approximate, directional only. COI: McKinsey sells the AI transformation consulting this finding implies demand for; the high-performer construct is defined on the outcome it is used to explain. Full evidence handling at Source Notes -
When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering — Stolze & Strässle (OST Eastern Switzerland UAS / smartive AG, arXiv 2608.26316, 2026-08-26, ESEM 2026 SEIP),
case-study: §4.4 (the upward re-scoping of oversight and operational explainability) and §4.5 (the supervisory-capacity lower bound). Cited here only against consideration 05's operators-to-overseers conversion; five interviews with no measured outcome — evidence notes at Layered Supervision -
Deploying AI from pilot to production: A practical blueprint for CIOs and technical leaders — Deploying AI from pilot to production: A practical blueprint for CIOs and technical leaders, Anthropic × Accenture, 2026-09-11, 38pp,
vendor-claim. Co-branded enterprise collateral: the authors sell the product the guide recommends, the "Getting started" section is a list of Claude product links, and every deployment figure is either an unattributed anecdote or a vendor case-study page. The prescriptive content — the insulation catalog, the Work-out/Assign split, the three components of ownership, the four-tier oversight ladder, the error-mode-distribution test — is practitioner judgment and is where the document's value sits; the numbers are Accenture's own surveys (Pulse of Change July 2026, AI-Ready Data for Advanced AI May 2026, Tokenomics September 2026), self-published without methodology. Parse warning: p.35's deployment blueprint is drawn as page graphics — docling recovered only its decorative timeline dots andpdftotextrecovers only the running head, so the page reads as present while carrying none of its content. The table on this page is a manual transcription made at ingest from the rendered page and is the only copy of that content in the vault; the source's own typo ("No new owners or decsions") is preserved in the raw with[sic]. The nine extracted tables parsed clean -
Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals — Damon A. Young (PolyWise Partners, sole author), Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals, arXiv 2608.23617 (submitted 2026-08-22; the paper's own dateline reads August 26, 2026), 23pp,
practitioner-opinion— a method proposal with no real-organization data of any kind. §1 (the secondhand failure statistics and their own caveats), §4.1–4.6 (the stage model, telemetry signatures, stall detection, failure signatures, intervention mapping, composite score), §5.2–5.4 (the six synthetic profiles, results, and the circularity admission), §6.1–6.5 (limitations), §8.4–8.5 (export-format degradation and the disposition dimension). COI: a consultancy proposing an instrument it would deploy, closing with a design-partner invitation; treat the framework as a marketable artifact, not a neutral method. Table 1 was reconciled cell-for-cell againstpdftotext -layoutand is exact — no collapse, no shift; the only wrapping artifact is the "Champion depen- dency" flag cell. Figure 1 (the stage diagram) was read as an image and agrees with the prose; Figure 2 is the toolkit's generated one-pager for the healthy/shallow pair, labelled in the source itself "Illustrative synthetic data · thresholds proposed, not validated." Instrument argument and the four-tradition gap at Telemetry vs. Survey Measurement.
Cited by 19
- Accenture×4
claude accenture pilot to production — Deploying AI from pilot to production: A practical blueprint…
- Standardize the Infrastructure, Not the Tools×4
Independent arrival at the same mechanism. Anthropic × Accenture prescribe the identical move for…
- Open Questions Backlog×3
Pilot To Production Gap ×2 (oldest 13d) — The blueprint's load-bearing claim is an ordering claim:…
- Telemetry vs. Survey Measurement×3
Where it lands on the survey arm. As a complement, on the AEI footing this page already describes:…
- AI-Assisted Error Analysis×2
Pilot To Production Gap — the same consequence-first reasoning applied to the automation boundary…
- AI Employee Framing×2
Pilot To Production Gap — the structural route to the same unowned outcome. Kropp et al. locate the…
- Build Instead of Buy Under Agentic Coding×2
No completion. The survey records a decision at decision time. Whether the in-house build shipped,…
- Conversation-to-Delegation Shift×2
Pilot To Production Gap — the same shift argued prospectively for enterprise workflows rather than…
- Document Parsing as the Retrieval Bottleneck×2
claude accenture pilot to production — Deploying AI from pilot to production, Anthropic ×…
- Failures That Look Like Success×2
Pilot To Production Gap — the same failure class at the scale of a deployment programme. Anthropic…
- Forward-Deployed Engineering as a Delivery Layer×2
Pilot To Production Gap — the condition the layer exists to bill against: an enterprise whose pilot…
- Organizational Complements to AI×2
Pilot To Production Gap — the practitioner-side account of why the complement gap stays invisible…
- Risk-Tiered Auto-Approval×2
Pilot To Production Gap — where this page's tiering sits in an enterprise deployment lifecycle.…
- The Tragedy of the Cognitive Commons×2
Pilot To Production Gap — the prescription this page's argument contradicts, written by parties who…
- The Enterprise AI Adoption Gradient
Pilot To Production Gap — "adoption is only the beginning of deployment" is that page's thesis with…
- Layered Supervision
Pilot To Production Gap — the operators-to-overseers conversion, at a different altitude. That…
- Product & Organization
Pilot To Production Gap — Anthropic × Accenture's account of why enterprise AI pilots don't predict…
- Psychological Costs of AI Adoption
Pilot To Production Gap — its Consideration 05 (converting operators into overseers) has an…
- Usage-Telemetry Classifier Validation
adoption telemetry enterprise ai production signals — Damon A. Young (PolyWise Partners, sole…
Related articles
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Telemetry vs. Survey Measurement
Perception lags reality: survey-based research (DORA) misses damage system telemetry catches — plus the family effect (…
- AI Adoption in Scientific Work
Three instrument families on one profession. Google ATLAS telemetry + survey: scientists are the economy's heaviest AI…
- Firm AI-Spend Intensity and Headcount Growth
Ramp × Revelio panel of 21,559 US firms: high-intensity AI-vendor spenders grow headcount ~10% (entry-level ~12%) over…
