H
Howardism
Plate IIProduct & Org中文HOWARDISM

Standardize the Infrastructure, Not the Tools

Shopify's inversion of the one-tool-per-job norm for AI: route every coding agent through a central LLM proxy so leadership gets cost control, per-team usage analytics, and model portability, while engineers keep free tool choice — buying optionality under uncertainty about which model or workflow wins, with MCP servers extending the same governs-access-not-engineers principle to internal systems; Accenture's Tokenomics figures price what the meter is for (42% of orgs have no single AI-cost owner; formal chargeback ties 32¢ of every token dollar to an outcome, 6× no allocation); and the first population-scale record of the model-mix decision actually being exercised — frontier-model token share 53%→45% in five weeks as firms impose company-wide defaults on cost-effectiveness grounds (Ramp, September 2026), outcome observed, mechanism still not

Article metadata
Publication details
Published:August 11, 2026
Filed:Concept
Domain:Product & Org
Reading:25 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Standardize the Infrastructure, Not the Tools

Sources#

Summary#

Farhan Thawar (VP & Head of Engineering, Shopify) states the rule as a deliberate exception to his own org's convention:

"At Shopify, we always have one tool for one job, except for with AI. Since we don't know yet which company, workflow, or model is going to win."

The mechanism is an internal LLM proxy — a single gateway every AI request passes through before reaching a model, whether it originates from Claude Code, Copilot, Cursor, Codex, or anything else. What the org standardizes is the layer underneath the tools; the tools themselves stay uncontrolled.

case-study tier, first-hand from a named leader at a named company, and unmeasured: no cost figure, no analytics finding, and no counterfactual are given.

What the layer buys#

Three things, per the source, and they are worth separating because they have different half-lives:

  1. Centralized cost control and usage analytics by team and project. This is the durable one, and it is the org-level instrument the vault otherwise lacks. Cost-per-Task Over Cost-per-Token establishes that the right unit of AI cost is the completed task and that harness choice moves it more than model choice — but a per-task accounting requires seeing every request, which is exactly what a gateway provides and a fleet of independently-billed tools does not.
  2. Model portability — "the ability to switch models as capabilities evolve without forcing engineers into a single workflow." The proxy makes the model a swappable dependency rather than a property of each engineer's toolchain.
  3. Preserved tool experimentation. Engineers are not funneled into one workflow, so the org keeps sampling the space while it is still unsettled.

The underlying claim is about optionality under uncertainty, not about efficiency. The standard argument for one-tool-per-job is that fragmentation costs more than it's worth; Thawar's counter is that the cost of fragmentation is temporary and bounded, while the cost of standardizing on the losing tool is neither. That reasoning holds only while the winner is genuinely unknown — it is an explicitly dated position, and the rule inverts back the moment the market settles.

The same principle applied to internal systems#

Shopify extended the pattern past model access: through MCP servers, engineers query Salesforce, Slack, GitHub, and internal wikis from AI assistants "with the same access controls as their normal auth flow."

The load-bearing phrase is the source's own summary of why this scales: "The infrastructure governs access, not individual engineers." That is the same architectural move as the proxy — put the control at a chokepoint the org owns, so that the thing being governed (spend, or reach into internal systems) is governed structurally rather than by policy each engineer has to follow. It is also the organizational counterpart to what Zero Trust for AI Agents argues at the protocol level, arrived at from a cost-and-analytics motivation rather than a security one, and reaching the same shape.

Worth noting what "the same access controls as their normal auth flow" does and does not settle: it makes the agent's reach equal to the engineer's, which is the right default and is precisely the property Agent Identity and Authentication treats as the hard problem. The source asserts it as a solved implementation detail and gives no mechanism.

And the population-scale prior runs the other way (2026-05). Remote MCP Authentication in the Wild censused 7,973 live remote MCP servers and found 40.55% exposing tools with no authentication mechanism at all, 29.00% on static tokens or API keys, and every one of 119 end-to-end-testable OAuth deployments carrying at least one confirmed authentication flaw. That does not contradict Shopify — an internally-run server behind a corporate identity provider is exactly the configuration a public census cannot see, and is plausibly the good case. It does mean the claim is the unusual outcome rather than the default one, and that "the infrastructure governs access" is load-bearing precisely because MCP by itself does not: on the measured population, wiring an assistant into Salesforce or Slack through a third-party MCP server inherits no access control at all. practitioner-opinion asserting a property against empirical measurement of the surrounding ecosystem — the mechanism the source declines to give is the whole question.

What the meter is worth, priced (Accenture Tokenomics, September 2026)#

Benefit 1 above is asserted by Shopify and unmeasured. Deploying AI from pilot to production (Anthropic × Accenture, vendor-claim) supplies the first numbers in the corpus attached to it, from Accenture's September 2026 Tokenomics research:

  • 42% of organizations rely on shared IT-and-finance accountability with no single owner for AI costs and outcomes.
  • Organizations with formal chargeback accountability for AI spending link 32 cents of every dollar of token spend to a quantified business outcome — six times the rate of those with no allocation.

Read carefully, this is not a claim about the gateway. A proxy produces per-team usage analytics; the 6× belongs to chargeback, which is the accounting policy that makes someone's budget absorb the bill. Shopify's layer is the precondition (you cannot charge back what you cannot attribute) and the document's own prescription — a named owner with decision rights, escalation authority and executive backing — is the other half. So the corpus now has a plausible mechanism on one side, a claimed outcome on the other, and nothing joining them: no source shows an org that built the meter and then did or did not turn it into chargeback.

The 42% is also the sharper of the two figures and cuts against the usual reading of this page. The default failure is not standardizing on the wrong tool; it is that nobody owns the bill at all, which no amount of substrate discipline fixes. Both figures are Accenture's own survey, self-published without methodology — treat them as the shape of the argument, not its size. See Pilot-to-Production Gap for the surrounding blueprint and its evidence limits.

Adoption by demonstration, not mandate#

The organizational half of the same account, and the part that generalizes past engineering: Thawar "didn't mandate AI adoption. He modeled it," sharing his own AI-assisted work framed as leverage rather than as capability —

"I didn't say look at how much work I did and how smart I am. I said, 'Look how lazy I am.'"

The claimed effect is spread outside engineering: sales reps building dashboards, finance building workflow tools without an engineering ticket, HR generating "n-of-1" software. That last phrase is the same phenomenon Implementation Abundance Inverts Product Work and Printing Press Software Democratization describe from the supply side — software worth writing for one user — observed here as a downstream effect of a tooling decision. No measurement accompanies it; "a boost in productivity across the engineering organization" is the strongest form the claim takes.

Independent arrival at the same mechanism. Anthropic × Accenture prescribe the identical move for enterprise rollouts and state the reason more precisely: identify champions and engineer the demonstration moments deliberately rather than wait for them, because "a portfolio manager showing a compliance specialist how she summarized a 200-page filing in three minutes converts more skeptics than any structured rollout." Two sources, different domains (a single engineering org vs. cross-industry enterprise deployments), same conclusion that demonstration beats mandate — and the same absence of measurement on both sides. The addition worth keeping is the word deliberately: Thawar's account reads as a leader modelling behavior, where the blueprint treats the demonstration as a staffed, funded rollout step with a named owner.

Where the standardization actually landed, in teams that never built a gateway (September 2026)#

Shopify's version of this principle is a vendor-adjacent account of a large platform team with the budget to build a proxy. Stolze & Strässle (ESEM 2026 SEIP, case-study, five interviews across a CTO's enterprise team, an energy-utility frontend group, a construction-software team, a digital agency and an industrial-technology team) is a non-vendor look at the same choice made by organizations with no platform layer at all — and the shape holds, one layer down.

In none of the five did tool choice get standardized. AI-tool governance stayed informal or absent in three of the five, and the accompanying survey splits nearly evenly across the whole population: 24 of 50 respondents reported clear organizational guidelines for AI-tool use and 23 reported informal arrangements or no monitoring at all. What did get standardized, deliberately and by every participant who described a practice, is the constraint layer: lint rules, CI checks, architectural-conformance checks, build-breaking conventions and shared steering artifacts. One participant's formulation of the rule is the one this page would recognize — "If a rule is relevant, it must be enforced through linting" [P4] — and two describe the build system becoming the arbiter, so violations break the build rather than being argued with a reviewer.

That is the same governs-the-substrate-not-the-engineer principle operating on constraints instead of on model access, from a source with no product to sell. It is corroboration of the principle's generality, not of any figure here: five interviews, a convenience-sampled and unpiloted survey, no measurement. Note also what it does not touch — neither open question below. The paper contains no gateway, no model-mix decision, no proxy, no usage analytics and no cost data of any kind, so the portability question and the what-is-the-telemetry-used-for question are exactly where they were. Full treatment at Layered Supervision.

The model-mix decision, exercised at market scale (Ramp, September 2026)#

Both open questions below exist because the corpus recorded the claimed benefits of a gateway and never an org acting on one. Ramp's September 2026 AI Index (Ara Kharazian, empirical, corporate-card and bill-pay records for ~70,000 US businesses; instrument and COI at Ramp) is the first source that records the outcome of the portability decision across a population — and, tellingly, still not the mechanism.

What it observes. Token share by model tier, weekly: frontier models (Opus, Fable, Sol) held 52.5% in the week of August 2 and 44.7% by the week of August 30 — the letter rounds this to "45% of token share… down from a 53% peak in August" — while standard models (GPT-5.6 Terra, Claude's Sonnet series) went 25.7% → 35.8% across the same five weeks. The lite and other tiers were roughly flat (16.3% → 14.2% and 5.4% → 5.3%), so the movement is a direct frontier→standard substitution rather than a general flight to the cheapest thing available. Kharazian states the mechanism as his buyers report it to him:

"We've heard from businesses who are imposing company-wide defaults that reduce usage of frontier models, saying standard models are still highly performant and also more cost effective."

"Company-wide defaults" is this page's principle, stated by someone with no gateway to sell. A default that applies company-wide governs the substrate: it binds what engineers get unless they act, which is the same move as the proxy above and as the lint rules the five teams in Layered Supervision standardized instead of standardizing tools. And it is a model-mix decision actually taken — at enough firms to move a market-wide token-share series eight points in five weeks.

What it does not observe, and why Q1 below stays open. Ramp sees dollars crossing a payment rail. It cannot see how a default is imposed: a gateway policy, a config default in a shared harness, a procurement rule, a budget that makes the choice for people, or simply picking the cheaper model in a vendor's own console. Nothing in the letter mentions a proxy, a router or a gateway, and "we've heard from businesses" is anecdote attached to an aggregate. The corpus therefore now has the outcome (orgs do switch model mix, deliberately and en bloc) without the mechanism (whether a portability layer is what makes switching cheap). Firms reaching the same place through vendor-side defaults and no portability layer at all is the cheaper hypothesis, and this instrument cannot exclude it.

The confound the same edition supplies. Effective blended cost fell 41% to $0.68 per million tokens from a March 2026 peak of $1.15, with both OpenAI and Anthropic announcing price cuts in the month before publication. The price lever a gateway is meant to pull was moving hard on its own, which makes attributing the mix shift to any org-side mechanism harder rather than easier — see Cost-per-Task Over Cost-per-Token, where the same series is read as a buyer-side verdict on the start-with-the-strongest-model default.

Cost as a design constraint, at population scale (McKinsey, August 2026)#

The meter this page argues for exists so that someone can act on what it reads. The state of AI in 2026 (McKinsey / QuantumBlack, 2026-08-25, empirical but self-reported, n=1,719) says roughly one organization in five has now reached the point where the reading changes behaviour: about 20% of respondents report that AI-related operating costs, including token costs, have constrained their organization's AI use — "broadly consistent across organizations of different sizes," 12% (consumer goods and retail, public sector) to 25% (technology) by industry (Exhibit 8), and about one in ten on each of chatbots, agents and coding agents taken separately.

Two cuts sharpen it. First, the constraint is not concentrated in the least sophisticated organizations: AI high performers report cost-constrained use of software coding agents at 18% against 6% of all others — the only tool class in Exhibit 14 where the leading cohort reports more constraint than everyone else. Second, managing that cost is itself one of the practices the survey scores: "actively manage AI solution costs (tokens, compute, storage)" separates high performers from the rest by only ~34% to ~25% (Exhibit 11, dumbbell chart with no printed labels — gridline-read, approximate), one of the narrowest gaps in the chart. So cost management is not what distinguishes the successful cohort; it is a hygiene practice a quarter of everyone reports and which only just correlates with outcomes.

That is the honest frame for this page's argument: a central gateway is how an organization notices the bill, and noticing is spreading, but noticing has not so far been the variable that separates the organizations getting value.

Connections#

  • Pilot-to-Production Gap — the enterprise-deployment frame this decision sits inside, and it resolves the commitment question oppositely. That blueprint names an uncommitted build-versus-buy as the primary source of engineering debt ("every decision deferred past the pilot creates engineering debt: parallel systems, integration patterns that never standardize") where Shopify deliberately declines to commit and buys optionality instead. The two are reconcilable — Shopify did commit, to the substrate, and left only the replaceable layer open — and the distinction that makes them compatible is which layer the commitment binds. It also supplies the chargeback figures above and the same adoption-by-demonstration mechanism
  • Layered Supervision — the same principle without the platform budget. Five teams with no gateway standardized the constraint layer (lint, CI, build-breaking conventions, shared steering artifacts) while leaving AI-tool choice and governance informal in three of five — non-vendor corroboration that what organizations actually standardize is the substrate that binds output, not the tool that produces it. Contains no gateway, no cost data and no bearing on either open question below
  • Firm AI-Spend Intensity and Headcount Growth — the home of the payment-rail instrument behind the market-scale section above, and the reason its numbers are read as directions rather than levels: Ramp measures its own VC-forward-skewed card base, and its dollar series revise upward for months after first print
  • Cost-per-Task Over Cost-per-Token — the accounting this gateway makes possible: per-task cost is only measurable if every request passes one meter
  • Orchestration Sets Token Economics — the reason model portability matters more than model choice: the harness, not the model, sets the bill
  • AI-Native Organization — the org-design frame this is an infrastructure decision inside; here the encoded layer is the gateway rather than the skill library
  • Agentic Work Systematization — the same standardize-the-substrate instinct one level down, at the level of reusable skills rather than model access
  • Zero Trust for AI Agents — the security-motivated version of "the infrastructure governs access, not individual engineers"
  • Remote MCP Authentication in the Wild — the empirical counterweight to the MCP access-control claim above: across 7,973 live remote MCP servers, 40.55% authenticate nothing and 119 of 119 testable OAuth deployments carry an authentication flaw, so "the same access controls as their normal auth flow" describes an org that built the chokepoint, not a property MCP supplies
  • Agent Identity and Authentication — the hard problem the MCP claim asserts away: agent access equal to the operator's auth flow, stated as a property rather than a mechanism
  • Implementation Abundance Inverts Product Work — the claimed downstream effect: non-engineers building "n-of-1" software once the substrate is available
  • Build Instead of Buy Under Agentic Coding — what the metered substrate gets spent on once it is cheap enough: 32% of a 1,719-respondent panel report declining a software purchase because agentic coding could build the feature instead. The build decision and the token bill land on the same budget line, and this page's gateway is where an organization sees both
  • Outsource Your Thinking, Not Your Understanding — the same source's warning about what this speed costs: the gateway meters tokens and reversion rates, neither of which surfaces comprehension debt

Open Questions#

  • Does a central LLM gateway actually change model-mix decisions, or only report on them? The claimed benefit is portability; no source in the corpus records an org exercising it. Partially answered (2026-09-22): September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 records the outcome across ~70,000 US businesses — frontier-model token share falling 52.5% to 44.7% over five weeks while standard-tier share rose 25.7% to 35.8%, with buyers telling Ramp they are "imposing company-wide defaults that reduce usage of frontier models" because standard models are "still highly performant and also more cost effective." Organizations demonstrably do exercise a model-mix switch, deliberately and company-wide, and the switch is large enough to move a market aggregate. The gateway half is untouched: a payment-rail instrument cannot see whether a proxy, a shared-harness config, a procurement rule or a vendor-side default enforces the decision, and the letter names no gateway anywhere. What is left of the question is precisely the mechanism — is a portability layer what makes the switch cheap, or do firms reach the same place without one? A confound also arrived with the answer: effective token prices fell 41% over the same window, so the mix could be moving on price alone. A third mechanism, added 2026-09-22 — and it is a genuinely new axis, because it is neither a gateway nor a price. Startup ARR is less secure than ever, new research shows (practitioner-opinion, press reporting of an unread VC survey) reports Madrona finding 77% of 150 enterprise IT professionals re-evaluate their AI vendors every six months or on a rolling basis, with Madrona's own gloss that "switching costs are lower and the re-evaluation cadence is relentless." That is a procurement calendar, not a portability layer: it makes switching cheap by making the decision routine and scheduled rather than by making the migration technically easy, and it reaches a vendor-level switch with no infrastructure at all. So the bullet's remaining question — is a portability layer what makes the switch cheap, or do firms reach the same place without one? — now has a live candidate for the without-one branch, and this source cannot adjudicate it either: it names no infrastructure, never asks what a re-evaluation costs to act on, and reports a stated cadence rather than an observed switch. Note also the unit mismatch with the Ramp reading — Madrona counts AI vendors (applications), Ramp counts model tiers, and nothing establishes that the same cadence governs both.
  • What does per-team AI usage analytics get used for once it exists — cost containment, capacity planning, or performance evaluation of engineers? The third would collide with everything Telemetry vs. Survey Measurement establishes about what instrumented output data can and cannot support. Partially answered (2026-09-22): the first of the three is now observed, at firm rather than team granularity. September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 reports businesses acting on AI-spend visibility by imposing company-wide model defaults on explicitly cost-effectiveness grounds — cost containment, and the only one of the three uses any source in the corpus has recorded. Two limits keep it open: the granularity is wrong (vendor-level firm spend on a card rail, not per-team usage), and the publisher sells the spend-visibility product, so the framing is marketing surface as well as measurement. Capacity planning and performance evaluation remain entirely unobserved. Extended 2026-09-22 by The state of AI in 2026: On the road to ROI (empirical, self-reported, n=1,719), which adds the population denominator the Ramp letter cannot: ~20% of respondents report AI operating costs, including tokens, actually constraining their organization's AI use (Exhibit 8), spread 12% to 25% by industry and "broadly consistent" across company sizes. That is the first estimate of how many organizations have reached the point where a cost reading binds rather than merely informs, and it is the cost-containment use observed at a third instrument. Two qualifications keep the bullet open in the same place. The constraint is not a beginners' problem — AI high performers report cost-constrained coding-agent use at 18% against 6% of others (Exhibit 14) — and, more to the point of this bullet, the survey asks about outcomes of cost pressure and never about the instrument: no question mentions a gateway, a proxy, chargeback or per-team analytics, so what an organization looked at before constraining itself is still unobserved. Capacity planning and performance evaluation of engineers remain unobserved at every instrument.

Sources#

  • When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering — Stolze & Strässle (OST Eastern Switzerland UAS / smartive AG, arXiv 2608.26316, 2026-08-26, ESEM 2026 SEIP), case-study: §4.2–4.3 (the constraint layer as the standardized object, the lint-promotion rule, the build as arbiter) and §4.5 plus §3.3 (the governance-posture split, 24/50 vs 23/50). Non-vendor but tiny — five interviews and a convenience-sampled, unpiloted survey; evidence notes at Layered Supervision
  • Inside AI-pilled engineering teams: Five lessons for scaling without losing the plot — Bessemer Atlas, 2026-06-10 (case-study for the practitioner material): §2 "How Shopify enabled AI tool experimentation without chaos". Thawar quotes are first-hand; the productivity and cross-functional-adoption effects are his characterization, unmeasured. The post's opening adoption percentages are a separate, lower tier — see the Source Notes entry
  • A First Measurement Study on Authentication Security in Real-World Remote MCP Servers — Zhou et al. (Fudan University; one author at Central South University), A First Measurement Study on Authentication Security in Real-World Remote MCP Servers, arXiv 2605.22333, 2026-05-21, empirical, no COI. Cited here only as the ecosystem prior against which this source's unmechanized MCP access-control claim should be read — §3.2's Table 2 split over 7,973 validated servers and Finding 3.1 (119/119 flawed). Parse warning and full treatment on Remote MCP Authentication in the Wild
  • Deploying AI from pilot to production: A practical blueprint for CIOs and technical leaders — Deploying AI from pilot to production, Anthropic × Accenture, 2026-09-11, 38pp, vendor-claim. Cited here for the Accenture Tokenomics (September 2026) chargeback and ownership figures, and for the adoption-by-demonstration corroboration. Both figures are the publisher's own survey, self-published without methodology. Full treatment and evidence limits on Pilot-to-Production Gap
  • The state of AI in 2026: On the road to ROI — Dan Tinkoff, Lieven Van der Veken & Michael Chui with Tara Balakrishnan, The state of AI in 2026: On the road to ROI (McKinsey / QuantumBlack, 2026-08-25, empirical, self-reported; online survey, 1,719 participants in 97 nations, fielded May 4 - June 8 2026, GDP-weighted). Cited here for Exhibit 8 (20% cost-constrained by industry), Exhibit 14 (high-performer coding-agent constraint) and Exhibit 11's cost-management practice row — the last of which is gridline-read from an unlabelled dumbbell chart and approximate. COI: McKinsey sells AI transformation consulting; the survey is self-report from its own panel. Full treatment at Build Instead of Buy Under Agentic Coding
  • September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 — Ara Kharazian, September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 (Ramp, 2026-09-09, empirical): the weekly token-share-by-tier series (recovered from the article's Datawrapper dataset endpoint, since the page carries no static chart images) and the prose paragraph reporting company-wide defaults away from frontier models. Cited here for organizational behavior, not for the model market. COI: Ramp measures its own corporate-card customer base and sells the spend-visibility product whose use this page's second open question is about; the instrument's aperture and revision behavior are documented at Firm AI-Spend Intensity and Headcount Growth and Ramp
  • Startup ARR is less secure than ever, new research shows — Julie Bort, TechCrunch, 2026-09-03 (practitioner-opinion, press reporting of an unread Madrona survey, n=150 enterprise IT professionals): cited here only for the 77% six-month/rolling vendor re-evaluation cadence, as the procurement-calendar mechanism for cheap switching that involves no portability layer. Stated process, not observed switches; unit is AI vendors, not model tiers
§ end
Cited by 19
Related articles
  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Organizational Complements to AI

    The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…

  • Returns to Expertise in Agentic Coding

    Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Engineer PM Convergence

    Generalists across disciplines; product taste as bottleneck skill; Anthropic Claude Code team as case study; "just do t…