Howardism · Vol. 03Plate II · No. 02
AI Economics & Labor, in order.
Notes37DomainAI Economics & LaborOpen Qs100Newest1 Oct 2026Oldest8 May 2026
Work, wages, org design, and the economics of AI-driven labor.
Map of Content for the ai-economics-and-labor domain — 32 concepts. AI's measured economic footprint: usage telemetry, labor-market effects, returns to expertise, organizational complements, and framing effects on accountability. Curated entry point; see Home for all domains.
- AI Adoption in Scientific Work — Three instrument families on one profession. Google ATLAS telemetry + survey: scientists are the economy's heaviest AI users (SOC 19 over-indexes 2.7x; 46.6% use AI daily), LLMs and 2,690 specialized models are complements (elasticity 0.6->0.2), and a self-reported 6.9 hours/week saved does not become discovery because the bottleneck moves downstream (43.5% say so; 45.7% spend over a quarter of it verifying; 48.8% tilt to safer questions). Against it, a latent class analysis of 3,785 PhD students finds attitudes arranged by task: 51.5% comfortable with AI summarising literature against ~30% for writing, analysis and experiment design, in a 44% "division of labour" profile. Beside both, publication traces (~22% of computer-science output carrying LLM-modified text by September 2024) - blind to analysis, but on writing they find behaviour where the survey finds refusal
- AI and Market Power — OECD AI Papers No. 62 on French and Portuguese firm microdata plus global patent and start-up databases: non-GenAI adopters hold 7.5×/3.2× the market share of non-users, but the premium is selection (dies once broadband, digitalisation and lagged productivity enter) and adopters gain no market-share rank or markup growth over five years; firm-level GenAI exposure is inverted-U in size and market share while monotone in productivity and tertiary education; global AI-patent concentration fell 32–60% over 2001–21 yet correlates positively with sales concentration within markets; AI patents raise markups only in ICT (+7.95% interaction); and GenAI start-ups take ~130% more VC and are ~21% more likely to be acquired by incumbents
- AI Brain Fry — Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognitive cost surface for both tool and employee framings
- AI Employee Framing — Kropp et al. (HBR May 2026, n=1,261): framing AI agents as "employees" vs "tools" cuts personal accountability −9pp, increases escalation +44%, reduces error catching −18%, no adoption gain
- AI-Exposed College Majors at Labor-Market Entry — What happens to new bachelor's graduates whose field of study was AI-exposed before ChatGPT. The Census Bureau's PSEO×LEHD records (Orr, Tucker & Warren, Sept 2026; ~29% of US bachelor's degrees 2016–2024, mostly public institutions) show the most-exposed decile of majors, dominated by computer science, losing 5pp of initial employment and 13% of initial full-quarter earnings, about half of it from sorting into retail and food service. Deciles 7–9 recover within 1–2 years; the top decile is still −2.0pp and −3.3 log points at two years. Exposure is fixed by major before the outcome, so this design sees the adjustment that occupation-conditioned studies cannot. A CPS study (Fairlie & Wu, NBER w35796) finds no summer-2026 break in unemployment for all 22–25 bachelor's holders. That is consistent with these results: the two use different outcomes and populations.
- The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay — The wage margin of AI exposure, read on two instruments that turn out to agree once the margin is matched: Indeed Hiring Lab's US advertised salaries (Sept 2026) show pay in the most GSTI-exposed occupations up ~46% since 2021 vs ~25% in the least, a +5.7% post-ChatGPT diff-in-diff premium that shrinks to +2.4% (n.s.) once seniority mix is held — because the entry-level share of exposed-occupation postings fell 29%→10% — while ADP payroll (Canaries Fact 6) finds little pay divergence at all; the premium is a new-hire, senior-offer, mostly-composition effect, not a raise for incumbents
- AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins — Deng, Liu, Toubia & Jain (arXiv 2609.29143): a preregistered market-research study (GPT-5.1 voice moderator N=139 vs same-guide static audio N=154, randomized; expert human moderators N=24, separate later sample) — adaptive AI probing beats a standardized static guide on breadth, depth and customer needs (7.48 vs 5.01 per interview) and matches humans on content at ~20% of the cost, but participants sound flatter than with a human; digital twins built from the richer AI transcripts predict held-out ad reactions no better than twins built from static answers, and twins reason more System-2 than the people they copy
- AI Usage Cadences — AEI Cadences report: continuous hourly telemetry reveals AI usage carries the rhythms of daily life — personal use spikes 35%→~50% on weekends, recipes 2.3× at 6pm, sleep advice pre-dawn, tax queries 8× around the Apr-15 deadline; off-hours work skews toward higher-wage occupations
- The Automation–Optimism Link — AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-quality dimensions (pay, security, job-finding, meaning, autonomy, human interaction), report their skills growing more valuable, and show no learning deficit — inverting the common delegation→deskilling-anxiety narrative
- Codified vs Tacit Knowledge Exposure — The mechanism axis under the age gradient in AI's labor-market effects: generative AI substitutes best for codified knowledge (formal, documented, taught) and complements tacit knowledge (acquired through practice and mentorship), so it lands first on whoever's comparative advantage is freshly-acquired schooling. First measured in the August 2026 Canaries revision via a GPT-5.4-mini index over 7,514 ADP job descriptions — young workers in high-codified occupations decline while experienced workers in high-tacit occupations grow faster — with the honest asymmetry that the codified half dies under an education control and the tacit half survives it
- Context Advantage, Not Taste — Andrew Ng's reframing of the residual human contribution: not 'taste' but an information asymmetry — 'so long as the human knows something the AI does not, human-in-the-loop is needed.' Recasts the wiki's central open question (is taste a ceiling or the next jagged valley?) as a category error, and makes the human role a closable engineering gap rather than a moat — a reading Ng himself softened in August 2026, calling the asymmetry "a long-term advantage" no one will close "in a few years"
- Controlled Variance: AI's Edge as Reduced Dispersion — Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants between AI voice interviewers and human recruiters — offer rate 8.70%→9.73% (+12%), job starts +18%, one-month retention +18%, no productivity decline, with humans making every hiring decision in both arms. The mechanism the authors name is controlled variance: the AI follows the firm's interview protocol more consistently (topic order τ 0.53 vs 0.33, question similarity 0.59 vs 0.43, significantly lower cross-interview variance) while still adapting per applicant and using richer vocabulary — AI wins by being less dispersed, not more capable. The wiki's only randomized causal estimate of AI substituting for a human in an expert conversational task
- Conversation Artifacts — AEI Cadences report: the 'artifact' (the primary output a user takes away) as a new unit of economic analysis — 93% of conversations produce one, artifact type predicts work/personal/coursework use, compute (tokens) scales with the artifact's economic value, and Claude's output sits ~1 education-year above the prompt
- Conversation-to-Delegation Shift — OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production'), measured by Codex's share of output tokens across three populations — 99.8% OpenAI / 63.3% organizational / 16.5% individual — with adoption spreading beyond developers; standard usage metrics (active users, chats) become less informative as the unit shifts from a conversation to a delegated workflow
- The Enablement–Regulation Axis — Chueri & Törnberg (VU Amsterdam / UvA, arXiv 2609.02296): 1,514,950 parliamentary speeches from 33 national parliaments, Jan 2023–Apr 2026, LLM-coded down to 5,317 AI-work speeches. Political economy expects technological disruption to produce demands for compensation; compensation is 2.3% of response-frame mentions (125 in total), against enablement and investment 55.2%, regulation and restriction 21.8%, training 20.6%. The conflict is over whether public authority should accelerate adoption or govern it, and the seven party families fall on that single axis — radical left ≈89% threat / ≈67% regulation, conservatives and the radical right ≈86%/79% opportunity and ≈76% enablement, with the radical right declining to politicize AI as a labor threat. Inside the regulation frame copyright leads at 36.1% while algorithmic management and human oversight is last at 13.0%
- The Enterprise AI Adoption Gradient — OpenAI's ChatGPT Enterprise study (Chatterji, Holtz, Rakholia, Tambe & Weeratunga, Aug 2026): admin records for 1,764 organizations / 17.4M messages linked to Compustat — output tokens 7x in nine months (4x inside the pre-June-2025 cohort); among US public firms adoption rises with scale (+6.9pp top revenue quartile, +9.8pp top 5%) and with FY2021 intangible stocks (SG&A 0.020***, capitalized software 0.008**, R&D 0.004***); but conditional on adopting, per-employee use falls with size while messages per active user is flat — larger firms buy breadth they do not use. Within firms, use spans every function with a negative seniority gradient, and 56.3% of active users touch documentation/technical writing. Descriptive only, buyers-only sample, all five authors OpenAI-affiliated or OpenAI-paid.
- Experimental Learning Impact of Generative AI — Randomized experiments on how AI use changes learning. Contractor & Reyes (n=211): AI access raises test scores +0.27 SD, ~76% persisting unaided a week later, but durable gains go to 'augmentation' users (AI as tutor) while 'automation' users' gains vanish. Bocconi x OpenAI 2x2 (n=1,053): a taught causal-reasoning skill survives the tool while the rubric rewards the tool. LMU preregistered trial (n=704): metacognitive feedback on one's own AI use cuts answer offloading (OR 0.47) and lifts unaided scores (OR 1.51); a points penalty does neither
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoretical exposure (tasks an LLM could do), reported exposure (what workers say AI can do today), and anticipated exposure (what they expect in 12 months) — plus their orderings (theoretical > reported > observed), the GDP/experience/automation gradients the AEI survey reveals, and Steele & Cruz's seven-instrument head-to-head showing the instruments cluster by data source rather than by construct, nominate eleven distinct occupations across twelve most-exposed slots, and flip even the sign of the exposure-salary relationship by vintage
- Firm AI-Spend Intensity and Headcount Growth — Ramp × Revelio panel of 21,559 US firms: high-intensity AI-vendor spenders grow headcount ~10% (entry-level ~12%) over the 24 months after adoption while low-intensity adopters show no change — an intensity-gated learning-curve effect, read against Indeed's senior-tilted postings rebound, the Ramp AI Index adoption-breadth cut, and ICONIQ's growth-gated twin (146% median headcount growth in 2026 for 100%+ revenue growers against 10-20% cuts at AI-citing public incumbents in the same year).
- The Household Production Boundary — Google ATLAS's most novel contribution — 86.5% of conversational AI usage happens outside formal work, human time allocation predicts where AI questions go (slope 0.77, ~50% of variance), high-friction bureaucracy over-indexes ~20× with half of those queries outside business hours, and 0.5–5% household time savings values at $15–149B/yr in the US that GDP cannot see by construction
- Human-AI Accountability Redesign — HBR five-pillar prescription: span-of-control redesign, role redesign, performance management reset, decision-rights/escalation/consequences, agentic-unit-not-human-role design — paired with CIVIC-AI's six-condition audit test for whether a completed redesign counts as augmentation at all (two layers: snapshot workflow integrity, longitudinal human development), and its verifiability/reversibility/stakes rule for where the human-agent boundary goes
- Market-Priced AI Exposure (the AI Premium) — Borri-Liu-Tsyvinski: market-implied AI exposure built from 380T tokens of realized OpenRouter consumption — an AI Factor, rolling firm-level AI Betas, and a priced 64 bps/week long-short premium concentrated on frontier/paid use; the implied skill map is orthogonal to task-based exposure measures, and tool-call tokens rising to 52% signal an agentic economy.
- Organizational Complements to AI — The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design changes (David's electrification analogy, Brynjolfsson's paradox) — OpenAI's Codex natural experiment (99.8% vs 16.5% usage of the same model) shows the gap is complements; also home to the HAT substitution model and Kalff & Simbeck's institutional complement.
- Owning Your Externalized Cognition — Garry Tan's ownership axis on skill files: once your judgment is written down as executable markdown it is an asset with a holder, and the same file is either portable career capital or an extraction, depending only on whose repo it sits in — the appropriation counterpart to the cognitive-commons erosion argument, asserted from a keynote stage with no measurement behind it
- Post-Scarcity Macroeconomics — Musk's claim that once digital intelligence acquires end effectors the economy goes quasi-infinite, so money 'won't matter' by 2036: the load-bearing argument is a deflation one — create money slower than output grows and prices still fall — which makes universal transfers non-inflationary and taxation moot; the transition path is the part he concedes he cannot describe
- Procedural Value in AI Decisions — Wang, Sturgis & de Kadt (LSE/Cornell, arXiv 2609.16390): a preregistered paired-profile conjoint on 1,919 US job seekers (30,704 profiles) varying who decides, how often the system errs, and what recourse the applicant has, independently. Human decision authority is worth +0.272 in choice probability, about what cutting wrongful rejections from 30% to 10% is worth (+0.285; difference 95% CI [-0.007, 0.033]); appeal +0.156, opt-out +0.129, explanation +0.128, a bias audit only +0.068. The load-bearing result is a preregistered equivalence null: the value of appeal, the opt-out and the audit does not rise as errors rise (+0.001 / +0.005 / +0.012 over the 10-30% range, 90% CIs inside ±0.05), so procedural preference is not just demand for error correction. Applicants rank rights they can invoke above system-level oversight, and 7.8% doubt AI hiring's legitimacy yet intend to apply
- Returns to Expertise in Agentic Coding — Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the actions and 5× the output per prompt, reach verified success ~2× as often, and abandon stuck sessions far less; every occupation lands within 7pp of software engineers; gains are concentrated novice→intermediate, with mastery adding little
- Seniority-Biased AI Adoption: The Junior Share at Adopting Firms — What happens to the seniority mix inside firms that adopt generative AI. Chandar & Klein Teeselink (Sept 2026) study 1.25B postings and 154M Revelio records across 41 countries with a cross-border peer-adoption instrument. By March 2026, foreign affiliates of adopting companies have a junior share 1.9pp lower than matched controls. Most of that comes from senior growth (+6.7%), not junior loss (−2.5%, n.s.), and it shows up in 23 of 31 countries. The design contradicts Ramp's junior-tilted adopters on the same Revelio seniority coding, and it shows that adopter-vs-non-adopter gaps and aggregate employment changes can differ in sign.
- The Solo-Authorship Rebound — Matsui (arXiv 2607.10780): across 300M+ OpenAlex works and 26 fields, the decades-long decline in solo-authored papers halts or reverses at ChatGPT's November 2022 release — positive trend break in 23 of 26 fields, largest in Engineering (+2.5 pp/yr) and Business (+1.9), absent in Chemistry and Physics and negative in Arts and Humanities. It survives conditioning on author history and is strongest among authors who had never published alone; solo papers stay near their authors' coauthored content while narrowing 23% in breadth and tilting toward computational work. A solo paper is proposed as an observable behavioral trace of AI substituting for a human collaborator — but the design is an interrupted time series with no untreated unit, roughly half the pooled break is venue composition, and the disciplinary ordering, not any single number, is the actual argument
- Task Crossover — OpenAI's Work at the Frontier (800K+ US ChatGPT work messages mapped to O*NET, July 2026): 16.8% of work messages and 43.5% of occupation-specific ones concern tasks historically belonging to another occupation — jobs reorganizing before job descriptions change. Borrowing and lending are separate directions (design borrows 35.2% and lends 1.7%; engineering lends 7.4%), financial calculation and software troubleshooting travel to all seven other groups, and crossover falls as workspace size rises (18.9% at 2–5 seats → 16.3% at 101+)
- Task Saturation: Broad but Shallow AI Diffusion — Google ATLAS's marquee work finding — AI reaches 68% of detailed occupations (88.4% of US employment) but only 21% of the tasks in the median occupation, with end-to-end automation the intent of just 6.5% of non-routine-cognitive conversations vs 26.9% for routine-cognitive; the extensive margin is gated by physicality, the intensive margin concentrates in non-routine cognitive work, and usage over-indexes most on the lowest-expertise cognitive tasks
- The Tragedy of the Cognitive Commons — Lovett (HRD Review, July 2026): professional expertise is a profession-level commons whose regeneration mechanism — entry-level work — AI is removing. Distinguishes Internalized Mastery (built through cognitive struggle) from Distributed Mastery (orchestrating AI), and names the Validation Tether: substantive oversight of AI requires the expertise AI adoption erodes. Its sharpest claim is that junior labor's operational necessity was the hidden governance mechanism all along — regeneration was a side effect of business, never a decision
Derived#
- Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect — Re-works the open question on Jabarian & Henkel's +12% AI-interviewer offer effect (information collection vs the interviewer's discretion to abort) with the vault's second AI-interviewer study, Deng, Liu, Toubia & Jain's market-research RCT. Its randomized AI-vs-static arms split the information side: the zero-discretion static guide is the weakest arm on every content measure, and on this page's arithmetic from Table 7 it is also the more dispersed in outcome (breadth CV 0.24 vs 0.12, 0.31 vs 0.24, 0.13 vs 0.06), so a uniform procedure does not give uniform information. The working ingredient is adaptive probing inside a fixed guide. The abort side does not move. No vault source measures the screen-out channel, and the 2026-08-17 bounds (+16% symmetric, r* = 5.8%) are still the best available. Also a named second system (GPT-5.1) showing the lower-dispersion signature against humans, confounded and n=24, which is weak evidence that the property is not specific to one agent
- Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale — Two-question synthesis on AI-era expertise economics. (1) Stone's specialists-broaden-quickly claim splits into two different goods: tool-in-hand performance breadth is measurably cheap — the concave expertise curve (novice→intermediate captures most of the verified-success gain), every occupation within 7pp of software engineers, and the management edge showing the expertise meta-skills (precision of framing, verify-specification, who-corrects-whom) transfer across domains, so an experienced specialist enters a new domain above the novice floor and AI-assisted onboarding compresses ramp further; but retained-capability breadth is unproven and the only randomized evidence cuts against it — automation-mode gains vanish when the tool is removed and skew to upper ability quartiles, while self-report hides the deficit. Cheap to perform, unproven to internalize; no source measures cross-domain ramp speed for experienced specialists directly. (2) Domain-expert-as-builder now has three evidence tiers: capability parity (measured, within-7pp), market existence (200K+ non-technical customers building ERPs/CRMs, vendor-claimed; AI responsibilities in 28–40% of business job descriptions), but primary-job building as population-level practice remains unshown — every measured population is selection-biased toward adopters, complements gate realized value, and ATLAS's composition shows experts pointing AI at their inexpert tasks rather than non-experts becoming builders. The gating variable is complements plus retained understanding, not capability
- Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence — Reconciles the Founder's Playbook orchestration framings with HBR Kropp et al.'s accountability evidence; "orchestration as workflow design" survives the critique; "orchestration as mental model of agents-as-coworkers" does not; operational checklist for the disciplined founder
- Does the Augmentation/Automation Split Govern Skill at Work? — Verdict: the split transfers to the workplace in sign but is unestablished as a measurement. Four workplace deskilling results sit on the automation row (Budzyń's endoscopists, Dell'Acqua's consultants, Vicente & Matute's 80.7%, Wiles et al.) and two on the augmentation row (Everett's collaborative clinical workflows, Brynjolfsson's support agents) — but all six reach the vault secondhand through one
practitioner-opinionframework paper, none classifies use mode within a design, and none takes an unaided post-measure after AI is withdrawn. The disanalogy that bites is not the elite sample or the one-week horizon: the lab's mechanism is reallocation with time-on-task fixed (−5.3pp writing / +4.4pp reading), while at work AI's headline effect is time saved and nothing in the vault measures what the freed hour buys. Workplace volume also produces a third mode a 35-minute proctored session cannot generate — skipped review (31.3% of PRs merged unreviewed; 81.1% of genuine leaked credentials drawing no reviewer comment) — which is not the automation arm, since the automating students still submitted work they had read. The AEI cannot arbitrate because 'automation share' welds Directive to Feedback Loop and excludes the classifier's own Learning mode by construction, and that axis is drifting toward the hollow arm (43–45% and rising, against the lab's 32%). The one randomized workplace case runs the other way: controlled-variance's recruiters redeployed expertise onto the un-automated stage (discounting the AI's interview signal, −0.047/−0.029) and paid in latency (20→24 days), not skill. - What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators — Answers two open questions as one: the +12% AI-interview offer effect and ATLAS's 22.6% classifier accuracy are both ratios whose denominator was assumed rather than measured. For the first, the paper's own figures bound the abort channel — the arms lose almost the same share of interviews (43% human vs 45% AI), so the symmetric completion-conditional correction moves the effect up to +16%, not down; only the one-sided correction the question proposes flips the sign (−10%), and the screen-out variable it would run on is an unvalidated LLM label whose codebook thresholds on a variable the treatment moves. For the second, yes: leave-one-annotator-out human agreement is the ceiling estimator (ATLAS has the annotations and never computes it), and renormalizing gives 73–82% of the human ceiling at occupation title and 85–94% at major group — while at O*NET-task granularity ATLAS proves no ceiling is estimable, which is why aggregating categories is the ceiling fix
Open questions 100 open
- WaitThe verification tax (45.7% of saved time, >25%) is measured once, on mid-2026 models. Does it fall as models get more reliable, or is it a roughly fixed share of any dividend — the price of using output you did not produce? Falsifiable by the next iteration of this program if it fields the same question on a later window.
- SourceThe survey holds career stage for all 637 respondents (356 senior / 234 mid / 47 early) and publishes no cut by it. Does the expertise gradient other instruments find — heavier, more effective use by the more experienced — appear inside a scientist population when that cut is run? Partially answered (2026-09-23): the Nature Graduate Survey analysis does run a career-stage cut on a scientist population and finds nothing — first-two-years-of-PhD gives −0.057 / 0.019 / −0.083 across the three non-reference attitude classes, none significant — and the authors read the null as evidence that attitudes were not formed by exposure-from-the-outset. Three qualifications keep the question open: the contrast is inside doctoral training (years 1–2 vs 3+), not senior-vs-junior; the outcome is an attitude class, not use intensity or effectiveness; and the same table puts age 34-or-younger at +0.297 (p<0.05) toward the refusing profile, so what little gradient there is runs the opposite way from the returns-to-expertise prior.
- WaitThe LLM/specialized-model complementarity (Level-3 elasticity 0.2) is a snapshot the authors expect to move as frontier LLMs absorb domain tasks in mathematics, genomics and life sciences. Does the elasticity rise in the next iteration — the signal that one family is starting to substitute for the other?
- SourceThe comfort ordering (literature accepted, writing / analysis / design resisted) is equally consistent with a normative account (those tasks carry authorship) and a reliability account (those are the tasks models are worst at). The survey cannot separate them because it never asks about legitimacy or about perceived accuracy. Does an instrument that varies stated model reliability while holding the task fixed, or that asks comfort and legitimacy as separate items, move the boundary?
- SourceRun the trace and the attitude instrument on the same field partition: does a field's rate of detectable LLM-modified text move with its researchers' stated comfort with AI writing, against it, or not at all? Liang et al. report by field (computer science ~22%, mathematics ~9%) and the Nature Graduate Survey data are public on figshare, so this is answerable today by anyone who can map the two field vocabularies — but neither paper does it, and the sign of the correlation decides between "attitudes lag behaviour" and "attitudes are field-specific norms that behaviour respects."
- SourceNative-English-speaker status is the largest coefficient in the covariate model (+1.047 toward 'status quo', −0.777 away from 'all-purpose', both p<0.01) and the paper never mentions it. Is acceptance of AI in science partly a language-access effect — L2 researchers accepting as language support what native speakers do not need — and does it survive controlling for country income and institutional resources?
- AI and Market Power3 open
- SourceIs the non-GenAI null an artefact of a binary adoption measure? The paper's own future-work list asks for "measures of AI intensity" instead of the ICT survey's yes/no, and Firm AI-Spend Intensity and Headcount Growth finds intensity is the whole effect. Would a spend- or intensity-graded adoption variable on the same French and Portuguese panels recover a market-power effect the dummy hides?
- SourceDoes the size/market-share inverted-U in GenAI exposure survive contact with observed GenAI adoption? Exposure here is occupational composition, not use. If small, skill-dense, highly exposed firms turn out not to adopt at rates matching their exposure, the "window of contestability" reading collapses into a statement about who employs analysts.
- SourceDiffusion or consolidation? The 21% GenAI acquisition premium is consistent with technology transfer and with killer acquisitions, and the paper cannot separate them without acquirer-type data. Does post-acquisition patenting or product continuation at acquired GenAI start-ups differ by acquirer size and market position?
- AI Usage Cadences3 open
- SourceTime-of-day rests on IP-inferred location; how much noise do VPNs, travel, and datacenter-routed API traffic inject into the "sleep advice pre-dawn" style claims?
- SourceThe weekend personal-use spike is largest in high-income countries — is that a genuine work/life boundary difference, or a composition effect (who uses Claude for what, where)?
- WaitContinuous sampling is new; are these cadences stable, or will they drift as the user base shifts toward lower-wage tasks (the report's own diffusion trend)?
- WaitDoes the top decile's earnings gap close by years three to four, or does it scar the way recession cohorts do (Oreopoulos et al.)? At eight quarters it is −3.3 log points. Deciles 7–9 have already recovered. Trigger: a later PSEO×LEHD vintage in which the 2023–2024 graduating cohorts reach 12–16 quarters after graduation.
- SourceIs the top-decile effect caused by AI, or by a computing-specific shock (the tech-sector correction) that happens to line up with the most exposed majors? Handa's usage measure moves business majors to the top, and there the drop vanishes. Falsifiable: an event study restricted to non-computing majors in the top quintile of any exposure measure. If Accounting, Finance, Journalism and Marketing graduates show no break at November 2022, the "AI-exposed majors" result is a computing result.
- SourceDoes the grade-signal channel operate in hiring? Falsifiable: link course-level AI-driven grade inflation (the Chirikov / Hausman designs) to graduates' initial employment, within major. If graduates from programs whose grades inflated most see larger hiring declines at the same exposure, signal erosion is part of the mechanism.
- WaitDoes unemployment finally break for later graduating classes? Fairlie & Wu's CPS null covers only the class of 2026, and the authors expect the classes of 2027 and later "might be more affected." Falsifiable: a significant summer-2027 coefficient, on the standard or the sidelined measure, for 22–25 bachelor's holders against 2022–26 and against young non-graduates. If a later CPS vintage stays null while the Census top-decile gap persists, adjustment through sorting, graduate school and self-employment is the whole story at entry. Trigger: CPS basic monthly files for June–August 2027.
- SourceDoes the AI-over-static breadth advantage survive a codebook written independently of the moderator objective, so the AI is not scored against the checklist it was told to cover? The length-controlled and blind pairwise checks point the same way, but neither uses an independent codebook.
- SourceIs the human arm's affect advantage (valence, arousal, dominance, p<.001) a property of a human moderator, or of a live synchronous call? A randomized design with a live human over the same browser voice channel, or a lower-latency full-duplex AI, would separate them.
- SourceThe codified gradient does not survive a college-share control and the tacit gradient does. Is codified/tacit a distinct axis at all, or an education index with a mechanism story attached? Falsifiable: rebuild the codified score from a substrate with no schooling content (O\NET on-the-job-training duration, apprenticeship and licensure requirements, work-experience zones) and re-run the same quadrant plot — if the early-career codified gradient survives, the axis is real; if it collapses, this page is measuring years of schooling. Related (2026-10-01), not an answer: Seniority-Biased AI Adoption: The Junior Share at Adopting Firms finds the junior/senior split at AI adopters within occupation groups, which occupational schooling cannot explain. So the age gradient is not only an education artifact. Whether the codified score adds anything beyond seniority is still untested. Related (2026-10-01), not an answer:* AI-Exposed College Majors at Labor-Market Entry holds schooling level fixed (bachelor's only) and finds that field-of-study exposure still predicts entry outcomes. Level of schooling alone does not explain the young-worker effect. Field content remains a candidate. Codification is still unscored.
- SourceThe index is one LLM's reading of employer-written job descriptions at one vintage, with no inter-rater check of any kind. Does it reproduce under a different scorer, a different prompt, or human raters — and does the quadrant assignment move enough to flip the sign of any age band? Nothing in the source varies it.
- WaitTacit knowledge is defined as what resists being written down, and the current wave of agentic tooling is an industrial effort to write work down — recorded workflows, skill files, agent memory. If capture succeeds, today's tacit premium is a codification backlog rather than a comparative advantage. Trigger to watch: the first occupation where a high-tacit score and a declining experienced-worker headcount appear together.
- SourceDoes the asymmetry regenerate faster than it transfers? The whole human role, under this frame, rests on the answer. Nobody in the corpus has posed it.
- SourceNg prefers the frame because it "gives us a clearer path to helping AI systems get better." That is a reason to adopt the frame, not evidence that it's true. What would distinguish a context asymmetry from a capability gap empirically? (Returns to Expertise in Agentic Coding is the closest thing to an instrument.)
- SourceIf the human's contribution is context injection, is the human replaceable by better context plumbing — memory, retrieval, continuous production telemetry — rather than by a better model? That would put the expiry of human-in-the-loop on the infrastructure roadmap, not the scaling curve. Ng's August 2026 answer, by assertion: the plumbing "does not exist and I don't think exists for the foreseeable future" — an opinion about the roadmap, not evidence about it.
- SourceNg writes from 0-to-1 consumer products. Does the frame survive contact with domains where the missing thing is a concept rather than a fact? Partially answered (2026-09-22) by training novices to think or giving them llms rct (
empirical, preregistered 2×2 RCT, n=1,053) — the first randomized test in this corpus in which a concept, not a fact, is experimentally supplied to the human side. Half the sample gets a ~6-minute causal-reasoning training (mechanisms, falsifiable boundary conditions, explicit cause→effect chains); half gets ChatGPT Edu; the cells cross. The concept takes: mechanism identification rises ~0.55 SD, falsification logic ~0.85 SD, idea diversity ~0.47 SD between solutions, and none of it is crowded out when a model is in reach. The concept does not pay: on the graded output its coefficient is negative (−0.111, p<0.01) while LLM access is worth +0.862 on a control mean of 2.09. So the frame survives, but only by narrowing to a conditional — what the human supplies is decisive only where the evaluator demands it. Here it was not: the rubric loaded positively on the model's coherence and idea count and negatively on falsifiability (−0.109), mechanisms (−0.166) and distance from the modal answer (−7.676, ≈ −0.29 points per SD of novelty). Two limits on how far to take this. The subjects are novices with no domain context at all, which is the opposite end of the population Ng writes about, and the task was chosen precisely because the model is already well-trained on it — so it tests the frame where the context gap is smallest by construction. And the study measures assisted output, never withdrawing the tool, so it says nothing about whether the concept was internalized.
- NowHow much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them. Partially answered (2026-08-17) by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators — and the direction is the opposite of the one the question anticipates. Summing this page's own interview-type shares shows the two arms lose nearly the same fraction of interviews to early termination (43% human — 25 screen-out + 9 disengaged + 9 unavailable — against 45% AI — 7 + 12 + 14 + 7 technical + 5 refusal). What differs is composition, not volume: the human arm's terminations are recruiter-initiated and conditioned on a stated disqualifier, the AI arm's are applicant- or machine-initiated and conditioned on nothing about fit. So the correction the question proposes is asymmetric — removing human screen-outs only gives 11.60% vs 10.47% (−10%, sign flipped); removing the AI's technical failures and refusals only gives 8.70% vs 11.06% (+27%); removing all early-terminating types in both arms gives 15.26% vs 17.70% (+16%, larger than the published +12%). The break-even bound: the abort channel accounts for the whole effect iff the differentially screened-out applicants would have converted at
r* = 1.03/18 = 5.8%, i.e. two-thirds of the human arm's own 8.70% base rate — implausible for applicants disqualified on a non-negotiable requirement (location, visa, rehire status), though the paper never publishes the early/midway/late screen-out split, and Late Screen-Out is by definition a near-complete interview. Three things keep it open, and all are data the paper does not report: (i)ritself, and the offer rate among early-terminated interviews by arm and type; (ii) per-arm evaluable-record rates over the randomized denominator — the 25%/7% shares are computed over 34,109 transcripts described as "a subset of all interviews conducted," and the AI arm plausibly interviews more of its assignees given time-to-interview fell 0.51 → 0.32 days; (iii) whether the AI agent possessed an abort action at all — the classifier's codebook requires that the "recruiter states the reason for ending the call due to disqualification," and the paper never says the agent was authorized to disqualify. Every correction above also conditions on a post-treatment variable, so they bound the channel rather than identify it; instrumenting the screen-out decision remains the only design that would. Evidence note (2026-10-01, AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins): a market-research RCT randomizes an adaptive AI moderator against a same-guide static interview and finds the gain comes from adaptive probing, not standardization. That splits the "information collection" side of this question into standardization vs adaptivity. It contributes nothing on the abort side, since no arm could screen out and none was measured, and its AI-vs-human contrast is not randomized. The decomposition asked for here is still unaddressed. Partially answered (2026-10-01) by Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect — the information side now splits in two, and the randomized half of that split points one way. In Deng et al.'s AI-vs-static arms, zero-discretion standardization is the weakest arm on content. On that page's arithmetic from Table 7 it is also the more dispersed in outcome: breadth CV 0.24 vs 0.12, 0.31 vs 0.24, 0.13 vs 0.06 across the three blocks where probing was not capped (descriptive, ceiling-compressed, no variance test). This page's channel is better named adaptive probing within a consistently followed guide than "standardization". In the hiring paper, number of exchanges, the feature closest to probing, is one of the positive offer predictors (correlational, human arm only). What remains is the whole abort side: no vault source measures a screen-out channel, the 2026-08-17 bounds (+16% symmetric, r* = 5.8%) stand as bounds, and items (i)–(iii) above are still unreported. The cleanest design that would settle it is a pre-interview eligibility screener identical in both arms, as the market-research study used. - SourceThe AI system is never identified — no model, vendor, or version, only that Google Cloud supplied infrastructure. Is "controlled variance" a property of a 2025-generation voice agent under this firm's prompt, or of AI-conducted interviews generally? Nothing in the paper lets a replication know what it is replicating. Partially answered (2026-10-01) by Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect — on generality only, and weakly. A fully named second system (GPT-5.1, temperature 0, on a market-research startup's platform) shows lower coverage dispersion than expert human moderators in a different task (breadth SD 0.91 vs 1.61, 0.80 vs 0.88, 0.35 vs 1.51). That contrast is not randomized (n=24, separate later sample). Identification has not moved: this study's agent is still unnamed.
- SourceRetention ≥1 month is both the quality proxy and the metric the recruiting firm is paid on by its clients. Does the AI advantage survive on an outcome the intermediary is not compensated for — client-side performance at 12 months, promotion, or wage growth? The four-month estimate, the longest horizon measured, already fails significance under recruiter clustering.
- NowHow much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them. Partially answered (2026-08-17) by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators — and the direction is the opposite of the one the question anticipates. Summing this page's own interview-type shares shows the two arms lose nearly the same fraction of interviews to early termination (43% human — 25 screen-out + 9 disengaged + 9 unavailable — against 45% AI — 7 + 12 + 14 + 7 technical + 5 refusal). What differs is composition, not volume: the human arm's terminations are recruiter-initiated and conditioned on a stated disqualifier, the AI arm's are applicant- or machine-initiated and conditioned on nothing about fit. So the correction the question proposes is asymmetric — removing human screen-outs only gives 11.60% vs 10.47% (−10%, sign flipped); removing the AI's technical failures and refusals only gives 8.70% vs 11.06% (+27%); removing all early-terminating types in both arms gives 15.26% vs 17.70% (+16%, larger than the published +12%). The break-even bound: the abort channel accounts for the whole effect iff the differentially screened-out applicants would have converted at
- Conversation Artifacts3 open
- SourceTokens are a proxy for both compute cost and output value, but verbose models inflate tokens per unit of intent (the same critique Conversation-to-Delegation Shift raises); how much of "compute tracks value" is genuine value vs. models simply emitting more?
- SourceThe reading-level "+1 year" gap may be register (terse prompts, polished replies) rather than substance; can it be separated from genuine elevation of content?
- SourceArtifact classification is first-party and single-model-graded; do the 30+ categories and the work/personal/coursework split survive independent replication?
- SourceThe token-share metric rewards verbose agentic output. How much of the 99.8% / 63.3% / 16.5% spread is a genuine work shift vs. agentic tools simply emitting more tokens per unit of human intent?
- SourceOpenAI-internal is a frontier preview by assumption. Does the external organizational curve actually trace the OpenAI path (the paper's implicit claim), or does it plateau where adoption frictions don't vanish? Partially answered (2026-09-22) by mckinsey state of ai 2026 road to roi (
empirical, self-reported; 1,719 respondents in 97 nations, fielded May 4 - June 8 2026): at organizational granularity the answer is both, split by firm size. The share of organizations scaling AI agents in at least one function rose 27% → 40% year over year among those above $1B in annual revenue, and 21% → 22% — flat — among those below it (Exhibit 2), with the same size gap on software coding agents (31% vs 17% at least scaling, Exhibit 3). So there is no single external curve to trace: the large-enterprise margin is advancing fast enough to be consistent with the frontier-preview claim, while the smaller-organization margin has plateaued for a year at just over one in five, which is what adoption frictions that do not vanish look like. Three limits keep it partial. The quantity is wrong — deployment phase reported by one respondent, not a share of work delegated, so it cannot be compared with the 99.8% / 63.3% / 16.5% token shares numerically. The time base is one year. And "at least scaling" is the respondent's own judgment on an instrument whose coding-agent rows carry its highest don't-know share. A second reading, on the right quantity and the wrong product (2026-09-23): how organizations use ai (empirical, OpenAI administrative records for organizations adopting 2024-01-01 to 2026-03-31) supplies the external curve in the same unit as the frontier-preview claim — output tokens — rather than in deployment phase. It rises steeply and does not plateau: ~7x aggregate June 2025 → March 2026, ~4x inside the fixed cohort that had adopted by June 2025, with every cohort accelerating simultaneously in early 2026. On that instrument the external organizational curve is deepening, not stalling, which leans toward the frontier-preview reading and away from the small-organization plateau McKinsey found. Two reasons it does not settle the bullet, and they pull opposite ways. The product is wrong: the series is overwhelmingly ChatGPT, not Codex (the authors say so and defer the agentic cut to this page's own source), so it measures the conversational base deepening rather than delegation spreading. And the population is wrong in the same direction as the original claim — ChatGPT Enterprise buyers only, skewed toward large US public firms, which is precisely the segment McKinsey found advancing. The organizations that plateaued are the ones this instrument cannot see. - Source"Asking is half of ChatGPT, doing is most of Codex" — but the two tools self-select different work. How much of the asking→doing contrast is the shift itself vs. routing pre-existing "doing" tasks to the tool built for them?
- SourceTime-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
- SourceGains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
- The augmentation/automation choice is endogenous to incentives (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half? (Andrew Ng's August 2026 "as most commonly used" assumes the mix is automation-dominated; the experiment's own logs run 49% augmentation / 32% automation, so the assumption is unmeasured, not confirmed.) Partially answered (2026-09-22) by training novices to think or giving them llms rct — narrowly, and on a lever the question did not name. A preregistered 2×2 (n=1,053) randomizes a ~6-minute causal-reasoning training against a placebo, crossed with ChatGPT Edu access, and finds the trained skill is not crowded out by the tool: every Causal × GPT interaction on reasoning style is positive and significant on top of the main effects (mechanisms +0.129, falsifiability +0.203, coherent logic +0.228), and causal training's effect on idea diversity (+0.47 SD between-solution) is undiminished when a model is in reach. So something upstream can change what novices do with an LLM — the first randomized demonstration of that in this vault. Three reasons it is a partial and not a retirement. (i) It is training, not incentive or interface design — the only interface-like element, the one-line nudge "Remember the skills you learned during the game," is bundled inside the training arm and cannot be identified separately. (ii) It never measures the mix: use mode is not classified, conversation logs were collected but not analysed, and the outcome is reasoning style in the assisted output, not augmentation-vs-automation behaviour. (iii) The deskilling clause is untouched — there is no unaided post-measure, no tool withdrawal, no retention test, and the authors state flatly that they cannot say whether the LLM's advantage reflects knowledge acquired or output procured. It also supplies the question's inverse, which is the more useful half: the incentive that exists today pushes the other way. In the same paper's lasso, the evaluation rubric loads negatively on falsifiability (−0.109), mechanisms (−0.166) and distance from the modal answer (−7.676, ≈ −0.29 points per SD) while loading positively on the coherent, idea-dense text the LLM produces — so under a standard rubric the trained behaviour is scored down and the outsourced one up. Any incentive-design answer has to beat that gradient first. Partially answered, on both named levers (2026-10-01): designing against deskilling metacognitive feedback (
empirical, preregistered, N=704) randomizes one interface lever and one incentive lever against the same classified use mode. The two split cleanly. - Interface works. Per-item metacognitive feedback on the learner's own AI use cuts answer offloading (OR 0.47, one-sided p=.026). It moves the request mix toward guided help and intermediate results (answer offloading 58.2% → 51.2% of assisted items) and lifts unaided transfer (OR 1.51, p=.030). This is the reversal of the deskilling half, at least back to the no-AI baseline (0.67 vs 0.65).
- Incentive fails. A points schedule penalizing answer requests reduces use (OR 0.39) but not offloading (OR 0.66, n.s.). Among assisted items, offloading's share rises to 62.8% because guided help is what gets dropped.
- SourceDoes the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer? (Promoted 2026-09-05 lint: the vault now holds the workplace-side pieces to test against — the observed automation-share drift on Anthropic Economic Index, the recruiter case in the controlled-variance study, the self-report on The Automation–Optimism Link, the offloading account on AI Brain Fry.) Partially answered (2026-09-05): Does the Augmentation/Automation Split Govern Skill at Work? — the sign transfers. Budzyń's endoscopists, Dell'Acqua's consultants, Wiles et al. and Vicente & Matute all reproduce this page's removal signature (performance high with the tool, no advantage without it) on non-students, while Everett's collaborative clinical workflows and Brynjolfsson's support agents supply the augmentation row — but all six reach the vault secondhand through The Tragedy of the Cognitive Commons, none classifies use mode inside a design, and none takes an unaided post-measure. The answer to the second clause is not "too unlike to transfer": the disanalogies that actually bite are that this design holds time-on-task fixed (its entire mechanism, −5.3pp writing / +4.4pp reading) while at work AI's headline effect is time saved, and that workplace volume generates a third mode this experiment cannot produce — abdication, the review step skipped outright (Acceleration Whiplash's 31.3% unreviewed PRs, Security Debt of Agent-Generated Code's 81.1% uncommented credential leaks), which the automation arm here is not, since those students still submitted work they had read. Still open, and no longer answerable by synthesis: it needs a workplace study that classifies use mode within-subject from logs and measures unaided skill after AI is withdrawn. Checked and recorded as a restatement, not a partial (2026-09-22): when does ai augment work workflow framework (CIVIC-AI workshop whitepaper,
practitioner-opinion, no study of its own) makes this question one of six conditions for calling a workplace deployment "augmentation" at all, and argues the transfer runs through a selection effect rather than an analogy: the tasks most amenable to AI delegation are "high-volume, procedurally defined and verifiable," which are "exactly those tasks through which junior workers develop their tacit judgement needed to evaluate outputs, catch failures and push back on algorithmic recommendations." It also writes down the instrument this bullet says is missing, in close to the same terms — a workflow record carrying total effort including verification and repair, plus repeated longitudinal measures of mastery, progression and agency, complemented by employee-level evidence because "formal adoption data may miss unofficial AI use." Nothing in it is measured, no use mode is classified anywhere, and no unaided post-measure is proposed, so the bullet is unchanged: what it gains is a second independent specification of the missing study, co-signed by a labour ministry that would be positioned to run it.
- SourceBinned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the report checks robustness with a ≥60%-of-tasks indicator, but the levels remain self-reported)?
- SourceReported exposure exceeds observed partly because the survey reaches heavy users; what does the reported/observed gap look like in a representative sample? Sharpened: realized-consumption measurement (Borri-Liu-Tsyvinski) adds a market-implied instrument built on 380T tokens of actual paid requests — but it is skewed toward developers/sophisticated users (OpenRouter is ~2% of global tokens), a different non-representativeness than the survey's heavy-user skew. The lesson: no current AI-exposure instrument is representative; each collection mechanism biases in its own direction, so the reported/observed/market-implied gaps are partly artifacts of who each method reaches. A representative census remains the open target. Sharpened again by Steele & Cruz, which makes the population question concrete across seven instruments at once — 2,000 MTurk respondents (Felten), GPT-4 as rater (Eloundou), whoever files AI patents (Webb), Crowdflower workers against a rubric (Brynjolfsson), 70 jobs hand-coded by two researchers (Frey), and Claude/ChatGPT users in 2025 (Massenkoff; Steele & Cruz). Every instrument biases toward whoever it reaches, and the head-to-head shows the consequence is not a level shift that a rescaling would fix: the instruments produce different job rankings, nearly disjoint most-exposed lists, and opposite signs on the exposure-salary gradient. A third sharpening, on the denominator itself (2026-09-22): google atlas ai in science september 2026 runs telemetry and a self-report survey on the same occupation — scientists — and states the harder version of this question in its Appendix 3: "There is no reliable population frame for scientists in either the UK or the US," so its 637-respondent panel is reported unweighted with no margin of error quoted, and the authors name selection toward AI-enthusiastic respondents as a reason both the adoption rate and the 6.9-hour time saving may be overestimates. For an occupation this well documented — publications, grants, job postings, professional bodies — no frame exists to weight to. So "a representative census remains the open target" is optimistic about the target: for many occupations the representativeness gap is not merely unmeasured but currently unmeasurable, and the telemetry side is no better off, since its science slice is a classifier's guess at who a user is.
- SourceDoes averaging across instruments reduce error or merely blend incompatible biases? Steele & Cruz's cross-model average is a diversification argument, not a validated one — no instrument in the set has been scored against realized labor-market outcomes, and the average's apparent stability partly reflects dropping the two instruments that disagreed most (Massenkoff for redundancy, Frey for anomaly). Falsifiable: score all seven, plus the average, against subsequent occupation-level employment and wage changes. Complication (2026-08-04): Indeed's postings data shows the exposure–outcome relationship changing sign between windows on a single instrument (2022–2026 negative, 2025–2026 positive), so any such validation scores the window as much as the instrument, and a scoring period must be pre-specified rather than chosen after the fact. Partially answered (2026-09-22): canaries in the coal mine august 2026 (
empirical, ADP payroll microdata through June 2026) runs six of these instruments — Eloundou 2024, Brynjolfsson 2018, Felten 2021 and its two 2023 variants, Tomlinson 2025 — against the same realized employment change, and gets 10–13% declines for the most-exposed quintile at ages 22–25 across all of them. So the answer to "does averaging blend incompatible biases" is conditional on what an analysis consumes: binned into quintiles and read at the extremes, the instruments are nearly interchangeable; ranked occupation by occupation, they are not. It is not the proposed test — no instrument is scored against another, no accuracy statistic is computed, and the window is fixed post hoc at Nov 2022 → Jun 2026, which the complication above says is itself a choice. See the section on it. Checked and largely declined (2026-09-22): Zhu (empirical) scores benchmarks rather than exposure instruments for construct validity, and the one line that transfers is a warning about the denominator rather than an answer — the capability battery these instruments assume is near-collinear (one factor at 74.5% of common variance) and its leading axis is substantially a release-date trend, so "AI capability" as an instrument input is a moving target whose movement is mostly calendar. The paper fixes its own scope to internal validity and runs no instrument against any realised outcome, so the falsifiable test above is untouched. Complication (2026-10-01), on the wage half: indeed ai exposure advertised pay scores one instrument (GSTI) against realized advertised pay and finds the answer turns on the composition control — +5.7% (p<.05) with occupation fixed effects, +2.4% (n.s.) with occupation × seniority — so a wage-validation test must pre-specify whether seniority mix counts as outcome or confounder, just as the window must be pre-specified. See The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay. Complication (2026-10-01): scored against realized entry outcomes by college major, Eloundou and Eisfeldt agree with Handa usage through the 80th percentile, then split: Handa ranks business majors above computing majors and the top-decile drop disappears (AI-Exposed College Majors at Labor-Market Entry). - WaitThe experience gradient rests on what workers believe AI can't do (judgment, relational work) — a belief that could be either durable comparative advantage or the next capability to fall. Which, and when? Partially answered (2026-09-22), on the first half only: canaries in the coal mine august 2026 supplies the first outcome measurement of the gradient the survey only elicited as belief. Laying a codified/tacit knowledge index over ADP payroll (Codified vs Tacit Knowledge Exposure), employment grows faster for experienced workers in high-tacit occupations through June 2026, and the complementarity coefficient from the Anthropic Economic Index usage split is positive and significant only at 41–49 (+0.024) and 50+ (+0.015) while the automation coefficient shrinks monotonically from −0.098 at 22–25 to −0.006 at 50+. The belief is currently correct: through mid-2026 the experienced are being complemented and the young substituted, and the tacit gradient survives an education control that kills the codified one. The "and when" clause is untouched, and is the whole question — four years of payroll cannot distinguish a durable comparative advantage from a capability that has not fallen yet, and the paper's own framing ("we do not make predictions about future impacts") declines to try. Stays
#oq/wait; the trigger to watch is the first vintage in which the tacit premium for experienced workers narrows rather than widens.
- SourceWhat operational mechanism converts intensive AI spend into hiring? The paper establishes the correlation (adopters, especially intensive ones, grow) but explicitly cannot say why — product acceleration, sales productivity, engineering leverage, support automation, faster analysis, or new business lines are all candidates, and the firms that cracked it have no incentive to share. Related evidence (2026-08-04): the monthly AI Index narrows what the top of the intensity distribution is buying — the $248-PEPM cohort is defined by paying model-serving and inference platforms, i.e. building on APIs rather than buying more seats. That is a characterisation of the spend, not of the mechanism, but it points the candidate list toward engineering/product leverage and away from enterprise chat rollout. A named channel, from inside the firms (2026-09-22): iconiq 2026 state of scaling observes ICONIQ's portfolio running backfill suppression rather than restructuring — attrition used to 'avoid, delay, or down-level backfills,' explicit no-backfill policies 'tied to AI-driven productivity improvements in sales, content, operational, and analytical functions,' and hiring freezes held while productivity gains are evaluated before incremental headcount is approved. This is the first mechanism in the corpus that would be invisible to this page's instrument: it changes net headcount with no layoff event and no hiring event, so a spend-and-headcount panel records only a flat line. It is also entirely unquantified — an investor's qualitative observation of its own book, no share of companies, no effect size — and it is a mechanism for headcount not growing, whereas this bullet asks what converts spend into growth. Worth carrying as a candidate for the other tail of the distribution.
- WaitDoes the effect diffuse beyond Information as adoption cohorts mature? Significant gains are, so far, an Information-sector phenomenon; the authors intend to update with later cohorts and post-24-month windows. Will professional services, finance, and non-technical sectors follow, or is the coding-agent workflow special? Independent corroboration of the sector boundary (2026-08-11): OECD AI Papers No. 62 finds the markup premium from AI patenting is significant only in ICT (
AI × ICT+7.95%, while standaloneAIturns negative with fixed effects) across ~600K firm-years in 21 European countries — a different outcome (markups, not headcount), a different instrument (patents, not spend), a different continent, and the same sector line. Their reading is the sharper version of the question: AI pays where it is the firm's output, not where it is an input. Not an answer to diffusion-over-time, but two instruments now agree on where the effect currently lives. - WaitIs the entry-level growth durable or a lead-indicator that later reverses? Gains compound through month 24 on thinning samples; whether the +12% entry-level result holds (or inverts toward the Brynjolfsson "Canaries" pattern) as high-intensity adopters mature past 24 months is unresolved. Countervailing signal (2026-08-04): Indeed Hiring Lab finds the May 2025 – May 2026 software-postings rebound is 71% senior roles, on a later window than this panel's average and on the demand flow rather than the headcount stock — not an answer (different unit, different population, no control group), but the first vault evidence pointing the other way on composition. Partially answered (2026-09-22): canaries in the coal mine august 2026 (
empirical, ADP payroll microdata through June 2026) settles the Canaries half of the comparison — the pattern this question asks whether the +12% might invert toward has itself continued and widened, to a 19% kept-pace shortfall for 22–25-year-olds in exposed occupations, driven by hiring rather than separations, and concentrated where AI usage is automative. It does not settle the durability question, and the reason is structural rather than a data gap: ADP firm identifiers are anonymized, so Canaries cannot observe adoption at the firm level at all and cannot say what happened inside intensive adopters past month 24. The two instruments measure different objects and are reconciled rather than averaged (see the section above). Contradicted on composition (2026-10-01), not on durability: Seniority-Biased AI Adoption: The Junior Share at Adopting Firms finds adopters' junior share falling through March 2026 on the same Revelio coding, and falling further with adoption intensity. The one overlapping cut — state-level adoption tiers from Anthropic Economic Index usage — leans against this panel: the young-worker decline is sharpest in leading-adoption states (about −19% for the most-exposed quintile) and smallest in emerging-adoption ones. Geography is a weak proxy for firm adoption, so this sharpens the question rather than answering it: the durable test remains a firm-level panel that can follow the same adopters past 24 months. A second Indeed series leans the same way (2026-10-01): indeed ai exposure advertised pay reports the entry-level share of salaried postings in the most-exposed occupations falling 29% → 10% (2021 → 2026) — economy-wide flow, not adopter stock, so still not the test; detail on The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay. - WaitIs the top-1% per-employee-spend decline a seasonal dip, a vintage artifact, or the start of a plateau in the intensity distribution's top tail? The three readings separate in the data: a seasonal dip reverses by October (Ramp names August engineer holidays and observes the same dip each November–December), a vintage artifact shrinks once August is revised on the same footing as July — this page's cross-edition arithmetic puts the like-for-like fall at roughly 2.6% against a published 9.7% — and a genuine plateau shows up as a second consecutive decline in an already-revised series. Trigger: the October and November 2026 editions of the Ramp AI Index, which republish the same cohort series.
- WaitIs the premium a durable risk price or an early-diffusion artifact? The authors flag the short, fast-moving sample and call it "the current price of AI exposure." Does the transition-risk premium persist, shrink, or invert as AI diffusion matures?
- SourceHow much does the developer skew move the answer? OpenRouter's slice is unrepresentative; would a representative realized-consumption panel (if one existed) price the same firms and skills, or is the frontier/intensive-margin concentration partly a sampling artifact of who uses OpenRouter?
- SourceWhy is the market-implied skill map orthogonal to every task-based measure (<2% variance)? Is market-implied exposure capturing genuinely different information (forward-looking rents, complement/substitute value rather than technical automability), or is it noisier — and which should labor-impact forecasts trust? Partially answered: Steele & Cruz's seven-instrument head-to-head (see Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated) removes the framing that made this look like an indictment — the task-based measures are largely orthogonal to each other too. Only two pairs correlate strongly and both share a data source (ρ=0.89 between the two Anthropic-usage instruments; GPT-4-rated and MTurk-rated theoretical capability next). Webb's patent measure, Brynjolfsson's ML rubric, and Frey's bottleneck model show "very little correspondence with each other or with later measures," and the exposure-salary gradient flips sign between the older and newer instruments. So being uncorrelated with the task-based family is not evidence of noise — there is no coherent family to be uncorrelated with. The second half stays open, and is now harder: no instrument in either camp has been scored against realized labor-market outcomes, so "which should forecasts trust" has no empirical answer yet. Context added 2026-09-22, from the capability side of the same question: Zhu (
empirical, single-author unrefereed preprint) runs the first construct-validity audit of the benchmarks an exposure instrument implicitly assumes, on a hash-pinned single-harness leaderboard snapshot of twelve of them over 421 model configurations, and finds the mirror image of the exposure picture: mean off-diagonal Spearman rho = 0.79 with every pair positive, one retained factor holding 74.5% of common variance, ten of twelve benchmarks in a single correlation cluster — and that one axis tracking model release date at R-squared = 0.505. So capability measurement is about as near-unidimensional as measurement gets, while exposure measurement is mutually orthogonal. The orthogonality on this page is therefore not inherited from a genuinely multidimensional capability frontier; it is manufactured somewhere in the mapping from capability to tasks, occupations and prices, which is where the search for an explanation belongs. Two limits on how far this can be pushed: those are benchmark scores, not exposure instruments, and the paper fixes its own scope to internal validity — adoption, revenue and any realised economic outcome are explicitly outside it, so it adds no instrument-to-outcome validation to the second half of this question. - WaitThe agentic premium is only "early evidence" (imprecise). Does a positive agentic premium survive a longer sample, and does the falling price-per-agentic-token (caching + cheap-model routing) erode the dollar-side signal even as token volume explodes? Context (2026-10-01), not an answer: Epoch (
empirical) measures the cheapest cost of a fixed benchmark score falling ~47% per quarter. The fall is fastest right after a level becomes SOTA (66% per quarter) and slowest on agentic coding (SWE-bench Verified, 27.5%). The frontier, paid-for usage where this page finds the premium concentrated is where prices fall fastest. Whether that erodes the dollar-side signal depends on how much of the frontier rate buyers actually capture, and nobody has measured that (The Price of Fixed Capability). - WaitDoes "Science most negative / interaction most positive" hold out of sample? The finding cuts against the intuition that AI automates cognitive work last; is the market right, or pricing a transient narrative?
- SourceThe "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign workflows quickly, or does the low cost of tool adoption mask slow, expensive process redesign (the real complement)? Partially answered — and the split is between the two halves of the question. Kalff & Simbeck find both happening at once in the same 410 firms: the low-threshold half diffuses faster than the organization, with 183 of 410 respondents using AI informally on personal devices regardless of employer policy, while the half that needs process redesign stalls exactly where the electrification analogy predicts — advanced analytics "seldom economically or logistically viable" without centralised data and standardised processes, and 20.2% of departments using no AI tool at all. So tool adoption does mask the absence of process redesign, but not by making it look fast: the two run on separate tracks, and the visible one requires no organizational change to happen. Still short of settling it — self-reported, cross-sectional, one function, one country, and no measure of redesign speed where it does occur. Extended (2026-09-22), with the first national-survey reading: when does ai augment work workflow framework (
practitioner-opinion, citing Singapore MOM statistics it did not collect) reports firms naming role redesign (18.9%) about three times as often as reduced headcount (6.2%) at 28.5% adoption — the redesign half of this question happening at population scale rather than in one favorable case. It does not close the question, and the authors supply the reason: employer-survey statistics "do not capture the full costs of verification, exception handling, rework, recovery, or unofficial AI use," so "role redesign" as a survey checkbox cannot distinguish genuine process redesign from re-labelled tool adoption — which is the exact distinction the bullet turns on. What it does establish is that the distinction is not merely this vault's worry: a labour ministry co-signs the claim that its own instruments cannot resolve it. One more corroboration of the slow half, weak in tier but strong in shape (2026-09-22): adoption telemetry enterprise ai production signals treats non-redesign as the field's normal condition rather than a defect — it sets its depth thresholds so that no cohort can read "healthy" at the recurring-use-to-workflow-integration boundary, on the stated belief that the ~2% workflow-embeddedness rate (ActivTrak, secondhand, 120,620 workers) is "an industry-wide cliff, not a cohort-specific defect."practitioner-opinionwith no real deployment data, so it is a design assumption and not a measurement; what it adds is that an instrument builder priced the redesign half as near-zero and made that the default. The intensive-margin half, measured rather than surveyed (2026-09-23): how organizations use ai (empirical, OpenAI administrative records, 1,764 organizations / 17.4M messages) shows tool adoption and organizational reach coming apart inside the same firms, in the direction this bullet predicts. Among ChatGPT Enterprise adopters, larger firms use less per employee — weekly active users per employee −0.032\\\ and output tokens per employee −0.667\\\ on lagged log employment — while messages per weekly active user is flat (−0.002, not significant). The entire size penalty is breadth, not intensity: the workspace is bought, and a smaller share of the workforce ever shows up. That is tool adoption masking the absence of reach, stated as a coefficient. The paper's own conclusion is this bullet's premise in the authors' words — "adoption is only the beginning of deployment" — and its growth decomposition shows the other half moving too: output tokens grew ~7x across all customers between June 2025 and March 2026 and ~4x within the cohort that had already adopted by June 2025, so roughly half of the growth is deepening inside existing adopters. Still short of settling it, and for a reason no sample size fixes: the paper measures no redesign of any kind — no process, workflow, or organizational-change variable exists in it, and no outcome does either — so it can show that reach lags adoption but cannot say whether redesign is what closes the gap. - SourceWhich complement is the true binding constraint — access/permissions, skills, or review capacity? The paper lists all; it doesn't decompose their relative weight. Sharpened, not answered: the list itself is incomplete. Kalff & Simbeck's German evidence adds an institutional complement (works-council co-determination under BetrVG §87(1) no. 6, EU AI Act high-risk classification) that is not internal to the firm at all and that determines which capability is adoptable rather than how well it is used — so any decomposition needs a fifth term whose weight varies by jurisdiction rather than by firm. Their own most-cited internal blocker is data centralisation, which is closest to "access." A candidate term the list does not have, from a framework that cannot weight it (2026-09-22): when does ai augment work workflow framework adds oversight-competence maintenance — reviewer competence "does not persist automatically. It requires organisations to deliberately assign workers enough substantive review work and enough exposure to AI failure to keep the verification skill current." On that reading "skills" and "review capacity" are not two items on the list but one item measured at two times, and the binding constraint is a stock that depreciates rather than a headcount — which changes what a decomposition would have to estimate (a depreciation rate, not a coefficient). It is
practitioner-opinionwith no data, so it reorders the candidates rather than weighting them. Its one distributional datum points at access/skills: among young Singaporean workers, 38% of those with secondary qualifications used new technology at work against 74% of degree holders. The first measured ranking of candidate complements, on the adoption margin only (2026-09-23): how organizations use ai regresses ChatGPT Enterprise adoption on fiscal-year-2021 intangible stocks per employee — three to four years before the FY2024–25 adoption window, so the measure cannot be an artifact of adopting — and the coefficients separate: SG&A stock 0.020\\\ (0.005), capitalized software 0.008\\* (0.003), R&D stock 0.004\\\ (0.001), all reconciled againstpdftotext -layout(the docling parse of that table is shifted). Accumulated organizational* capital beats the literal software-investment measure by 2.5x and the innovation measure by 5x, which points the decomposition at this bullet's "access" and "skills" terms and away from technical infrastructure. Three discounts keep it partial, and the first is decisive. It ranks predictors of adoption, not of value* — the question asks which complement binds the gain*, and this paper measures no output, productivity or quality outcome at all. Outside technology and high-R&D sectors only the R&D coefficient keeps significance, so the ordering's generality is unestablished. And the population is US public companies that bought one vendor's product, with the non-adopter group contaminated by a disclosure-motivated random-sampling step. Nothing here touches the institutional fifth term or the oversight-competence depreciation rate; Compustat has no line for either. - SourceIf complements, not capability, gate value, does model progress have diminishing near-term returns until orgs catch up — and how long is that lag for agentic AI specifically? Bounded, not answered (2026-09-23): how organizations use ai supplies the first long usage series from inside enterprise buyers and it runs the wrong way for a simple diminishing-returns story — output tokens grew roughly sevenfold between June 2025 and March 2026, about half of that inside firms that had already adopted, and the acceleration in early 2026 hit every adoption cohort simultaneously, which the authors read as a product- or market-level development rather than the normal post-adoption ramp. A cohort-independent acceleration is evidence against "orgs are the rate limiter and capability waits on them" over a nine-month window. Three reasons it does not close the bullet: tokens are an input measure, not value, and the paper has no outcome variable, so rising consumption is equally consistent with a lag in realized returns; the window is nine months against a question about multi-year catch-up; and the series is overwhelmingly non-agentic by the authors' own Appendix A1, so it says nothing about agentic AI specifically, which is what this bullet asks.
- SourceHAT's P2 (middle-management vulnerability) is the vault's most-quoted-but-least-tested substitution claim, and nothing here can settle it: Ramp × Revelio resolves seniority only to entry-level / non-entry / manager-plus, which cannot distinguish "middle layers thinned first" from "manager-plus grew more slowly than the bottom." Falsifiable with a level-resolved employment panel (Revelio or matched employer-employee data cut by reporting depth, not seniority band) tracking layer counts before and after intensive AI adoption. The same instrument would settle P4, since it could compare regulated against unregulated industries on the same measure — though P4 now has its first field evidence from German HR (see the ledger row), which confirms the direction on self-report and leaves the panel test outstanding.
- WaitCorollary 5 claims irreversibility: once an AI agent is strictly cheaper on risk-adjusted grounds, the optimal allocation never reverts, given fixed human costs and negligible switching costs. The vault has no case either way, and this is falsifiable only by a future event — a documented instance of a firm re-staffing with humans a role it had already automated, for reasons other than a regulatory shock or a rise in AI risk (both of which the corollary's own conditions exempt). Watch for it in the same firm-level adoption panels; a reversal with human costs and regulation unchanged would falsify the corollary, and a long clean run of non-reversal would be weak support for it.
- SourceThe "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign workflows quickly, or does the low cost of tool adoption mask slow, expensive process redesign (the real complement)? Partially answered — and the split is between the two halves of the question. Kalff & Simbeck find both happening at once in the same 410 firms: the low-threshold half diffuses faster than the organization, with 183 of 410 respondents using AI informally on personal devices regardless of employer policy, while the half that needs process redesign stalls exactly where the electrification analogy predicts — advanced analytics "seldom economically or logistically viable" without centralised data and standardised processes, and 20.2% of departments using no AI tool at all. So tool adoption does mask the absence of process redesign, but not by making it look fast: the two run on separate tracks, and the visible one requires no organizational change to happen. Still short of settling it — self-reported, cross-sectional, one function, one country, and no measure of redesign speed where it does occur. Extended (2026-09-22), with the first national-survey reading: when does ai augment work workflow framework (
- SourceThe ownership variable is asserted as a boolean (your repo vs the company's), but employment contracts, work-for-hire doctrine, trade-secret law, and non-competes already govern externalized judgment. Does any jurisdiction or litigated case actually treat an employee-authored skill file as portable personal property rather than work product — and has any employer yet claimed ownership of one?
- SourceTan's compounding curve (week 4 flywheel, week 12 library-that-answers) is a personal anecdote against telemetry showing skills are copied once and rarely maintained (Agentic Work Systematization). Does any longitudinal measurement of individual skill-file libraries show quality or coverage improving with age, as opposed to accumulating?
- WaitHis first objection bets that better models raise the value of a personal library while Harness Shrinkage as Models Improve predicts scaffolding gets absorbed. These are separable — harness vs library — but no source tests the library half. Does a model generation that absorbs harness complexity also shrink the measured advantage of personal context?
- WaitMusk's deflation prediction is falsifiable and dated: does the price level of manufactured goods and AI-delivered services fall as robot deployment scales, or do input constraints (energy, land, minerals) keep it rising? Trigger: goods-vs-services price divergence through 2030.
- SourceThe end-effector premise is the checkable half of the abundance case — does the physically-gated share of tasks that Task Saturation: Broad but Shallow AI Diffusion measures actually fall as humanoid deployment scales, and at what rate?
- NowIf validation capacity is a commons and the Stockfish threshold is reached unevenly across domains, the commons is destroyed before the threshold arrives in the domains that still need validators. Is there any domain where the ordering has been observed? Partially answered (2026-08-17) by The Data Wall and the Validation Commons Are One Supply Constraint — a qualified negative with one measured instance and a structural correction. No domain in the corpus shows commons-scale depletion, and Lovett says so about his own framework ("a structural prediction … not an observed outcome"), with active counter-evidence (Danish null effects, Ramp's +12.0% entry-level headcount at intensive adopters). The closest measured instance is screening colonoscopy — endoscopists' independent detection accuracy fell after adopting AI-assisted detection, in a domain where the threshold plainly has not arrived — but that is atrophy of the existing stock, not failure of the regeneration mechanism, and it sits in the framework's low-vulnerability set. The structural correction is the important part: the threshold arrives first exactly where a sound cheap verifier exists (Lean, hidden test cases, execution feedback), which is exactly where human validators were never load-bearing — so arrival order and need order are correlated, and the ordering risk this bullet fears is not the general case. It is real one level down, at the sub-task boundary, and formal math already shows the shape: with Lean checking every step the residual human job is checking the formalization, not the proof. Still open: nobody has measured whether that residual capacity is depleting anywhere. The settling instruments are named on the derived page.
- SourceEvery procedural feature in this design was costless: the level said "available" and never said what invoking it costs in delay, effort or uncertainty. Does the flat moderation survive when appeal is priced — e.g. randomizing "re-review within 3 days" against "re-review within 6 weeks," or attaching an explicit success rate to it? An error-correction account should revive as soon as the feature's expected corrective yield is made legible, and the equivalence bound is tight enough (±0.021 at 90%) that a real slope would show.
- SourceThe paper names two rival explanations for the flat moderation it cannot separate — applicants value procedure for standing and dignity, or applicants discount appeal's efficacy exactly as errors mount (and/or read its availability as a signal that true performance beats the displayed rate). These predict identical AMCEs and opposite policy conclusions. Does a design that measures believed appeal-efficacy alongside choice, or that states the appeal success rate as its own randomized attribute, split them?
- NowThe vault now holds a stated premium for human decision authority (+0.272, US Prolific job seekers) and a revealed 78.4% choice of an AI interviewer (Philippine entry-level applicants, humans deciding in both arms). The object-of-choice reconciliation above accounts for the sign, but not the magnitude: is the residual attributable to population and stakes, or would US job seekers facing a real application also take the AI? Answerable by synthesis over the two pages plus the vault's other stated-vs-revealed pairs before any new source arrives.
- WaitThe forward test the report itself names: do the returns to expertise persist, narrow, or invert as models improve? A decrease would mean models are absorbing the judgment users currently supply.
- SourceOutcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the whole success gradient — is real outcome vs. who-narrates-success-in-the-transcript?
- SourceThe study excludes headless / SDK / IDE usage (a "substantial share"). Does the returns-to-expertise pattern hold in non-interactive and pipeline use, where there is no human steering mid-session at all?
- WaitIs "intermediate captures most of the benefit" stable, or an artifact of current model capability — i.e., will the concave curve flatten further (everyone converges) or steepen (mastery starts to separate again) as models get better?
- SourceIs the composition gap between Ramp and this paper caused by the adoption measure (vendor spend against AI job ads) or by the population (US firms against multinational affiliates)? Falsifiable: apply both adoption definitions to the same Revelio firm panel. If spend-defined adopters tilt junior and postings-defined adopters tilt senior on identical firms, the measure is the explanation.
- SourceDoes the recomposition carry into wages, with the senior premium rising and junior pay falling at adopters? The design differences out wages by construction. Falsifiable: an employer–employee panel with pay (Nordic registers, or ADP with firm identifiers) cut by adopter status and seniority.
- WaitDoes the junior-share gap at adopters keep widening past March 2026, or plateau as the paper's "initial reorganization"? Trigger: a later vintage of the same design, which the authors frame as reusable.
- Task Crossover3 open
- SourceCrossover is measured on consumer-surface ChatGPT messages from Business-account users. Does it hold in agentic/API/Codex usage, where work is delegated rather than typed — or does the occupational boundary reassert itself when the unit is a task handed to an agent? Partially answered (2026-09-23) — the population is now observed, the measure is not. how organizations use ai (
empirical, same lab, August 2026) covers exactly the excluded population: ChatGPT Enterprise, 1,764 organizations and 17.4M messages, with task classification on 973 organizations and 8.7M messages. Its task composition by job title class shows the shape crossover predicts — role-specific over-indexing (engineers on debugging, finance staff on financial and tax work) layered on a core of documentation, technical digital work and communication that every role performs, with the paper stating that the role-specific tasks "do not overpower the small set of core tasks that are performed by all roles." It does not answer the question as posed, for a structural reason. The enterprise study uses a proprietary 60-category task taxonomy with no O\NET mapping and no occupational-boundary construct, so no within/cross/generic split and no crossover ratio can be computed from it — and it is still conversational* usage, not the agentic/API/Codex traffic the second half of the bullet asks about (the authors defer Codex to their companion paper). Retagged in effect rather than in tag: the blocking constraint is no longer that the population is unobserved but that the observing instrument uses the wrong taxonomy, which is a re-classification of data OpenAI already holds. - SourceThe dataset records what people attempted and nothing about outcome — the authors say so directly. Is borrowed work done as well as the specialist would have done it, and where does crossover stop being role expansion and start being unreviewed amateur output? A crossover measure joined to a quality or review-coverage measure would settle it, and nothing in the corpus currently does.
- SourceIs the workspace-size gradient about specialist availability (OpenAI's substitution story) or about permission and norms (a large firm's marketer may be allowed to touch less)? The two predict opposite things as small firms grow, and 2.5pp across a descriptive seat-count proxy is thin evidence for either.
- SourceCrossover is measured on consumer-surface ChatGPT messages from Business-account users. Does it hold in agentic/API/Codex usage, where work is delegated rather than typed — or does the occupational boundary reassert itself when the unit is a task handed to an agent? Partially answered (2026-09-23) — the population is now observed, the measure is not. how organizations use ai (
- WaitATLAS is a two-week snapshot with no time dimension, while the AEI reports automation share rising. Does median task saturation move at all over a year, and in which direction? Still no second window (2026-09-22): ATLAS's September 2026 follow-on google atlas ai in science september 2026 is new analysis — a science filter and a new task taxonomy over "approximately 15 million anonymized interactions … sampled in early April 2026" (§2.2.1), i.e. this same corpus — plus an interactive explorer, not a second collection. The program states it is long-term and will publish further iterations; the trigger for this question remains a release whose logs are drawn from a later window.
- SourceThe expertise inversion (2.6× on lowest-expertise non-routine cognitive tasks) is measured on consumer surfaces. Does it hold on enterprise and agentic-coding traffic, where the task mix is deliberately harder? Not yet, and the September 2026 science follow-on restates the exclusion rather than closing it (2026-09-22): google atlas ai in science september 2026 re-cuts the same consumer corpus and repeats that paid/enterprise API is out of frame and that agentic tools such as Antigravity are excluded, "incorporating agentic data … is an important next step" (§6.1). What it does add is a harder-task cut within the consumer surfaces: the ~360K interactions its science classifier isolates score +26% on the same domain-expertise classifier (and +19% tokens, +11% turns) against the average work conversation, so a deliberately sophisticated slice of the same traffic sits above baseline on expertise. That is the population end, not the task end — it does not touch the Q1-versus-Q4 task gradient this question asks about. The survey half has career stage for all 637 respondents and publishes no cut by it.
- WaitAutor & Thompson predict opposite wage effects depending on whether AI absorbs an occupation's expert or inexpert tasks. ATLAS's snapshot points at inexpert. What signal would show the crossover if it happens? Partially answered (2026-09-22): canaries in the coal mine august 2026 (
empirical, ADP payroll microdata through June 2026) names the signal and takes its first reading. The signal is the relative base pay of AI-exposed occupations, split by age and by job-stayer vs. new hire, read against the employment divergence in the same cells. Its first reading is a near-null: a roughly 20-point employment divergence for 22–25-year-olds with "little difference in compensation trends by age or exposure quintile," slightly slower real pay growth for stayers in exposed jobs, and no exposure relationship in young workers' starting pay. The authors' own interpretation is that the two Autor–Thompson effects may be offsetting in the short run, with wage stickiness the alternative — so the null is consistent with no crossover and with a crossover already underway. Two things keep this from closing the question. ADP base salary excludes bonuses, overtime, commissions, equity and tips, "components that are largest in precisely the most exposed, high-income occupations," so the margin where a crossover would surface first is the one the instrument cannot see; and four years is short against wage adjustment. What it does establish is that the crossover has not yet shown up as a wage decline in exposed occupations, which is the direction expert-task absorption would push. Second reading, on the offer margin (2026-10-01): Indeed Hiring Lab (empirical, the platform's own salaried postings through mid-2026) reads the same signal on advertised pay and agrees with ADP once the margins are matched — a modest senior-offer premium in exposed occupations, ≈none at entry, and a pooled +5.7% premium that shrinks to +2.4% (n.s.) with seniority held, because the entry share of exposed postings fell 29% → 10%. Falling entry hiring plus a senior-offer premium is the inexpert-automation signature, i.e. still no crossover; but the post-2022 gap is nearly flat across levels (≈5/6/7 points), so the signal is weak. Reconciliation table on The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay.
- SourceIs the post-ChatGPT senior-offer premium real once seniority is held and the sample is powered for it? Model 3's +2.4% is not significant and the level split's terms mostly are not. Falsifiable: a matched-title, matched-level panel of advertised or realized new-hire pay by exposure, with bonus/equity included, over a window ending 2027 or later.
- SourceHow much of the 29% → 10% entry-share collapse in exposed occupations predates November 2022? The post reports only 2021 and 2026 endpoints; a 2022-baseline split would say whether the tilt is a post-ChatGPT event or a continuation of the post-COVID tech correction that Canaries' pre-trend decomposition also has to handle. Related (2026-10-01), not an answer: on graduates rather than postings, AI-Exposed College Majors at Labor-Market Entry finds the two margins split. Initial employment for the most-exposed majors is flat before 2020 and breaks right after November 2022. Initial earnings were already below their pre-pandemic level relative to less-exposed majors by 2022. So on that instrument the hiring break is post-ChatGPT and part of the pay decline is not.
- SourceSelection vs. treatment: tenure controls attenuate but don't eliminate the enthusiast-selects-into-delegation story. Does a within-person design (sentiment before/after adopting automated workflows) hold the effect?
- SourceSelf-reported "no learning loss" cannot detect real atrophy; is there an objective skill measure that agrees, or does measured skill diverge from felt skill (the AI Brain Fry direction)? Partially answered: Contractor & Reyes's randomized experiment supplies exactly the objective, unaided skill measure this survey lacks — and gives a both answer. It agrees that learning can persist under AI (augmentation users hold +0.29 SD test gains a week later, unaided), so felt-and-measured can align. But it also finds the divergence the question feared: automation users' gains vanish once AI is removed — and "automation share," the survey's own axis, pools both types, so a flat self-reported learning curve can hide a real deskilling half. (Different population — elite undergrads in a proctored lab, not workers — so this sharpens rather than closes the workplace-atrophy question.) Extended (2026-09-05) by Does the Augmentation/Automation Split Govern Skill at Work?: the pooling is structural, not incidental — automation share is Directive + Feedback Loop, and the same collaboration-mode classifier's Learning category is excluded from the axis by construction, so no re-analysis of this survey can separate the arms. Only a decomposition into Directive vs Feedback Loop with Learning re-admitted as a mode would. And the felt/measured gap runs both ways: in the experiment self-assessed knowledge showed no treatment effect (β=0.02) against a measured +0.27 SD, and untreated students overpredicted their own gain fivefold (+25.2pp against an actual 5.1pp) — so self-report is miscalibrated in magnitude and sign depending on whether the respondent has used the tool. Extended again (2026-09-22) by training novices to think or giving them llms rct (
empirical, preregistered 2×2 RCT, n=1,053), which sharpens the question by showing how easily an "objective measure" answers a different one. Its outcome is genuinely objective and human — a 1–5 score from three of twenty trained, condition-blind expert raters — and LLM access lifts it by +0.862 on a control-group estimate of 2.09. But there is no unaided post-measure: the tool is never withdrawn, so the score grades assisted output, not skill, and the authors state outright that they cannot tell whether the model's advantage reflects "knowledge that participants acquired and retained, or output they procured without acquiring anything." The felt side is the sharper omission for this bullet: the study collected the matching self-report — confidence overall, confidence relative to peers, self-assessed knowledge breadth and depth — both before and after treatment, and reports the post-treatment wave nowhere. The felt-versus-measured comparison this question asks for exists in that dataset and is unpublished. Extended (2026-10-01) by designing against deskilling metacognitive feedback (empirical, preregistered, N=704; full treatment on Experimental Learning Impact of Generative AI). This study measures both sides of the comparison: a predicted score before the unaided test, item-level confidence, and the unaided score itself. Felt and measured did not diverge in the feared direction. Every condition under-predicted its unaided score, the AI-only arm included (+0.53 items of 6, against +0.93 for no-AI). The between-condition difference was not significant, and confidence discriminated correct from incorrect answers above chance everywhere, with no condition effect. That cuts against the illusion-of-competence reading this bullet borrows from AI Brain Fry, though only for an assistant that gives answers only on request. The more useful finding for this page is a dissociation. Metacognitive feedback changed behaviour (answer offloading OR 0.47, unaided score OR 1.51) without measurably changing judgment. A survey item about felt learning is therefore the wrong place to look for whether a delegator is regulating their offloading, because the regulation shows up in the request log and not in self-assessment. Still a lab session with an immediate test, not a worker panel, so the workplace clause is untouched. - SourceThe sample is heavily computer/math + management and 88% men; how much of the automation–optimism link survives in a representative population? Partially answered in a population about as far from this one as the vault contains: Jabarian & Henkel surveyed 2,764 Filipino entry-level customer-service applicants (60% female, wages ≈$280–435/month) and found 47% expect AI's workplace impact on themselves to be positive against 19% negative, with the belief predicting delegation choice — 77% of optimists, 72% of balanced, and 65% of pessimists chose an AI voice agent over a human recruiter to interview them. So the belief→delegation association survives a low-wage, non-Western, majority-female, non-technical population. Two things it does not settle. The direction is still unidentified (this is choice given belief, not sentiment given usage), and the same paper shows the association inverts by position: the recruiters, whose own task was the one being automated, split 68% "AI will have a significant personal impact" but only 12% "generally positive" — a quarter of the applicants' rate, inside the same firm. Optimism may track being served by AI rather than delegating to it, and this survey cannot separate those. Extended (2026-09-23) by ai hiring procedure conjoint (Wang, Sturgis & de Kadt, arXiv 2609.16390,
empirical, preregistered conjoint, n=1,919 US job seekers), which measures a served-by population directly and splits the evidence two ways. First, exposure does not breed acceptance: 74.9% of these respondents had already been evaluated by an AI system, they still paid +0.272 in choice probability for a human decision-maker over AI deciding alone, and prior AI-evaluation experience did not condition that premium (interaction -0.013, p = 0.36). So "being served by AI" does not by itself generate the optimism this bullet hypothesizes — though the construct differs (a forced choice between hiring systems, not expectations about one's own labor-market future), so it narrows rather than closes the question. Second, and more usefully, it undermines the behavioral half of the annotation above: the Filipino 77/72/65% figures are choices made by people with a live application at stake, and this study shows that class of measure diverges from acceptance. Intention to apply averaged 3.24 against belief in the legitimacy of AI hiring at 2.85 (paired difference 0.389, d_z = 0.42, t(1918) = 18.56, p < 0.001), with 7.8% (n = 150) combining below-midpoint legitimacy with above-midpoint intention to apply. People apply to employers whose procedures they doubt, because the alternative is not applying — so a delegation choice collected under those conditions measures compliance, not endorsement, and the belief→delegation link this bullet tracks needs a design where refusing carries no cost.
- SourceThe classifier is validated at stage two and unvalidated at stage one: precision and recall are both computed inside the 19,411 dictionary-retrieved candidates, so no number bounds how many AI-work speeches the multilingual dictionaries never surfaced. Does a recall audit — LLM-coding a random sample of the ~1.5M non-retrieved speeches, or a second retrieval pass with an orthogonal (embedding-based) recall stage — change the response mix, and specifically does it raise regulation, whose workplace-governance content is most likely to be phrased without naming a technology?
- SourceParliamentary speech and enacted policy are measured on different objects, and this design measures only the first. Does the enablement–regulation axis predict legislative output — do the countries and periods whose speech tilts toward regulation actually pass more workplace-AI statute than the enablement-tilted ones, net of party composition? Falsifiable against any coded AI-policy output dataset covering the same 33 countries over 2023–2026.
- WaitCompensation's absence is currently explained by the absence of stable AI-loser constituencies, which is a prediction: if measurable displacement arrives, compensation's 2.3% share should rise and the axis should rotate back toward the compensation politics the literature expected. Trigger: a country in this sample with a documented AI-attributable employment decline, re-coded on the same pipeline.
- SourceThe FY2021 complement stocks predict adoption; nothing here shows they predict value. Does a high SG&A or capitalized-software stock at adoption forecast deeper deployment (WAU per employee, task breadth) or better firm outcomes two years later — or does it only forecast who signs the contract? The same Compustat linkage run forward on the 2024 cohorts would settle it.
- WaitLarger adopters show lower per-employee usage entirely through breadth (WAU/emp −0.032\\\*) with per-active-user intensity flat (−0.002, n.s.). Is that a rollout-speed artifact that closes as seat deployment catches up with headcount, or a durable ceiling on how much of a large workforce a centrally administered workspace ever reaches? Falsifiable by re-measuring the same cohorts at week 52 and week 104 against week 26.
- Source"Other / unknown" is the largest single category of weekly active users (≈37.5% of the firm-level job-title panel). Are unclassifiable-title users a random slice of the workforce, or systematically different — contractors, frontline staff, non-English titles — in a way that biases the functional and seniority composition estimates? The classifier validation in Appendix C reports top titles per class but never characterizes the residual.
- SourceThe \$15–149B range rests on an assumed 0.5–5% time saving because no causal estimate exists for AI in the household. What experiment would measure actual household time savings, and does the effect survive contact with one?
- SourceATLAS argues gains skew to women (30% more productive household time) but could reverse given the AI adoption gender gap. Which effect dominates in current data?
- SourceIf AI substitutes household self-service for purchased professional services (tax prep, legal advice, therapy), measured GDP falls while welfare rises. Is that substitution detectable yet in the market-services data Coyle cites?
- SourceRoughly half the pooled break is venue composition (+1.72 → +0.75 pp/yr inside continuously observed venues), and OpenAlex changed its author-disambiguation pipeline in July 2023 — inside the window. Does the break survive in a corpus with stable curation and stable disambiguation (Scopus, Web of Science, or a single large publisher's internal records) over the same period?
- SourceSolo papers narrow 23% in content breadth while showing no movement toward new territory. Is that scope discipline (the author does what they can verify alone) or capacity limit (the LLM covers the execution but not the range a second mind supplied) — and does the quality of solo output diverge from coauthored output on citations, replication or retraction? The paper measures quantity and content and explicitly not quality.
- WaitThe break's attribution rests on a cross-field ordering the author concedes is an ordering, not a test. If the halt is LLM-driven it should track LLM capability, so the ordering should shift as models improve — fields whose execution work is newly automatable (lab protocol design, instrument control) should join late. If it is a one-time re-sorting, the ordering freezes and the halt decays.
- SourceMechanism 2 has never been directly measured in a workplace. Does a cohort that entered an AI-heavy profession after 2023 show lower unaided task accuracy than a matched earlier cohort at the same tenure? The paper specifies the design (no-AI assessment stratified by cohort and AI exposure); nobody has run it. Related (2026-09-05): Does the Augmentation/Automation Split Govern Skill at Work? searched the vault for a substitute and confirms there is none — the four workplace results with the removal signature (Budzyń, Dell'Acqua, Wiles, Vicente & Matute) are all borrowed through this page, and the vault's own workplace instruments measure delegation, throughput, error rates and retention but never unaided skill. It also names a second design that would bear on the same gap: classify use mode within-subject from usage logs, then measure performance after AI is withdrawn. Related (2026-09-22), not an answer: psychological costs ai adoption software engineering (
case-study) is the corpus's first workplace source in which practitioners state the mechanism themselves — anticipated skill atrophy, and an explicit worry about eventually being unable to critically review AI output — and the first to inventory the counter-practices they adopt against it. It measures no unaided accuracy of any kind, its cohort is 19-of-21 veterans whose capability predates the exposure, and self-reported fear of deskilling is exactly the self-report this question was written to route around. It shifts the mechanism from inferred to voiced, which is worth one line and nothing more. - SourceThe cohort evidence is a snapshot ending Sept 2025 in the most AI-exposed occupations. Does the 22–25 employment decline persist, reverse, or re-sort as agentic tooling matures — and if entry-level postings recover, does that restore the developmental content of the work or just its headcount? Recovery of positions and recovery of the regeneration mechanism are not the same event, and only the first is currently instrumented. Partially answered (2026-08-04): Indeed Hiring Lab instruments the first half — postings in the most-exposed occupation did rebound (US software development +15% since February 2025 against overall postings −7%) — and the composition answers the sub-question in this page's favour: 71% of the May 2025 – May 2026 increase is senior roles, 37% AI-titled, with the author himself conceding the market "could still be experiencing a seniority-biased technological change." So the recovery is real and is not entry-level, on this instrument. It leaves the harder half untouched: nothing there measures the developmental content of any role, the data are one job board's vacancy flow analyzed by that job board, and Ramp's firm panel finds entry-level headcount growing fastest on a different unit. Extended (2026-09-22) — the first clause is now settled and the second is not, which is why this stays open. canaries in the coal mine august 2026 (
empirical, ADP payroll microdata through June 2026) answers "persist, reverse, or re-sort" without ambiguity: it persists and widens, from a 15% kept-pace shortfall at the July 2025 vintage to 19% at June 2026, with the divergence continuing for more than three years and through a period in which interest rates stabilized and fell. It also re-sorts in a specific way rather than diffusely — the loss concentrates where AI usage is automative (−0.098 per SD at 22–25, shrinking monotonically to −0.006 at 50+) and is absent where it is augmentative, and it runs through hiring rather than separations, which is this page's Mechanism 1 stated as a measurement. The snapshot-ending-Sept-2025 premise in the question above is therefore obsolete; treat the first clause as answered. The second clause is untouched and unreachable from this instrument: payroll records carry headcount, hiring rates, separations and base pay, and carry nothing about whether the roles that remain still teach. Until something measures developmental content, the question this page actually cares about has no evidence on either side. Extended (2026-10-01): Indeed corroborates persistence through 2026 on postings — entry share in the most-exposed occupations 29% → 10% since 2021; see The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay. Extended (2026-10-01), to firms and outside the US: Chandar & Klein Teeselink (empirical, Revelio, instrumented) finds the junior share at AI-adopting affiliates −1.9pp against matched non-adopters by March 2026. The share is negative in 23 of 31 countries and significant in seven (US, UK, Spain, Poland, Brazil, Mexico, Saudi Arabia). It comes mainly from senior growth (+6.7%) rather than junior cuts (−2.5%, n.s.). So on this instrument the regeneration slot is shrinking as a share at adopters, not necessarily as a count. The developmental-content clause is still untouched; see Seniority-Biased AI Adoption: The Junior Share at Adopting Firms. Extended (2026-10-01), to graduates by major: Census PSEO×LEHD records show the most-exposed decile of majors still −2.0pp employed and −3.3 log points in earnings two years after graduating, while deciles 7–9 recover. The move into retail and food service is the nearest proxy yet for first-job content, but it is a sector code, not a measure of learning; see AI-Exposed College Majors at Labor-Market Entry. Counter-evidence (2026-10-01), reconciled rather than superseding: Fairlie & Wu (CPS,empirical) find no significant summer-2026 unemployment break for all 22–25 bachelor's holders, including a 'sidelined' measure. It measures a different outcome (unemployment, not occupation headcount or earnings) in a wider population, so the persistence clause stands; see AI-Exposed College Majors at Labor-Market Entry. Related (2026-09-22), and deliberately not counted as partial: when does ai augment work workflow framework draws this bullet's distinction independently — entry-level PMET openings in Singapore rose 32,500 (Dec 2025) → 32,800 (Mar 2026) with fresh-graduate outcomes "broadly resilient," which its authors read as snapshots that "do not yet indicate broad-based deterioration in entry-level opportunities, but neither do they establish the longer-term effects of AI on human development." Positions and regeneration are separated there for the same reason they are separated here, and the instrument proposed to close the gap is the one this bullet says is missing: repeated longitudinal measurement of learning, progression and agency inside the workflow, not headcount at the door. It answers nothing — the figures are third-party MOM statistics reaching the vault secondhand in apractitioner-opinionframework paper, a single small economy, one quarter apart, and no developmental content is measured in either jurisdiction. - SourceThe framework predicts differential depletion by its five factors. Do software engineering, financial analysis and legal research actually diverge from medicine and engineering on validation-capability measures — or does regulatory intensity turn out to be weaker protection than the model assumes? Related (2026-09-22), and deliberately not counted as partial: canaries in the coal mine august 2026 publishes occupation-level detail at last — the most-exposed quintile's largest ADP occupations are chief executives, customer service representatives, accountants, HR specialists, wholesale sales reps, systems software developers and computer/IS managers, with bookkeeping and billing clerks further down (Online Appendix Table A.6, ranks 1–25 verified; the table's tail is parse-damaged and is not cited). That overlaps two of the framework's three predicted high-vulnerability professions and misses the third entirely — legal research does not appear, while clerical and customer-facing work the framework does not theorize about dominates the list. More to the point, the paper never cuts employment change by profession, licensure or regulatory intensity, so it supplies an exposure ranking and no test of the five factors. What it does supply is a warning about the question's framing: the five-factor model predicts depletion in high-skill, low-regulation professions, and the measured decline is concentrated in occupations that are exposed rather than professional. Partially answered (2026-10-01), on the input side only: Chandar & Klein Teeselink cut the within-occupation junior share at AI adopters by profession. The significant declines are legal −4.9pp, computer & mathematical −3.3, business & financial operations −3.2, plus sales and management. Healthcare practitioners and architecture & engineering have intervals spanning zero (chart reading). That is the framework's predicted high/low split, profession by profession, on entry-level headcount. It is not validation capability, legal does not survive Bonferroni, and the country-level labor-rigidity correlation (+0.20, n.s.) measures general labor law, not licensure.
- SourceMechanism 2 has never been directly measured in a workplace. Does a cohort that entered an AI-heavy profession after 2023 show lower unaided task accuracy than a matched earlier cohort at the same tenure? The paper specifies the design (no-AI assessment stratified by cohort and AI exposure); nobody has run it. Related (2026-09-05): Does the Augmentation/Automation Split Govern Skill at Work? searched the vault for a substitute and confirms there is none — the four workplace results with the removal signature (Budzyń, Dell'Acqua, Wiles, Vicente & Matute) are all borrowed through this page, and the vault's own workplace instruments measure delegation, throughput, error rates and retention but never unaided skill. It also names a second design that would bear on the same gap: classify use mode within-subject from usage logs, then measure performance after AI is withdrawn. Related (2026-09-22), not an answer: psychological costs ai adoption software engineering (