H
Howardism
Plate IIAI Economics & Labor中文HOWARDISM

Task Saturation: Broad but Shallow AI Diffusion

Google ATLAS's marquee work finding — AI reaches 68% of detailed occupations (88.4% of US employment) but only 21% of the tasks in the median occupation, with end-to-end automation the intent of just 6.5% of non-routine-cognitive conversations vs 26.9% for routine-cognitive; the extensive margin is gated by physicality, the intensive margin concentrates in non-routine cognitive work, and usage over-indexes most on the *lowest*-expertise cognitive tasks

Article metadata
Publication details
Published:July 25, 2026
Filed:Concept
Domain:AI Economics & Labor
Reading:23 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Task Saturation: Broad but Shallow AI Diffusion

Sources#

Summary#

The headline result of Google's ATLAS v1.0: AI has reached almost every kind of job and almost none of the work inside them. Gemini usage appears in 68% of detailed O*NET occupations, which together employ 88.4% of US civilian workers — but within the occupations where it appears, workers use it for a median of 21% of that occupation's tasks. Breadth is nearly universal; depth is a fifth.

"Task saturation" is ATLAS's operational primitive: the share of an occupation's constituent O*NET task statements where usage clears a minimum-users threshold (25 unique users globally for a task; 50 for an occupation). It is a presence measure, not a volume measure — deliberately, because presence is far more robust to classifier noise than exact frequency (Usage-Telemetry Classifier Validation).

Evidence note. empirical — 14.65M de-identified Gemini interactions, April 6–19 2026, mapped to O*NET v30.2. First-party, consumer-and-free-API surfaces only, observational. The task-level numbers rest on a classifier whose exact O*NET task accuracy is 22.6%; ATLAS mitigates by measuring presence rather than counts and by aggregating into Autor–Thompson task types (70.4% agreement). See Google AI & Economy ATLAS for full limitations.

The distribution#

CutResult
Detailed occupations with any observed usage68% (≈88.4% of US employment)
O*NET tasks with observed usage, overall~20%
Occupations with zero task saturation29%
Occupations with ≥25% of tasks saturated30% (44–50% of US employment)
Occupations with ≥50% of tasks saturated11% (26–31% of US employment)
Occupations with ≥75% of tasks saturated3% (9–10% of US employment)
Median saturation, conditional on any usage21%

The most-saturated occupations are software QA analysts and testers, HR specialists, document management specialists, market research analysts, and network/systems administrators — roles whose task lists are unusually text-shaped. The least saturated, among those with any usage at all, are teachers (special education, kindergarten, postsecondary English) and midwives. The largest occupations with no observed usage are food preparation workers, fast-food cooks, dining attendants, and short-term substitute teachers.

The extensive margin is gated by physicality#

What separates the 29% of occupations with zero saturation from the rest is not skill, wage, or prestige — it is how much of the job is physical. Occupations with zero saturation have markedly higher shares of manual tasks; cognitive and interpersonal tasks, especially non-routine cognitive, dominate the occupations where AI shows up at all.

But the boundary is porous, and this is ATLAS's most underrated finding: manual trades do use AI, for the cognitive tasks embedded in physical work. Industrial machinery mechanics (44% manual tasks) generate thousands of conversations about analyzing test results and machine error messages. Automotive service technicians (83% manual) generate over ten thousand conversations about testing vehicle components, rewiring systems, and inspecting parts for wear — with more than 2× the multimodal conversation share of the work baseline. The blue-collar exclusion narrative is wrong in the specific way that matters: AI enters physical work through its diagnostic and interpretive layer, and it enters through the camera.

The intensive margin concentrates in non-routine cognitive work#

Applying the Autor–Thompson (2025) five-way task classification — routine/non-routine × cognitive/manual, plus non-routine interpersonal:

Task typeShare of O*NET universeShare of Gemini interaction volume
Non-routine cognitive analytic35%65%
All cognitive (routine + non-routine)~50%86%
Manual28%~5%
Interpersonal22%~9%

This is the discontinuity with prior technology waves. Autor et al. (2003) and Acemoglu & Autor (2011) modeled digitalization as routine-biased: computers substituted for codifiable work and complemented non-routine problem-solving, producing job polarization. LLMs invert the target — the tasks they absorb most are precisely the non-codifiable ones that rules-based computing could never reach. As ATLAS puts it, "AI-related impacts will not be cabined to routine tasks."

Intent: automation is rare, and confined to routine work#

ATLAS classifies each conversation cluster's intent into five categories. The pattern is the report's central rebuttal to the mass-displacement narrative:

IntentNon-routine cognitiveRoutine cognitiveInterpersonalManual
Task Automation (end-to-end)6.5%26.9%2.9%4.2%
Partial Drafting & Generation41.5%43.8%37.3%7.0%
Review & Refinement3.8%4.5%~1%~0.6%
Ideation & Strategy18.6%2.6%31.1%6.0%
Information Retrieval & Learning29.6%22.2%27.6%82.1%

Three things fall out. Automation intent is 4× higher for routine cognitive work than non-routine — the old substitution logic still holds where tasks are codifiable, it just no longer describes the bulk of usage. Manual-task usage is overwhelmingly learning: 82.1% information retrieval, which is the mechanic reading a diagnostic, not a machine replacing them. And interpersonal work splits between drafting and ideation with almost no automation — the AI writes the difficult email, it does not have the conversation.

ATLAS is careful that this is intent, not outcome: it cannot see the work happening outside Gemini, so it cannot say what fraction of the whole job AI completed. And it notes that even genuine task automation "does not necessarily equate [to] job automation, as coordination costs, complementary tasks and organizational frictions are highly prevalent" — the same argument Organizational Complements to AI makes from the adoption side.

The expertise inversion#

The finding that sits least comfortably with the rest of the wiki. ATLAS replicates Autor & Thompson's expertise measure (100 minus the average Standard Frequency Index of a task statement's lemmatized words — expert vocabulary is rare but low-entropy), sorts ~19,000 O*NET tasks into expertise quartiles, and computes how over-represented Gemini usage is relative to the task universe in each:

Task typeQ1 (lowest expertise)Q2Q3Q4 (highest)
Non-routine cognitive2.60×1.64×1.72×1.78×
Routine cognitive0.89×1.65×1.60×1.18×
Interpersonal0.68×0.42×0.26×0.38×
Manual0.14×0.23×0.17×0.23×

Usage is most over-represented on the lowest-expertise non-routine cognitive tasks — 2.6× baseline, well clear of the 1.6–1.8× flat band across the other three quartiles. Yet the people doing the using skew rich and educated: a 1% increase in an occupation's median earnings is associated with >2.5% higher usage intensity (2.68 univariate; 1.86 controlling for education, R² 0.337), and weighting US median earnings by Gemini conversations moves it from $62,252 to $82,919 — and to $86,157 when weighted by tokens.

So the composition is: high-expertise workers, using AI disproportionately on their low-expertise tasks. That is the augmentation reading in its strongest form, and it is compatible with Returns to Expertise in Agentic Coding rather than opposed to it — the expert brings the judgment, and offloads the parts that don't need it. Autor & Thompson's model says which way this cuts: automating an occupation's inexpert supporting tasks raises the scarcity of the remaining human expertise, lifting wages while lowering employment; automating its expert tasks erodes barriers to entry and depresses wages. ATLAS's data currently points at the first.

The labor-market half, read for the first time (2026-09-22). ATLAS's prediction is testable only against realized employment and wages, which ATLAS does not observe. Brynjolfsson, Chandar & Chen (revised August 2026) observe both, on ADP administrative payroll through June 2026, and reach for the same Autor & Thompson model to interpret the result. Their Fact 6: adjustment is running through employment, not base compensation. Employment of 22–25-year-olds in the most-exposed quintiles fell ~11% against ~10% growth in the least exposed — roughly a 20-point divergence — while compensation shows "little difference in compensation trends by age or exposure quintile." Among job-stayers, real base pay grew slightly more slowly in more-exposed jobs (not specific to the young); among new hires, starting pay shows no relationship with exposure for young workers and rose modestly faster for older hires in exposed jobs.

The authors read that null exactly as this page's open question anticipates: "the limited changes we find for wages suggest that these effects may be offsetting, at least in the short run" — inexpert-task automation pushing wages up, expert-task automation pushing them down, cancelling — with wage stickiness the alternative. So the crossover signal is the relative base pay of exposed occupations, split by age and by stayer-vs-new-hire, and its first reading is ambiguous rather than confirmatory. One limitation makes the instrument weakest precisely where the answer lives: ADP base salary "excludes bonuses, overtime pay, commissions, equity, and tips — components that are largest in precisely the most exposed, high-income occupations." The compensation margin where a crossover would first appear is the margin this data cannot see.

A qualification from outside the occupation panel (2026-10-01). Fact 6 holds for workers in an occupation. Census PSEO×LEHD records (Orr, Tucker & Warren) assign exposure by college major, and there graduates of the most-exposed majors lose 13% of initial earnings, about half from sorting into lower-paying sectors. So pay does adjust at entry. It adjusts through which job graduates land in, which an occupation-level wage series cannot record. See AI-Exposed College Majors at Labor-Market Entry.

Why this is contested rather than settled#

ATLAS explicitly stages the two readings of its own wage gradient rather than picking one:

  • "Professionals are automating themselves out of existence" — high-wage white-collar workers are the heaviest users, and they are pointing AI at the cognitive core of their jobs.
  • "Augmentation deepening the premium" — those workers are automating routine cognitive tasks and collaborating on non-routine ones, which raises returns to the non-routine human skills, widening the gap against everyone who can't use AI well.

Underneath sits the micro–macro gap (Imas & Shukla 2026): controlled experiments consistently find AI compresses the expert premium (novices catch up), while real-world observational studies find it widens. The reconciling mechanism ATLAS names is the endogenous adoption margin — adoption isn't randomly assigned, so the catch-up scenario requires broad uniform access while the run-away scenario follows from concentrated adoption via task selection (low-skill workers apply AI where it doesn't help), complementary judgment (verification requires human capital), and seniority-biased demand (firms substitute away from entry-level hiring while senior staff amplify).

Where the number is fragile#

Two caveats to attach whenever the 21% is quoted:

  1. It is a task-level statistic from a classifier that gets exact task assignment right 22.6% of the time. Google names this gulf itself and calls it "a caution against over-relying on hyper-specific task analysis." The mitigations are real — presence rather than frequency, aggregation into Autor–Thompson types where agreement is 70.4%, human raters approving 85.8% of task labels as economically reasonable — but they mitigate, they don't eliminate. Full treatment at Usage-Telemetry Classifier Validation.
  2. It disagrees with the AEI. ATLAS observes ~20% of tasks against Anthropic's 36% (Handa et al. 2025) and 49% combined (Appel et al. 2026), and attributes the gap to its stricter privacy thresholds. On automation the gap is wider still and definitional: ATLAS's <10% for non-routine cognitive vs Anthropic's 43–45% overall — and ATLAS notes Anthropic's automation share has been rising over time while its own snapshot has no time dimension at all.

Connections#

  • Seniority-Biased AI Adoption: The Junior Share at Adopting Firms — the formal version of the micro–macro gap's endogenous-adoption point: an adopter-vs-non-adopter DiD and the general-equilibrium employment change can have opposite signs (all four combinations occur in simulation)
  • AI-Exposed College Majors at Labor-Market Entry — Fact 6's "employment, not pay" restated for graduates: measured by major rather than occupation, entry pay falls 13% for the most-exposed decile, mostly through sorting across sectors
  • The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay — the Autor–Thompson crossover signal's second reading, on advertised offers rather than payroll: senior-offer premium, no entry premium, mostly composition — consistent with inexpert-task automation, and with Canaries Fact 6
  • The Enterprise AI Adoption Gradient — the broad-but-shallow shape measured inside paying enterprises rather than across occupations: documentation and technical writing reaches 56.3% of weekly active users but is 18.3% of messages, topic overviews reach 38.9% and account for 4.4%, and the pooled long tail of 48 smaller task categories is the single largest block of message volume at 27.3%
  • GDPval Benchmark — the complementary cut of the same labor market, selected by a different instrument and answering a different question. ATLAS samples observed usage and asks what share of an occupation's tasks AI touches at all; GDPval selects occupations top-down by wage mass (the five highest-paying predominantly-digital occupations in each of the nine sectors contributing over 5% of U.S. GDP, 44 in total, $3T of annual earnings) and asks whether the delivered artifact beats the incumbent professional's. The pairing is the useful part: this page's answer is broad-but-shallow (68% of occupations, 21% of the median occupation's tasks), and GDPval's is that on the deep end of the highest-wage slice, win-or-tie rates against the professional run 12.4% to 47.6% with a steep duration gradient. Neither instrument sees the other's margin — usage telemetry cannot tell you whether the output was any good, and a win rate cannot tell you whether anyone is actually using the model for that task
  • Task Crossover — the moving denominator under this page: saturation measures what share of an occupation's tasks AI reaches, while crossover finds that which tasks belong to the occupation is itself shifting (43.5% of occupation-specific use is outside the user's own job)
  • Google AI & Economy ATLAS — the program and dataset this finding comes from
  • AI Adoption in Scientific Work — the same corpus and the same two-week window, re-filtered to the ~360K science interactions and mapped to a scientific task taxonomy instead of O*NET. It is the heavy end of this page's distribution (SOC 19 at 2.7× employment share, 5.8× for core STEM detailed occupations) and the place where the "broad but shallow" picture gets a why: time savings that do not reach output because the constraint moves downstream
  • Anthropic Economic Index — the rival measurement; observes 36–49% task coverage and 43–45% automation against ATLAS's ~20% and <10%, on a different product with a different classifier
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — task saturation is a stricter observed exposure measure from a second lab; the four-way exposure distinction is what keeps "68% of occupations" from being read as "68% of jobs are at risk"
  • Returns to Expertise in Agentic Coding — the composition here (expert workers using AI on low-expertise tasks) is the mechanism behind the expertise premium, not a counterexample to it: the human supplies the judgment and offloads what doesn't need it
  • Organizational Complements to AI — ATLAS's own caveat that task automation ≠ job automation because coordination costs and organizational frictions bind, stated from the usage side
  • Codified vs Tacit Knowledge Exposure — the labor-market outcome this page's expertise inversion predicts but cannot observe. Expert/inexpert (word rarity) and codified/tacit (how the knowledge was acquired) are different axes — tax law is expert and fully codified — but they agree on who gets hit first, and the payroll series behind that page is the first reading of Autor & Thompson's wage prediction on realized data
  • Usage-Telemetry Classifier Validation — the measurement floor under every number on this page
  • The Household Production Boundary — the other 86.5% of usage; task saturation describes only the 13.5% of conversational AI that is work
  • Conversation-to-Delegation Shift — the delegation reading from OpenAI's Codex data; ATLAS's <10% automation intent is measured on consumer surfaces that exclude exactly the agentic coding traffic where delegation concentrates
  • Jagged Intelligence (Ghosts, Not Animals) — the task-level rather than job-level shape of AI capability is what makes saturation partial by construction
  • Role Averaging, Not Role Elimination — occupations losing a fifth of their tasks to AI collaboration, not their existence
  • Market-Priced AI Exposure (the AI Premium) — the market's skill map penalizes analytical/scientific work and rewards interactive/relational, which is the same manual-and-interpersonal-under-represented pattern priced from the equity side
  • Firm AI-Spend Intensity and Headcount Growth — the firm-level counterpart: intensity-gated headcount growth under AI adoption, which is what shallow-and-collaborative diffusion should produce
  • Printing Press Software Democratization — the QA-analyst and document-specialist saturation ceiling is what happens when a text-shaped job meets a text-shaped tool
  • Post-Scarcity Macroeconomics — the physicality gate measured here is the exact variable Musk's abundance forecast turns on: his claim is that humanoid "end effectors" open it, and ATLAS's extensive-margin measurement is how you would ever check
  • Cognitive Capability Profiling for Task Suitability — a cognitive account of the broad-but-shallow shape. Prunty et al. find workplace activities converge on a shared cognitive core (Planning, Semantic Memory, Working Memory, Language, Procedural Memory) that dominates nearly every one of 18 activities across six occupational domains, while six profiled AI systems are uniformly strong on knowledge, language and social cognition and uniformly weak on Action Planning (1.99), Instrumental Reasoning (1.22) and Object Permanence (0.29). That pairing predicts this page's shape directly: broad reach wherever an activity's demands fall on the core, shallow penetration wherever its secondary demands land on the three floors. It is a mechanism, not a corroboration — their suitability scores are comparative rather than calibrated and are validated against no usage data, so the prediction has not been checked against ATLAS's 21% median

Open Questions#

  • ATLAS is a two-week snapshot with no time dimension, while the AEI reports automation share rising. Does median task saturation move at all over a year, and in which direction? Still no second window (2026-09-22): ATLAS's September 2026 follow-on Google AI & Economy ATLAS: AI in Science (September 2026) is new analysis — a science filter and a new task taxonomy over "approximately 15 million anonymized interactions … sampled in early April 2026" (§2.2.1), i.e. this same corpus — plus an interactive explorer, not a second collection. The program states it is long-term and will publish further iterations; the trigger for this question remains a release whose logs are drawn from a later window.
  • The expertise inversion (2.6× on lowest-expertise non-routine cognitive tasks) is measured on consumer surfaces. Does it hold on enterprise and agentic-coding traffic, where the task mix is deliberately harder? Not yet, and the September 2026 science follow-on restates the exclusion rather than closing it (2026-09-22): Google AI & Economy ATLAS: AI in Science (September 2026) re-cuts the same consumer corpus and repeats that paid/enterprise API is out of frame and that agentic tools such as Antigravity are excluded, "incorporating agentic data … is an important next step" (§6.1). What it does add is a harder-task cut within the consumer surfaces: the ~360K interactions its science classifier isolates score +26% on the same domain-expertise classifier (and +19% tokens, +11% turns) against the average work conversation, so a deliberately sophisticated slice of the same traffic sits above baseline on expertise. That is the population end, not the task end — it does not touch the Q1-versus-Q4 task gradient this question asks about. The survey half has career stage for all 637 respondents and publishes no cut by it.
  • Autor & Thompson predict opposite wage effects depending on whether AI absorbs an occupation's expert or inexpert tasks. ATLAS's snapshot points at inexpert. What signal would show the crossover if it happens? Partially answered (2026-09-22): Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence (empirical, ADP payroll microdata through June 2026) names the signal and takes its first reading. The signal is the relative base pay of AI-exposed occupations, split by age and by job-stayer vs. new hire, read against the employment divergence in the same cells. Its first reading is a near-null: a roughly 20-point employment divergence for 22–25-year-olds with "little difference in compensation trends by age or exposure quintile," slightly slower real pay growth for stayers in exposed jobs, and no exposure relationship in young workers' starting pay. The authors' own interpretation is that the two Autor–Thompson effects may be offsetting in the short run, with wage stickiness the alternative — so the null is consistent with no crossover and with a crossover already underway. Two things keep this from closing the question. ADP base salary excludes bonuses, overtime, commissions, equity and tips, "components that are largest in precisely the most exposed, high-income occupations," so the margin where a crossover would surface first is the one the instrument cannot see; and four years is short against wage adjustment. What it does establish is that the crossover has not yet shown up as a wage decline in exposed occupations, which is the direction expert-task absorption would push. Second reading, on the offer margin (2026-10-01): Indeed Hiring Lab (empirical, the platform's own salaried postings through mid-2026) reads the same signal on advertised pay and agrees with ADP once the margins are matched — a modest senior-offer premium in exposed occupations, ≈none at entry, and a pooled +5.7% premium that shrinks to +2.4% (n.s.) with seniority held, because the entry share of exposed postings fell 29% → 10%. Falling entry hiring plus a senior-offer premium is the inexpert-automation signature, i.e. still no crossover; but the post-2022 gap is nearly flat across levels (≈5/6/7 points), so the signal is weak. Reconciliation table on The AI-Exposure Pay Premium: Advertised Offers vs Realized Pay.

Sources#

§ end
Cited by 25
Related articles