H
Howardism
Plate IIEntities中文HOWARDISM

Google AI & Economy ATLAS

Google's recurring economic-research program measuring Gemini usage across the economy — ATLAS v1.0 (July 2026) maps 14.65M de-identified interactions from Gemini App, AI Mode, and the Gemini API onto BLS/O*NET occupations and ATUS household activities across 150 countries and 143 languages; the direct methodological rival to the Anthropic Economic Index, and the first such program to publish its classifier-validation numbers; a September 2026 update adds an open-access interactive explorer and a science-focused follow-on report, both re-analyses of the same April 2026 corpus rather than a second collection

Article metadata
Publication details
Published:July 25, 2026
Filed:Entity
Domain:Entities
Reading:13 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Google AI & Economy ATLAS

Sources#

What it is#

ATLAS — Activity, Task, Landscape, and Adoption Study — is Google's ongoing economic-research program measuring how AI is actually used, read off Gemini usage logs. ATLAS v1.0 (July 23, 2026) is the first release: 14,653,926 de-identified interactions sampled from the Gemini App, Google AI Mode, and the Gemini API between April 6 and April 19, 2026, clustered and mapped onto BLS/O*NET occupations and tasks for work usage and the American Time Use Survey lexicon for non-work usage. Coverage: 800+ occupations, ~4,000 work tasks, 300 household activities, 150 countries, 143 languages.

It is the structural counterpart to the Anthropic Economic Index: same instrument class (privacy-preserving classification of a lab's own conversation logs), same question (where is AI landing in the economy), different product, different user base, and — as it turns out — different numbers.

Authors are drawn from Google and Google DeepMind; Zanna Iscenko and Scott Strand are the corresponding authors, with James Manyika and Fabien Curto Millet among the senior names. Diane Coyle (Cambridge) and David Autor (MIT) are credited for guidance and review, and Coyle contributes a signed guest comment on the household production boundary — an unusual move that puts named external economists inside a first-party vendor report.

Evidence note. empirical — large-scale measured usage with a documented pipeline, published classifier-validation results, and stated limitations. But it is first-party: Google measuring Google's own surfaces, over a two-week window, with Gemini models doing the summarizing, clustering, and classifying. Not independently reproducible; the raw logs cannot be released. The evidence tier reflects the methodology, not independence.

The pipeline#

  1. Work / non-work split. An automated classifier routes each interaction into one of two pipelines. This binary gates everything downstream (93.7% accuracy on a balanced synthetic set).
  2. Summarize, then cluster. Conversations are summarized individually (full text discarded), then grouped by OCTO (Observation Clustering and Taxonomy Organisation), a bespoke DeepMind clustering and hierarchical-taxonomy tool, and re-summarized at the cluster level.
  3. Map to official taxonomies. Work clusters → BLS 2018 SOC + O*NET v30.2 occupations and task statements. Non-work clusters → BLS 2024 ATUS Activity Lexicon (three tiers).
  4. Bespoke overlays. Intent (five categories), task expertise (Autor–Thompson Standard Frequency Index replication), Autor–Thompson routine/non-routine × cognitive/manual/interpersonal task types, and multimodality. Gemini 3.1 Flash-Lite does the classification.
  5. Validate. Three ways — synthetic ground truth, inter-rater agreement, human approval. See Usage-Telemetry Classifier Validation.

Privacy governance is four-layered: DLP filters strip PII before processing; internal log identifiers are replaced with mathematically unlinked UUIDs; two rounds of summarization discard the underlying text; and k-anonymization drops any cluster representing fewer than 10 unique users. Neither raw conversations nor individual summaries are retained in the final dataset.

How Google says it differs from the AEI and OpenAI#

The report states its own methodological deltas against Anthropic (Handa et al. 2025; Massenkoff et al. 2026) and OpenAI (Chatterji et al. 2025) — worth recording because they are the reasons the numbers diverge:

  • Wider surface pool — a standalone chat app, an AI search experience, and a developer API pooled together, rather than one product family.
  • Scale — 15M interactions, with clustering rebuilt for that volume.
  • Recursive nested taxonomies — SOC and O*NET traversed in a single pipeline rather than classified separately.
  • LLM-assisted category annotation — taxonomy category descriptions are expanded by a model to give the classifier more context.
  • Randomized classifier options — option order is shuffled to defeat the documented position bias in LLM classification (the same bias family LLM-Judge Validation measures).
  • Synthetic-data validation — accuracy measured against a generated ground truth, since real logs have none.
  • ATUS mapping for non-work — the substantive expansion; prior work treated non-work usage thinly.
  • Penetration adjustment for cross-country comparison — raw usage divided by a StatCounter Gemini-share proxy, so the map reads as AI diffusion rather than Google's regional footprint.

Global diffusion (§5)#

ATLAS's third analytical block, and the part most specific to having a globally-deployed consumer surface:

  • Adoption scales with wealth. Penetration-adjusted conversations per capita against log GDP per capita gives a slope of ~0.9 — a 1% increase in GDP per capita is associated with a 0.9% increase in usage. (Anthropic's international index found 0.7 for Claude.) The relationship survives restricting to free services, so it is not simply ability to pay.
  • The concentration is severe. The lowest-usage quintile of countries holds 17% of world population and generates 2% of conversations; the top quintile holds 11% of population and drives 30% — 2.8× its proportional share.
  • Internet access is a real but partial bottleneck. Normalizing by internet users rather than population nearly triples Sub-Saharan Africa's adoption metric, but does not close the gap.
  • The interest–adoption gap. Relative Google Trends search interest in AI is highest in South and Southeast Asia and East Africa — exactly the regions in the lowest usage quintiles — and lowest in Western Europe and Japan, which lead per-capita usage. Google's reading: in mature markets AI has already stopped being a thing you search for and become background utility.
  • The work-share inversion. By absolute per-capita work conversations, high-income countries lead. By share of a country's conversations that are work-related, the ranking flips: Africa surges into the top quintile, the US and EU fall to the bottom. Google offers three candidate explanations and endorses none — goal-directed usage under metered data costs, leisure dilution in rich countries, or an artifact of excluding enterprise Gemini subscriptions that are more common in the US and EU.
  • API usage is far more concentrated than conversational usage, in established tech hubs. The stated reason is that APIs need engineers, cloud infrastructure, and capital, and are billed per token — with tokenization bias making non-Latin-script languages structurally more expensive per unit of text.
  • Language. 143 languages clear the privacy threshold; English is just over a third of conversations, Spanish 12%, Arabic ~7%, Portuguese ~6%. Non-primary-language use is 26% for work vs 24% for non-work — near-parity, which disproves the hypothesis that users code-switch into English for high-stakes tasks. It is highest in volunteer (21.9%), religious (20.5%), and civic (18.5%) activity, i.e. driven by the sociolinguistic character of the activity, not its economic value. But it costs: non-primary English conversations run 9–12% more turns and 18–20% more tokens after fixed effects.
  • Multimodality skews the other way. Non-OECD work conversations are roughly twice as likely to include a generated image or video; media generation is most common in Africa, least in Europe.

The September 2026 update (and what it is not)#

On 2026-09-15 Zanna Iscenko and Scott Strand announced the program's first update since v1.0 (blog.google). Three parts, and the distinction between them matters for anything citing ATLAS:

  1. A new interactive, open-access data experience — the ATLAS data points published as a browsable explorer (usage rates by occupation, household usage, country adoption) rather than as a PDF's figures.
  2. A science-focused research report, AI in Science: Early Insights, with Google DeepMind and MIT FutureTech — see AI Adoption in Scientific Work. It combines a new science filter over the existing v1.0 corpus with two genuinely new datasets: an inventory of 2,690 specialized scientific AI models and a survey of 637 US/UK scientists fielded by More in Common in July–August 2026.
  3. A statement that ATLAS is a long-term program — "we'll work with partners in academia and elsewhere to identify new areas of research and deliver new insights" — which is the first explicit commitment to future iterations.

It is not a second measurement window. The science report's telemetry is "approximately 15 million anonymized interactions … sampled in early April 2026" — the same April 6–19 corpus as v1.0, re-filtered to ~360,000 science interactions. Every v1.0 limitation on this page carries over unchanged, including the enterprise/paid-API exclusion and the absence of agentic surfaces, which §6.1 of the science report restates in its own words.

The blog carries ATLAS-general cuts that are in no report. India's arts/design/media occupations at 19% of work AI usage (1.6× the global average); US computer and mathematical occupations at 30%, double the rest of the world; manual-task usage at 7% of work AI use in Brazil and Germany against 4% in Japan; Brazil and the UAE adopting above what their GDP per capita predicts. These are blog statements about the v1.0 dataset with no published method behind them — cite them to the announcement at vendor-claim weight, not to either paper.

Where ATLAS and the AEI disagree#

The two programs measure the same construct and get materially different levels. ATLAS names most of these itself:

QuantityATLAS v1.0Anthropic Economic Index
Share of O*NET tasks with observed usage~20%36% (Handa et al. 2025); 49% combined (Appel et al. 2026)
Automation vs. collaboration<10% automation intent for non-routine cognitive work; 26.9% for routine cognitive43–45% automation, 52–57% augmentation — and rising over time
GDP elasticity of adoption0.90.7

The task-share gap ATLAS attributes to its stricter privacy thresholds. The automation gap is largely definitional: Anthropic uses a binary augmentation/automation split, ATLAS a five-category intent classifier where "Task Automation" means end-to-end execution of the core task and everything short of that (drafting, review, ideation, retrieval) counts as collaboration. The two are not measuring the same cut, which is itself the finding — the headline "is AI automating or augmenting?" number is an artifact of where you draw the line. Treat direction as corroborated across labs and levels as method-dependent.

Limitations (author-stated)#

  • No paid Gemini API content — so enterprise usage via Google Cloud is under-represented. Paid API request counts by country are used for geography only.
  • No Workspace, AI Overviews, Translate, Maps, Flow, Antigravity, or Gemini Notebook — products with billions of users, and the ones where agentic coding and world models would show up.
  • Behavioral interactions, not outcomes. A completed conversation is not evidence that the user's goal was met, time was saved, or value was created.
  • Current adoption, not potential. Non-adopters are invisible; this is the frontier of use, not the frontier of usefulness.
  • Probabilistic classification. Granular occupation and activity findings carry more uncertainty than major-group trends — see Usage-Telemetry Classifier Validation for how much more.
  • Two weeks in early April 2026, which overlaps US tax season — ATLAS flags this as inflating tax-adjacent occupations.
  • Observational throughout. No causal identification of anything.

Reports in this wiki#

Connections#

Open Questions#

  • ATLAS excludes paid API, Workspace, AI Overviews, and Antigravity — the surfaces where agentic and enterprise usage concentrate. Does the "shallow, collaborative, non-automating" picture survive when v2 includes them, or is it an artifact of measuring the consumer surfaces?
  • ATLAS and the AEI disagree by 2–4× on automation share and task coverage. Would running both classifiers over both labs' logs reconcile the gap, or is cross-lab usage measurement structurally incomparable?
  • The work-share inversion has three candidate explanations (goal-directed usage under data costs, leisure dilution in rich countries, excluded enterprise subscriptions) and ATLAS endorses none. Which one is it?

Sources#

  • Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy — Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy (Iscenko, Strand et al., Google & Google DeepMind, July 23, 2026); §2 Data and Methods, §5 Global Diffusion, Appendices A–B
  • Google AI & Economy ATLAS: AI in Science (September 2026) — AI in Science: Early Insights (Codreanu, Imas, Mateos-Garcia et al.; Google, Google DeepMind, MIT FutureTech; September 2026, 42pp, empirical with vendor COI), plus the 2026-09-15 announcement post reproduced at the head of the raw document. Cited here for §2.2.1 (the shared April 2026 corpus), §6.1 (exclusions restated), Appendices 1–3, and the blog's program statement and ATLAS-general country/occupation cuts
§ end
Cited by 14
Related articles