Sources#
- Helping People Choose Careers in the Age of AI
- How AI is expanding what people do at work
- How Organizations Use AI: Evidence from ChatGPT
- Work at the Frontier: How AI is expanding what people do at work
Summary#
Task crossover is OpenAI Economic Research's name for a measured pattern: work historically associated with one occupation appearing in the AI use of people in another. Across 800,000+ work-related messages from US ChatGPT users mapped onto O*NET work activities, 16.8% of all work messages — and 43.5% of occupation-specific ones — concern another occupation's tasks.
The claim underneath the number is methodological, and it is aimed at the field's dominant approach. Most AI-exposure work "begin[s] with a fixed list of tasks and ask[s] which ones a model can perform." Crossover's rejoinder: the list itself is changing. The salesperson explores the dataset that used to go to an analyst; the marketer troubleshoots the site without waiting for a developer. The report's own conclusion states the consequence for measurement plainly — "if AI changes which workers perform particular tasks, measures based only on existing job descriptions will gradually diverge from how work is actually organized."
Evidence note.
empirical, with a first-party caveat matching Returns to Expertise in Agentic Coding and Conversation-to-Delegation Shift: OpenAI measuring its own product, self-classified, no external replication. Descriptive only — the report is emphatic that it "does not estimate AI's effect on employment or productivity" and cannot say "how the same people would have allocated their work without AI."
The three-way split#
The design's load-bearing step is what it sets aside. Activities "broadly shared across occupations" — writing emails, scheduling — are labelled Generic and excluded from the headline ratio:
| Category | Share of work-related messages |
|---|---|
| Generic (shared across most occupation groups) | 61.5% |
| Within occupation | 21.8% |
| Cross-occupation | 16.8% |
| Cross-occupation as a share of non-generic messages | 43.5% |
Both numbers belong together: 43.5% quoted without the generic filter overstates by ~2.6×, while 16.8% alone hides how concentrated crossover is once undifferentiated work is removed. Across the eight groups, crossover runs 11–30% of all messages and 28–77% of occupation-specific messages.
Who borrows, and who lends#
The two directions are close to independent — an occupation can pull in outside tasks, supply tasks others take up, both, or neither:
| Occupation | Borrows (own messages that are other occupations' tasks) | Lends (average share of other occupations' messages that are its tasks) |
|---|---|---|
| Design | 35.2% | 1.7% |
| Marketing | 24.3% | 8.9% — highest outward share |
| Engineering | 18.5% | 7.4% |
Design borrows heavily and lends almost nothing. Engineering is nearly the reverse — a source of work others take on. Marketing is high in both directions at once. The report's summary: "Designers use AI to combine work from several fields. Engineering tasks are frequently taken up by workers elsewhere. Marketing does both."
By cross-occupation share of occupation-specific messages, the heaviest borrowers are customer experience 77%, design 75%, HR 69%, legal 56%, marketing 53% — a majority in five of the eight groups.
Which tasks travel#
Some activities recur across nearly every occupation. Table 2 of the report ranks them by how many of the other seven groups they reach:
| Task | Home occupation | Groups reached | Highest share in |
|---|---|---|---|
| Calculate financial data | Finance | 7/7 | Sales: 14.3% |
| Troubleshoot computer applications or systems | Engineering | 7/7 | Customer experience: 7.6% |
| Discuss goods or services information with customers | Customer experience | 7/7 | Marketing: 70.8% |
| Create marketing materials | Marketing | 5/7 | Design: 13.6% |
| Communicate with government agencies | Legal | 7/7 | Customer experience: 21.4% |
Read the percentages carefully — each is a share within the messages that matched that home occupation, for users in the named group (so 70.8% means: of marketers' customer-experience-related messages, 70.8% are this one task). In absolute terms the flows are smaller: 9.8% of finance-related messages from non-finance users are financial calculation, 6.1% of engineering-related messages from non-engineering users are troubleshooting.
Aggregated by source (Fig. 4), marketing tasks account for 28–29% of non-generic messages among sales and design users and 26% among customer-experience users; engineering tasks account for 28% among design users and 20–22% among customer-experience and finance users.
The workspace-size gradient#
Crossover falls as the workspace gets bigger: 18.9% in 2–5 seat workspaces down to 16.3% at 101+ seats — about 2.5pp, ~13% relative. OpenAI's reading is substitution: in a small business "there is no specialist to delegate to," while in a larger organization the same worker "may be more likely to rely on an established team, workflow, or internal service." AI is most useful as a generalist where specialist resources are scarce.
Three caveats the report attaches, and they matter:
- The gradient holds only for typical-volume users (middle 50% by message count). The top quartile shows no monotonic trend; the bottom quartile falls sharply but is imprecise — the largest-workspace cell holds just 197 messages.
- Heavy users may have "stable AI-supported workflows that are relatively similar across organizational settings," or may simply be doing more iteration inside their own occupation. The report picks neither.
- Workspace seats are not company size, and the comparison is descriptive — seat count proxies specialization, industry, occupation mix, and organizational maturity all at once.
What this does to exposure measurement#
The most consequential claim, and it applies to nearly every number in Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated. Observed, theoretical, and reported exposure are all computed per occupation against a task list held fixed by O*NET. If 43.5% of occupation-specific AI use is another occupation's work, then:
- Observed exposure lands on the wrong occupation. A marketer's website troubleshooting scores as engineering-task usage under a task-to-occupation mapping, but the person is a marketer.
- Theoretical exposure prices a stale bundle — the tasks an occupation held before the borrowing started.
- The denominator is moving. Task Saturation: Broad but Shallow AI Diffusion measures how deeply AI penetrates a reached occupation's tasks; crossover says which tasks belong to the occupation is itself in motion.
None of this makes exposure measures wrong — it makes them a snapshot of an allocation AI is actively rearranging. Exposure asks what share of this job's work could AI do; crossover asks whose job is this work now.
The second, independent problem arrived a week earlier from a different direction. Steele & Cruz put seven exposure instruments on the same occupations and found they barely agree — nearly disjoint most-exposed lists, and an exposure-salary gradient whose sign flips between older and newer instruments (Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated). So the instruments disagree with each other even holding the task list fixed, and crossover says the list itself is moving. The two failure modes are independent and compound: fixing one would not repair the other.
There is one independent corroboration, cited by the report rather than produced by it: Yang et al. (2026) find Perplexity queries regularly fall outside users' inferred primary occupation. A second platform, a different inference method, same direction.
The organizational reading#
Crossover is the population-scale, telemetry-side measurement of a shift the wiki has documented mostly from inside frontier labs. Engineer PM Convergence records Anthropic teams where "everyone codes" and roles dissolve; Printing Press Software Democratization argues software authorship spreads to whoever has the need. Crossover puts numbers on the same move across eight occupation groups of ordinary business users — and adds a structural finding neither had: the convergence is strongest where the org chart is thinnest, a claim about organizational slack rather than about model capability.
It also sharpens the caution on Role Averaging, Not Role Elimination — and here the report agrees rather than resists. Its own framing: "AI may make a task easier for an outsider to attempt while specialists remain critical for expert-level judgment and review," and its closing recommendation is that workers "may need training to evaluate AI-assisted work outside their established expertise." That is the Validation Tether arriving at the same place from the opposite direction: crossover is the mechanism by which work reaches people who lack the Internalized Mastery to substantively validate it.
The limits, stated by the authors#
Unusually complete, and each one bounds a different claim:
- The unit is a message — "not an hour of work, a completed project, or a job." Message counts are not time, effort, or value.
- No outcome data at all: "We do not observe whether the output was used, how good it was, how much time it saved, whether the user could have completed the task without AI, or whether a specialist reviewed it." Every quality question this page raises is out of reach by construction.
- Population: occupations come from self-reported ChatGPT Business Department/Role, while the analyzed messages come from those users' individual accounts. Business and Enterprise are separate product populations and the report says estimates "should not be generalized to Enterprise users." Not representative of the US workforce.
- Classification is embedding-mediated twice: message → IWA → DWA (drawn from that IWA or its five nearest neighbours), then a DWA counts as inside an occupation's boundary at cosine similarity ≥ 0.80 to the closest boundary DWA. A message with several tasks gets one primary label.
- Task source ≠ provenance: the assigned occupation "describes which occupation's O*NET tasks the DWA is closest to, not whether the work was delegated from, or actually originated in, that occupation."
The excluded population, observed (OpenAI Enterprise, August 2026)#
The report above is explicit that its estimates "should not be generalized to Enterprise users." How Organizations Use AI: Evidence from ChatGPT (Chatterji, Holtz, Rakholia, Tambe & Weeratunga, arXiv 2608.12236, empirical) is the same lab measuring exactly that excluded population — 1,764 ChatGPT Enterprise organizations, 17.4M messages, with a task-classified subsample of 973 organizations and 8.7M messages at the six-month-post-adoption mark. Full treatment at The Enterprise AI Adoption Gradient.
What it shows is the shape crossover predicts without being able to measure crossover itself. By job title class, task prevalence "varies in ways that align with job responsibilities" — engineers over-index on technical digital work and debugging, finance and accounting staff on financial and tax tasks, sales and marketing roles on sales and marketing — but those role-specific tasks "do not overpower the small set of core tasks that are performed by all roles." Documentation and technical writing reaches 56.3% of weekly active users and technical digital work 49.6%, across every job title class and every seniority level. A shared production core with role-specific specialization layered on it is what a population with substantial crossover looks like when you stop asking whose task it was.
But it is not a crossover measurement, and the gap is structural rather than a data limitation. The enterprise study assigns messages to one of 60 proprietary task categories, not to O*NET work activities, and it has no occupational boundary construct — no home occupation, no within/cross/generic split, no cosine threshold. Nothing in it can produce a 16.8% or a 43.5%. What it establishes is that the Enterprise population is now observable on task composition by role, and that the instrument needed to answer this page's first open question is one taxonomy change away from existing.
Connections#
- The Enterprise AI Adoption Gradient — the Enterprise population this page's method excluded, measured by the same lab on a 60-category proprietary task taxonomy: role-aligned specialization sitting on a shared core of documentation, technical work and communication, with no occupational-boundary construct and therefore no crossover ratio
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — crossover attacks the fixed-task-list assumption under all four exposure measures: the tasks an occupation "has" are being reallocated faster than the mapping updates. That page's seven-instrument head-to-head is the independent second attack: the instruments disagree with each other even with the task list held fixed
- Task Saturation: Broad but Shallow AI Diffusion — the complementary depth measure; crossover says its denominator (which tasks belong to an occupation) is itself in motion
- Conversation-to-Delegation Shift — the same lab's prior usage telemetry, on the intensity axis (asking → delegating) rather than the boundary axis (my work → someone else's); heavy users break the simple pattern in both
- Agentic Coding Work-Composition Shift — Anthropic's within-tool version: what a session is for shifts over time inside one occupation, where this measures work shifting across occupations
- Engineer PM Convergence — the frontier-lab account of roles merging, here measured at population scale and outside tech
- Role Averaging, Not Role Elimination — the report's own hedge is this page's thesis: outsiders can attempt the task, specialists remain necessary for judgment and review
- Printing Press Software Democratization — the direction-of-travel claim with engineering's 7.4% outward share as a measurement of it
- Organizational Complements to AI — the workspace-size gradient read as complements: AI substitutes least for specialists where the specialists already exist
- Returns to Expertise in Agentic Coding — the open tension: expertise amplifies the agent, yet crossover is people working outside their expertise; both rest on usage telemetry, and neither observes outcome quality
- The Tragedy of the Cognitive Commons — the theoretical frame treating this exact pattern as a risk: work moving to non-specialists is the mechanism by which specialist validation capacity and regeneration pathways erode
- Usage-Telemetry Classifier Validation — the error bar under every number here; both the user's occupation and the task's home occupation are model-mediated judgments
- OpenAI — the publisher; Work at the Frontier is announced as a recurring series
- Cognitive Capability Profiling for Task Suitability — a candidate mechanism for crossover, measured from the cognitive side. Prunty et al. elicit capability-importance profiles for 18 O*NET-derived work activities from 410 workers across six occupational domains and find the domain-specific matrices correlate cell-for-cell at r = 0.53–0.77 (mean 0.63): a shared cognitive core — Planning, Semantic Memory, Working Memory, Language, Procedural Memory — dominates nearly every activity, with domains separating only on secondary capabilities (Warehouse +2.1 on Spatial Reasoning, Manufacture +2.5 on Planning, Hospitality +1.7 on Theory of Mind). If occupations differ mainly in secondary demands, the tasks that cross occupational boundaries should be exactly the ones loading on the core, which is what 43.5% of occupation-specific use being another occupation's work looks like from underneath. Shared dependency worth noting: both take their work-activity vocabulary from O*NET and inherit whatever that taxonomy gets wrong
Open Questions#
- Crossover is measured on consumer-surface ChatGPT messages from Business-account users. Does it hold in agentic/API/Codex usage, where work is delegated rather than typed — or does the occupational boundary reassert itself when the unit is a task handed to an agent? Partially answered (2026-09-23) — the population is now observed, the measure is not. How Organizations Use AI: Evidence from ChatGPT (
empirical, same lab, August 2026) covers exactly the excluded population: ChatGPT Enterprise, 1,764 organizations and 17.4M messages, with task classification on 973 organizations and 8.7M messages. Its task composition by job title class shows the shape crossover predicts — role-specific over-indexing (engineers on debugging, finance staff on financial and tax work) layered on a core of documentation, technical digital work and communication that every role performs, with the paper stating that the role-specific tasks "do not overpower the small set of core tasks that are performed by all roles." It does not answer the question as posed, for a structural reason. The enterprise study uses a proprietary 60-category task taxonomy with no O*NET mapping and no occupational-boundary construct, so no within/cross/generic split and no crossover ratio can be computed from it — and it is still conversational usage, not the agentic/API/Codex traffic the second half of the bullet asks about (the authors defer Codex to their companion paper). Retagged in effect rather than in tag: the blocking constraint is no longer that the population is unobserved but that the observing instrument uses the wrong taxonomy, which is a re-classification of data OpenAI already holds. - The dataset records what people attempted and nothing about outcome — the authors say so directly. Is borrowed work done as well as the specialist would have done it, and where does crossover stop being role expansion and start being unreviewed amateur output? A crossover measure joined to a quality or review-coverage measure would settle it, and nothing in the corpus currently does.
- Is the workspace-size gradient about specialist availability (OpenAI's substitution story) or about permission and norms (a large firm's marketer may be allowed to touch less)? The two predict opposite things as small firms grow, and 2.5pp across a descriptive seat-count proxy is thin evidence for either.
Sources#
- Work at the Frontier: How AI is expanding what people do at work — OpenAI Economic Research, Work at the Frontier: How AI is Expanding What People Do at Work (2026-07-27),
empirical. The full 16-page report; primary source for every figure on this page. Doclingverify: okon all checks — Table 2 (recurring cross-occupation tasks) parsed as a clean grid and was reconciled against the surrounding prose before citing. Figures 1–6 and A1 are images; all values quoted here come from the report's prose, captions, or Table 2, never from a chart read. - Helping People Choose Careers in the Age of AI — Steele & Cruz, arXiv 2607.15506 (2026-07-16),
empirical. Cited here only for the cross-instrument disagreement (§4.4, §5.1, Table 5); parse warnings and a source-internal contradiction are recorded on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated. - How AI is expanding what people do at work — the announcement post for the same study (2026-07-27),
empirical. ~1,100 words, superseded by the report above for every claim. Retained because it carries OpenAI's framing language and the link to the companion AI Jobs Transition Framework. Two provenance notes: the web clipper wrotepublished: 2026-07-31, corrected to 2026-07-27 at compile against four independent dated reports; and the post's heatmap survives in the clip as an empty HTML table (rendered as an image), which is what made the full report worth fetching. - How Organizations Use AI: Evidence from ChatGPT — Chatterji, Holtz, Rakholia, Tambe & Weeratunga, How Organizations Use AI: Evidence from ChatGPT, arXiv 2608.12236 (2026-08-12, 69pp,
empirical). Cited here for §3.2 (the worker-characteristics and task-classification samples), §4.4.1 and Figure 7 (overall task prevalence and message shares), and §4.4.3 and Figure 9 (task composition by job title class). Figure values are two-pass chart reads. No independent author — three OpenAI staff and two academics contributing as paid OpenAI contractors. Full treatment and parse notes at The Enterprise AI Adoption Gradient
Cited by 16
- Cognitive Capability Profiling for Task Suitability×2
The wiki has been circling this convergence from the other side. Task Crossover finds 43.5% of
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated×2
A practitioner's number, placed on the taxonomy (Ng, August 2026). Andrew Ng's public arithmetic…
- Open Questions Backlog×2
Task Crossover: Crossover is measured on consumer-surface ChatGPT messages from Business-account…
- Agentic Coding Work-Composition Shift
Task Crossover — the across-occupation counterpart: this page tracks what a session is for shifting…
- Conversation-to-Delegation Shift
Task Crossover — the same lab's telemetry on the orthogonal axis: this page measures intensity…
- Engineer PM Convergence
Task Crossover — the convergence measured at population scale and outside tech: 43.5% of…
- The Enterprise AI Adoption Gradient
Task Crossover — the population that page's method explicitly could not reach: crossover is…
- AI Economics & Labor
Task Crossover — OpenAI's Work at the Frontier (800K+ US ChatGPT work messages mapped to ONET, July…
- OpenAI
Task Crossover — OpenAI Economic Research's Work at the Frontier series (July 2026): 800K+ work…
- Organizational Complements to AI
Task Crossover — the complements argument read off a firm-size gradient: outside-occupation task…
- Printing Press Software Democratization
Task Crossover — democratization with a number on it: engineering tasks account for 7.4% of other…
- Returns to Expertise in Agentic Coding
Task Crossover — the unreconciled tension: expertise is what amplifies an agent here, yet crossover…
- Role Averaging, Not Role Elimination
Task Crossover — the averaging measured (customer-experience workers spend 77% of…
- Task Saturation: Broad but Shallow AI Diffusion
Task Crossover — the moving denominator under this page: saturation measures what share of an…
- The Tragedy of the Cognitive Commons
Task Crossover — the empirical pattern this frames as a risk: work moving to whoever encounters the…
- Usage-Telemetry Classifier Validation
Task Crossover — a measure that is classifier-mediated twice over: both the user's occupation and…
Related articles
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Organizational Complements to AI
The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…
- Conversation-to-Delegation Shift
OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production…
- Task Saturation: Broad but Shallow AI Diffusion
Google ATLAS's marquee work finding — AI reaches 68% of detailed occupations (88.4% of US employment) but only 21% of t…
- AI Adoption in Scientific Work
Three instrument families on one profession. Google ATLAS telemetry + survey: scientists are the economy's heaviest AI…
