H
Howardism
Plate IIAI Economics & Labor中文HOWARDISM

Experimental Learning Impact of Generative AI

Randomized experiments on how AI use changes learning. Contractor & Reyes (n=211): AI access raises test scores +0.27 SD, ~76% persisting unaided a week later, but durable gains go to 'augmentation' users (AI as tutor) while 'automation' users' gains vanish. Bocconi x OpenAI 2x2 (n=1,053): a taught causal-reasoning skill survives the tool while the rubric rewards the tool. LMU preregistered trial (n=704): metacognitive feedback on one's own AI use cuts answer offloading (OR 0.47) and lifts unaided scores (OR 1.51); a points penalty does neither

Article metadata
Publication details
Published:July 16, 2026
Filed:Concept
Domain:AI Economics & Labor
Reading:35 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Experimental Learning Impact of Generative AI

Sources#

Summary#

Zara Contractor and Germán Reyes (Middlebury College, arXiv 2607.08849, July 2026) run the study most of the AI-and-learning corpus was missing: a randomized controlled experiment with an objective, proctored, unaided skill measure. 211 undergraduates learn an unfamiliar technical topic (blockchain, carbon capture, or CRISPR) and write an analytical essay in a 35-minute lab session, randomly assigned to AI-allowed or AI-forbidden; one week later all of them return and are tested and write again without any AI or resources. The headline: AI access raises immediate test scores by 0.27 SD, about 76% of that gain still shows up a week later unaided, and essay quality rises only after AI is removed. But the durable part is not uniform — it belongs almost entirely to students who used AI to explain concepts (augmentation), while students who used it to draft their text (automation) post large short-run essay gains that vanish completely once AI is gone. This is the near-controlled test of "outsource your thinking, not your understanding" and the objective-skill measure the automation–optimism link's self-report could not supply.

Evidence note. empirical and causal (randomization, F-test balance, double-lasso controls, ITT + TOT/2SLS) — the strongest design in this vault's learning cluster. But scope it honestly: elite liberal-arts undergraduates (mean GPA 3.68, SAT 1386, >80% already AI-adopters), a proctored 35-minute lab task, off-the-shelf ChatGPT (GPT-4o), one topic, one week. It measures learning per unit of time with time-on-task held fixed (AI changed how the hour was spent, not its length — plausibly because the lab offered few competing uses of time). The authors explicitly do not claim that real-world AI adoption raises learning overall: outside the lab students choose how long to study, and many use AI to save time. It informs, but does not settle, workplace skill-atrophy.

The design in one paragraph#

Between-subjects randomization at the lab level (eight time slots × two parallel labs). Session One: baseline 5-question test → 35-minute learning phase (read, search, draft a ~500-word essay; AI-allowed group gets a logged-in ChatGPT tab, AI-forbidden gets Google) → unaided 5-question post-test. Compliance enforced by proctors, ChatGPT logs, and interface screenshots. Session Two (~7 days later, everyone unaided): 10-question test + a 20-minute essay on the same topic, complementary prompt. Learning is measured two ways: knowledge tests (factual/conceptual recall) and essays (higher-order analysis), the latter graded blind by 311 master's/PhD Prolific graders plus an LLM grader, averaged, with objective linguistic features (length, readability, lexical diversity, textual similarity, and Pangram AI-detection) alongside.

First stage was strong: assignment lifted AI use by 67.3pp off a near-zero control; 68% of the treated used AI; 88% found it helpful. In the logs, the most common use was explaining concepts (57%), then drafting (31%), then summarizing (17%).

Durable test-score gains, concentrated in the middle — and skewed to the able#

  • Immediate: +6.7pp on a 56.3% control baseline = +0.27 SD (ITT; TOT +10.0pp / +0.40 SD). More than double the 0.10 SD median of Kraft (2020)'s 747 education RCTs, comparable to the 0.29 SD of structured human tutoring, ≈ a one-SD jump in GPA, ≈ $8,600/student of school spending.
  • Retention (one week, unaided): +5.1pp = +0.27 SD, ~76% of the immediate effect persists (the two are statistically indistinguishable). Durable, partially decaying knowledge.
  • Shape: the gain sits in the middle of the score distribution — the tails (0/1 and near-perfect scorers) barely move (Figure 4).
  • Who gains: larger in the upper GPA/SAT quartiles (bottom quartile ~0.05 SD; upper quartiles 0.19–0.40 SD) — so AI access may widen learning gaps, a genuine tension with the floor-raising story of software democratization.
  • Felt vs. measured: self-assessed knowledge shows no treatment effect (β=0.02). Students learned more but did not feel they had — the mirror image of the felt-vs-measured worry the AEI self-report leaves open.

Higher-order skills surface only after AI is removed#

In Session One, treated essays are longer, simpler, and 12.3pp more AI-flagged (a 96% jump), but overall quality rises only slightly and imprecisely (+0.17 SD, p=0.25) — because the essay is partly AI-written, blending output with learning. Notably, no homogenization: within-group essay similarity is flat, contra the convergence Brynjolfsson et al. (2025) found in workplace writing (open-ended prompts admit many valid answers).

The clean read comes from Session Two, written unaided a week later (Figure 6): the AI-detection effect collapses to 0.0pp and the stylistic differences fade — proving the Session One style shift was AI text entering essays, not durable change — while quality gains emerge: writing style & clarity +0.30 SD (p=0.016) and relevance to prompt +0.26 SD (p=0.041). The learning was real and it raised higher-order skill, not just fact recall — but it only became visible once the AI crutch was taken away.

Augmentation vs. automation: the controlled test of "outsource thinking, not understanding"#

The paper's most load-bearing result. Classifying treated users from their ChatGPT logs: 49% augmentation (AI works with the student — tutor, explainer), 32% automation (AI does the work — drafts the essay), 8% mixed, 11% off-topic. Validation that the split is real: automation users' essays are 53.6% AI-flagged vs 20.9% for augmentation users, and automation users spend 17% less time reading/searching.

Session 1 essay qualitySession 2 essay quality (unaided)Session 2 test (unaided)
Automation users+0.54 SD (p=0.029)+0.02 SD — vanishes+0.19 SD (p=0.33, n.s.)
Augmentation users+0.05 SD+0.22 SD (p=0.22)+0.29 SD (p=0.06)

Automation's Session One advantage is AI-produced output that does not survive its removal; augmentation's gains persist unaided, consistent with skill accumulation. This is a within-experiment demonstration that AI's short-run productivity boost and its long-run learning effect can point in opposite directions, decided by how the tool is used — the exact shape of Karpathy's thesis, and consistent with Strömberg et al. (2026) (AI raises homework, lowers exams, losses concentrated among outsourcers), Bastani et al. (2025) (base GPT-4: −0.19 SD later), and Shen & Tamkin (2026) (engineers who learned a library via AI scored worse unassisted). The meta-analysis (Figure 5, grand mean 0.18 SD across 22 estimates) sorts the whole literature along this axis: losses cluster where AI could do the practice for you, gains where AI plays coach.

Mechanisms#

  • Time reallocated, not saved. No effect on total learning time (unlike the large workplace time-savings of Noy & Zhang 2023), but the mix shifts −5.3pp writing / +4.4pp reading-and-searching — effort moves from producing text to absorbing it (task execution → stewardship, à la Copilot users spending half their time verifying).
  • Enjoyment up 13% (+0.66 pts, p=0.031): AI made learning more engaging, not mechanical.
  • Cheating up but not the driver: rule violations rose +12.6pp combined (p=0.005), yet a back-of-envelope bound attributes at most ~2.2pp (about a third) of the test-score gain to cheating — the effect is mostly real learning.

Beliefs: experience corrects the magnitude, not the direction#

Both groups expect AI to help, but only the treated gauge how much correctly. Control students predict a +25.2pp own gain — about 5× the actual 5.1pp; treated students, who used it, predict +3.8pp, close to the estimate. Perceived gains track actual gains across subgroups. And students already hold the right model: 69% name both a help and a harm channel, 54% say the effect "depends on how AI is used," and their two most-named mechanisms (AI explains/tutors 40%; AI shortcuts the work 42%) map exactly onto augmentation vs automation. Students grasp the mechanism from the outset; firsthand use only calibrates the magnitude — a belief-updating result that rhymes with the professional-misjudgment literature (Becker et al. 2025 / METR: developers predicted +24% speedup, measured −19%).

The practitioner statement of the automation half (Ng, August 2026)#

Andrew Ng — co-founder of Coursera, founder of DeepLearning.AI — says the deskilling half of this result out loud, as a general claim, to a general audience (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion): "frankly AI models are terrible for learning… Students score higher on homeworks when they use AI. Yay, higher homework scores. But retention, their long-term performance is much worse because their AI do the work for them." His mechanism is this page's: "when you ask AI to do work for you you're cognitive offloading to AI which is great because that's how society moves forward and gets work done but human retention is much worse." His datum is himself: six months after asking a model how a front-end/back-end component works, "I don't remember the answer… I ask AI again."

Set against the design above, the claim is right about one arm and wrong as stated. Homework up, retention down is the automation row of the table — Session One essay quality +0.54 SD, Session Two +0.02 — and Strömberg et al.'s homework-up-exams-down pattern in the meta-analysis. It is not the augmentation row, where gains persist unaided (+0.29 SD on the test a week later), nor the pooled ITT, where ~76% of the immediate gain survives. Ng's qualifier — "LLMs as they are most commonly used are terrible for learning. I'm not saying there's no way to use it in a way that is good for learning" — is where the two meet, and it turns his claim into an empirical one about the mix: in this experiment's logs the most common use was explaining concepts (57%), and the augmentation share (49%) exceeded the automation share (32%). Whether that mix holds outside a proctored lab, where grades reward output, is the third open question below; Ng's "as most commonly used" assumes the answer. His "the data is very clear" cites nothing, and the meta-analysis on this page (grand mean 0.18 SD across 22 estimates, sign decided by use mode) is what the data actually looks like.

The conflict of interest runs in the unusual direction: he makes the claim while announcing LearnVector, a venture building one-to-one learning experiences that his diagnosis makes necessary. That does not make the diagnosis wrong — it matches the randomized result for automation-mode use — but it is a reason the claim arrives without its augmentation half.

What it settles — and what it doesn't#

It settles, for this population and task, that off-the-shelf AI can produce durable, objectively measured learning gains, dissolving the "AI must erode learning" prior — but conditions the result on use mode. It does not settle: whether the gains hold outside a proctored lab where time is endogenous; whether they hold for non-elite learners (they skew to the able, suggesting not uniformly); or whether the same augmentation/automation split governs workplace deskilling, where the analogue evidence is still mixed (Brynjolfsson's support agents learn; Budzyń's endoscopists deskill).

The sibling RCT: train the human instead of restricting the tool (Bocconi × OpenAI, 2026-08)#

The design above asks what AI access does. A second randomized experiment asks the question this page's third open question actually poses — whether an upstream intervention can change what novices do with the model — by crossing LLM access with a taught cognitive skill. Asirvatham, Betti, Brown, Camuffo, Chatterji, Fumagalli, Gambardella, Mariani, Pandey, Ramos, Salvucci & Simic, Training novices to think, or giving them LLMs? (Training novices to think, or giving them LLMs? Evidence from an RCT, CEPR DP21882, August 2026, empirical): a preregistered 2×2 with 1,053 first-year undergraduates at Bocconi, randomized at the level of 13 intact class sections (Control 249 / Causal 256 / GPT 197 / Causal×GPT 351, the combined cell oversampled for power on the interaction). Arm one is a ~6-minute causal-reasoning game — 12 binary-choice items in four sections, each with a worked example and feedback, against a placebo that poses the same 12 items with no scaffolding and no feedback; the words "causal reasoning" are never used and the scenarios are city-policy, not business, to avoid anchoring. Arm two is ChatGPT Edu (GPT-4o) in a second browser tab. Both arms then write a ≤180-word consultant recommendation on a real merchandising problem in 45 proctored minutes, scored 1–5 on awareness and usage by three of twenty trained, condition-blind master's raters.

COI, and it is structural. Three of the twelve authors are OpenAI-affiliated or OpenAI-contracted, the treatment is ChatGPT, the paper is published on cdn.openai.com, and the measurement stack is OpenAI's — the causal-reasoning attributes are LLM-scored against a rubric, ideas are extracted by GPT-5.2 (with Claude Opus 4.6 as second judge), similarity to experts is cosine distance in text-embedding-3-large. The one outcome outside that loop is the human-rated awareness/usage score, and it is the one showing the largest pro-LLM effect, which cuts the other way. The counterweight worth naming: the result least useful to an LLM vendor — that only the human training arm raises idea diversity — is reported as the paper's headline finding.

The training worked, and the tool did not crowd it out. Relative to control, the causal arm raises mechanism identification ~0.55 SD (+0.241, s.e. 0.028) and falsification logic ~0.85 SD (+0.869, s.e. 0.062); GPT alone moves neither (−0.008 and +0.029, both n.s.). GPT instead raises coherent logic ~0.45 SD (+0.214, s.e. 0.043), where causal training adds only ~0.15 SD. Critically, every Causal × GPT interaction is positive and significant — mechanisms +0.129 (p<0.01), falsifiability +0.203 (p<0.05), coherent logic +0.228 (p<0.01) — on top of the main effects. Trained novices with a model in reach still reason like trained novices; the skill is complementary to the tool, not substituted by it. (Causal participants also finished the training game ~2.7 min faster and scored ~11 points higher on it, Figure A.1 — a pre-GPT measure, so it grades the training, not the tool.)

But the evaluation rewards the tool, not the training. GPT raises the awareness/usage score by +0.862 (s.e. 0.079, p<0.01) against a control-group estimate of 2.09 on the 1–5 scale — about 1.0 SD of the outcome's own dispersion (SD 0.843) — and moves solutions measurably closer to three independently interviewed domain experts (+0.026 / +0.028 / +0.044 cosine similarity on means of 0.794 / 0.753 / 0.806, all p<0.01). Causal training scores negative (−0.111, s.e. 0.040, p<0.01) and the interaction is indistinguishable from zero. Post-double-selection lasso decomposes the GPT effect: 0.862 → 0.490 → 0.448 → 0.362 as content, diversity and text-style controls enter, so ~42% (the paper's prose rounds this to "about half") survives everything, i.e. GPT improves substance and not only polish. Inter-rater agreement is itself a treatment outcome: ICC 0.534 / 0.493 in the GPT arms against 0.275 / 0.322 in the non-GPT arms — LLM-assisted text is not just better-scored but easier to score.

The negative result is where the page earns its place: the rubric penalizes exactly what the training produces. In the same lasso, coherent logic (+0.355) and number of ideas (+0.092) load positively while falsifiability (−0.109), mechanisms (−0.166) and between-solution idea diversity (−7.676) load negatively, all p<0.01 or better. Scaled by each variable's own SD, the conformity premium (−0.29 points per SD of distance from the modal answer) is roughly 1.6× the reward for within-solution variety (+0.18 points per SD) — and causal training's +0.47 SD of between-solution diversity, run through that coefficient, is −0.14 points, which is essentially the whole negative Causal coefficient. The causal arm is the only treatment that raises idea diversity (within-solution +0.55 SD, between-solution +0.47 SD; GPT's effects are +0.12 SD and zero), and GPT neither adds to it nor destroys it. The authors' reading, which this page adopts: the training did not fail — the assessment did not ask for what the training produced.

What it cannot say, in the authors' own words: "Whether the second component reflects knowledge that participants acquired and retained, or output they procured without acquiring anything, we cannot say." There is no unaided post-measure — no removal of the tool, no retention test, no second session. On this page's axis that is decisive: the Bocconi design measures assisted output, where Contractor & Reyes measure what survives withdrawal. It therefore cannot produce an automation row, and nothing here contradicts the vanishing-gains result; the two experiments measure different halves of the same sentence. Two further scope limits: randomization is at the class level (13 clusters, standard errors clustered on mother tongue × high-school type × class), and the task is deliberately one on which "LLMs are likely to be well-trained" — a documented marketing framework with a codified rubric, i.e. the most favourable possible terrain for the model and the least favourable for a novice's own concept.

The third RCT: leave the tool alone, change the offloading decision (LMU Munich, 2026-09)#

Contractor & Reyes observe the use mode and Bocconi trains the user beforehand. Maier, Schwabe, Schneider & Feuerriegel, Designing Against Deskilling (Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants, arXiv 2609.20143, 2026-09-17, empirical, a CHI 2027 submission) is the first design in the vault that manipulates the use mode inside the session, and it does it without restricting the assistant. The study is a preregistered (AsPredicted) online experiment with N=704 UK Prolific adults (mean age 39). It runs a 2×2 plus a no-AI control: 146 control, 137 AI-only, 144 metacognitive feedback, 138 effort-based reward, 139 combined. Participants did 10 fraction-arithmetic learning items next to a GPT-OSS-120B chat panel, then 6 unaided transfer items (new operands, same operations). The assistant was prompted to give method guidance by default and a complete answer only on explicit request. An LLM-plus-numeric-check classifier then labeled each item by the highest help level requested: none, guided help, intermediate result, or complete answer. "Answer offloading" means requesting the complete answer, so this study classifies the automation row itself instead of inferring it.

The two interventions target the decision itself, not what the model will say:

  • Metacognitive feedback. After each item, an LLM-written three-part card: How you used the assistant, What that meant for your practice, and Something to notice, plus a bar placing the item between "All with you" and "Mostly the assistant". It is informational, not instructional: it never says how much help to ask for next.
  • Effort-based reward. A correct item scores 10 points with no help or guidance only, 5 after an intermediate result, and 1 after a complete answer.

Results (Table 3, reconciled against the prose):

HypothesisOR [95% CI]one-sided pVerdict
H1: AI access lowers unaided score0.80 [0.42, 1.51].244not supported
H2a: feedback cuts answer offloading0.47 [0.22, 1.00].026supported
H2b: feedback raises unaided score1.51 [0.98, 2.33].030supported
H3a: reward cuts answer offloading0.66 [0.31, 1.40].139not supported
H3b: reward raises unaided score0.76 [0.50, 1.18].890not supported

The feedback × reward interaction is null on both outcomes. The participant-level ANOVAs in Appendix H reach the same verdicts: feedback cuts offloading by −5.05pp and raises the score by +5.28pp, and the reward's offloading p is.0501.

What makes it worth a section:

  • Pricing drove out the wrong help. The reward cut any assistant use sharply (OR 0.39, p<.001), and guided help trended down (OR 0.69, p=.060). But among the items where participants did use the assistant, answer offloading rose to 62.8% of requests, against 58.2% in AI-only (Figure 5). Feedback cut it to 51.2% and moved the rest into guided help (34.8%) and intermediate results (13.9%).
  • The two arms felt different. The reward was the most effortful arm: perceived effort was higher under the reward than under feedback in both phases, Δ≈0.6 on a 9-point scale, p=.010. The authors' design rule follows: scaffold productive engagement rather than pricing it.
  • Feedback breaks repetition, not first use. The feedback arm was as likely as AI-only to use the assistant at all (66.7% vs 68.6%) and to offload at least once (38.2% vs 39.4%). The difference is persistence: after offloading one item, participants offloaded the next one 69.7% of the time without feedback and 55.9% with it (interaction OR 0.54, p=.040).
  • The deskilling link is correlational. Across conditions, each +10pp of answer offloading goes with 32% lower odds of solving a test item unaided (OR 0.68 [0.62, 0.74]). Offloading itself was not randomized, and the authors flag prior ability as a possible confound.
  • Feedback changed behaviour, not judgment. Metacognitive accuracy (predicted vs actual score) and sensitivity (confidence AUC) did not differ between conditions. Every condition under-predicted its own score, including AI-only (+0.53 items) and control (+0.93) (Appendix I.3).
  • Offloading depends on the person, and the fix works for everyone. The participant random intercept carries 82% of the latent variance in offloading. Low perceived confidence (OR 0.32 per point) and low need for cognition (OR 0.52) predict it; trust in AI does not. None of these moderates the feedback effect, so there is no sign of intervention-generated inequality.

The H1 null disagrees with the literature on access, and the paper names the likely reason. Bastani et al. (−0.19 SD), Liu et al., and this page's own automation row all find unrestricted access harms unaided performance. Here, access did not significantly harm it. The authors attribute this to the default: complete answers came only on explicit request. Fewer than 40% of AI-only participants ever asked for one, against 61% of Liu et al.'s participants who used their assistant mainly for answers, and most of Bastani's unrestricted transcripts. This is a scope condition, not a supersession. Neither result overturns the other, because they tested different assistants. It turns "does access deskill?" into a question about the interaction default. It also sits beside Contractor & Reyes's positive ITT, where off-the-shelf ChatGPT was used mostly to explain concepts (57%). Both suggest the damage is concentrated where answer delivery is frictionless.

How far it reaches.

  • Effect size. Both confirmatory effects clear one-sided tests only: the H2a interval touches 1.00 and the H2b interval crosses it. The feedback arm's unaided score (0.67) is at, not above, the no-AI control's (0.65). The authors state this themselves: reducing answer offloading preserves learning, but no intervention yet beats a no-AI baseline.
  • Duration. A single session with an immediate test on fraction arithmetic, so durability is untested.
  • Classifier. It was validated only on synthetic dialogues: 198 development and 198 held-out, with perfect agreement. There is no human audit of the 3,969 real participant messages.
  • Answer-on-request default. The study depends on it, so the feedback effect is measured on an assistant already built to make offloading deliberate. Whether it survives an assistant that volunteers answers is the authors' own stated limit.

Connections#

  • The Tragedy of the Cognitive Commons — the framework this experiment's augmentation/automation split is the cleanest measurement of: Lovett's "augmentation without internalization" predicts exactly the vanishing of automation users' gains, and names it the mechanism no employment statistic can detect

  • AI-Exposed College Majors at Labor-Market Entry — where classroom offloading may reach the labor market. The Census graduate study proposes that AI-inflated grades (Chirikov 2026: largest rises in homework-heavy, AI-exposed courses) stop signalling skill, and that employers respond by hiring fewer graduates. If so, the automation that suppresses learning in these RCTs also degrades the credential

  • Organizational Complements to AI — where the HAT substitution model lives, and the mechanism-level check on its prediction P7 (workers respond to substitution pressure by upskilling): whether that response works is the augmentation/automation split measured here, since automation users' gains vanish once AI is removed — a distinction the model's private-effort formulation has no term for

  • The Household Production Boundary — the scale of the demand this experiment speaks to: education is 20.7% of global AI conversations and over-represented 5.8× against time actually spent on it in the US, higher still in developing countries — making the augmentation-vs-automation distinction measured here the one that decides whether that volume compounds or hollows

  • The Automation–Optimism Link — the primary complement. The AEI survey found heavy delegators self-report no learning loss; this experiment supplies the objective, randomized measure that survey could not — and it both agrees (augmentation users' learning persists unaided) and exposes the divergence the survey can't see (automation users' gains are hollow, and "automation share" pools both types)

  • Outsource Your Thinking, Not Your Understanding — the near-controlled test of the thesis: augmentation (AI helps you understand) builds durable skill; automation (AI does the thinking) leaves nothing once removed — Karpathy's principle, randomized

  • AI Brain Fry — both put an objective, measured number on AI's cognitive effect (there, oversight fatigue → +11%/+39% errors; here, learning → +0.27 SD, or hollow gains for automators) against the softer self-report signal; the deskilling half of this paper is that page's mechanism in a learning task

  • Returns to Expertise in Agentic Coding — the heterogeneity rhymes: gains skew to higher-ability students, and augmentation ≈ using AI to deepen the understanding that Anthropic's study finds is what amplifies an agent; both say the benefit accrues to whoever brings (or builds) understanding

  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the belief half: like reported/anticipated exposure, students' perceptions of AI's effect are directionally right but miscalibrated in magnitude until firsthand use corrects them

  • Printing Press Software Democratization — the tension: democratization predicts AI raises the floor, but here gains concentrate in the upper ability quartiles, hinting AI may widen learning gaps rather than close them

  • The Data Wall and the Validation Commons Are One Supply Constraint — this experiment used as one of three corroborating measurements that validation capacity can erode before any capability threshold arrives (beside Budzyń's endoscopists and Vicente & Matute's 80.7%). What it supplies that the others don't is randomization; what keeps it from settling the commons question is population and duration — one week, elite undergraduates, a proctored task rather than a career

  • Configurable Human Participation — the system-side mirror, published the same week: HAS-Bench's agency scale runs on the same augmentation-vs-automation axis (A1–A2 automation-oriented vs A3–A5 augmentation-oriented), and both land the same shape of result — the value of AI–human collaboration is decided by its mode and timing, not its amount (there: more agency at A4 breaks tasks A3 solved; here: automation-mode use yields hollow gains that vanish with the tool)

  • Andrew Ng — states the automation half as a general claim ("terrible for learning"), with a stake in the remedy; the section above sorts his claim against the two arms

  • The Three Loops of AI-Native Building — the builder's side of the same author's learning claim: Ng's typing app is cheap implementation deliberately spent on a non-offloading tool for a child, in the interview where he calls LLMs terrible for learning

  • Does the Augmentation/Automation Split Govern Skill at Work? — the workplace transfer of this page's use-mode split, tested against everything the vault holds on working adults. The sign replicates (Budzyń, Dell'Acqua, Wiles and Vicente & Matute on the automation row; Everett and Brynjolfsson's support agents on the augmentation row) but all six are secondhand and none splits use mode inside a design. The binding disanalogy is not the elite sample or the one week — it is that this experiment's mechanism requires time-on-task held fixed, and that workplace volume generates a third mode (skipped review) a proctored session cannot

Open Questions#

  • Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
  • Gains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
  • The augmentation/automation choice is endogenous to incentives (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half? (Andrew Ng's August 2026 "as most commonly used" assumes the mix is automation-dominated; the experiment's own logs run 49% augmentation / 32% automation, so the assumption is unmeasured, not confirmed.) Partially answered (2026-09-22) by Training novices to think, or giving them LLMs? Evidence from an RCT — narrowly, and on a lever the question did not name. A preregistered 2×2 (n=1,053) randomizes a ~6-minute causal-reasoning training against a placebo, crossed with ChatGPT Edu access, and finds the trained skill is not crowded out by the tool: every Causal × GPT interaction on reasoning style is positive and significant on top of the main effects (mechanisms +0.129, falsifiability +0.203, coherent logic +0.228), and causal training's effect on idea diversity (+0.47 SD between-solution) is undiminished when a model is in reach. So something upstream can change what novices do with an LLM — the first randomized demonstration of that in this vault. Three reasons it is a partial and not a retirement. (i) It is training, not incentive or interface design — the only interface-like element, the one-line nudge "Remember the skills you learned during the game," is bundled inside the training arm and cannot be identified separately. (ii) It never measures the mix: use mode is not classified, conversation logs were collected but not analysed, and the outcome is reasoning style in the assisted output, not augmentation-vs-automation behaviour. (iii) The deskilling clause is untouched — there is no unaided post-measure, no tool withdrawal, no retention test, and the authors state flatly that they cannot say whether the LLM's advantage reflects knowledge acquired or output procured. It also supplies the question's inverse, which is the more useful half: the incentive that exists today pushes the other way. In the same paper's lasso, the evaluation rubric loads negatively on falsifiability (−0.109), mechanisms (−0.166) and distance from the modal answer (−7.676, ≈ −0.29 points per SD) while loading positively on the coherent, idea-dense text the LLM produces — so under a standard rubric the trained behaviour is scored down and the outsourced one up. Any incentive-design answer has to beat that gradient first. Partially answered, on both named levers (2026-10-01): Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants (empirical, preregistered, N=704) randomizes one interface lever and one incentive lever against the same classified use mode. The two split cleanly.
  • Interface works. Per-item metacognitive feedback on the learner's own AI use cuts answer offloading (OR 0.47, one-sided p=.026). It moves the request mix toward guided help and intermediate results (answer offloading 58.2% → 51.2% of assisted items) and lifts unaided transfer (OR 1.51, p=.030). This is the reversal of the deskilling half, at least back to the no-AI baseline (0.67 vs 0.65).
  • Incentive fails. A points schedule penalizing answer requests reduces use (OR 0.39) but not offloading (OR 0.66, n.s.). Among assisted items, offloading's share rises to 62.8% because guided help is what gets dropped.

The question stays open on three counts. First, scale: one online session, fraction arithmetic, immediate test. Second, the default: every arm ran on an assistant that gives answers only on request, which may itself be the stronger lever, and the H1 null suggests it is. Third, real grades: the reward was symbolic points with pay held constant, so the Bocconi rubric gradient above, a real grade, is untested.

  • Does the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer? (Promoted 2026-09-05 lint: the vault now holds the workplace-side pieces to test against — the observed automation-share drift on Anthropic Economic Index, the recruiter case in the controlled-variance study, the self-report on The Automation–Optimism Link, the offloading account on AI Brain Fry.) Partially answered (2026-09-05): Does the Augmentation/Automation Split Govern Skill at Work? — the sign transfers. Budzyń's endoscopists, Dell'Acqua's consultants, Wiles et al. and Vicente & Matute all reproduce this page's removal signature (performance high with the tool, no advantage without it) on non-students, while Everett's collaborative clinical workflows and Brynjolfsson's support agents supply the augmentation row — but all six reach the vault secondhand through The Tragedy of the Cognitive Commons, none classifies use mode inside a design, and none takes an unaided post-measure. The answer to the second clause is not "too unlike to transfer": the disanalogies that actually bite are that this design holds time-on-task fixed (its entire mechanism, −5.3pp writing / +4.4pp reading) while at work AI's headline effect is time saved, and that workplace volume generates a third mode this experiment cannot produce — abdication, the review step skipped outright (Acceleration Whiplash's 31.3% unreviewed PRs, Security Debt of Agent-Generated Code's 81.1% uncommented credential leaks), which the automation arm here is not, since those students still submitted work they had read. Still open, and no longer answerable by synthesis: it needs a workplace study that classifies use mode within-subject from logs and measures unaided skill after AI is withdrawn. #oq/source Checked and recorded as a restatement, not a partial (2026-09-22): When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration (CIVIC-AI workshop whitepaper, practitioner-opinion, no study of its own) makes this question one of six conditions for calling a workplace deployment "augmentation" at all, and argues the transfer runs through a selection effect rather than an analogy: the tasks most amenable to AI delegation are "high-volume, procedurally defined and verifiable," which are "exactly those tasks through which junior workers develop their tacit judgement needed to evaluate outputs, catch failures and push back on algorithmic recommendations." It also writes down the instrument this bullet says is missing, in close to the same terms — a workflow record carrying total effort including verification and repair, plus repeated longitudinal measures of mastery, progression and agency, complemented by employee-level evidence because "formal adoption data may miss unofficial AI use." Nothing in it is measured, no use mode is classified anywhere, and no unaided post-measure is proposed, so the bullet is unchanged: what it gains is a second independent specification of the missing study, co-signed by a labour ministry that would be positioned to run it.

Sources#

  • Experimental Evidence on the Learning Impact of Generative AI — Contractor & Reyes, Experimental Evidence on the Learning Impact of Generative AI (arXiv 2607.08849, 2026-07-09), empirical. §4 immediate effects, §5 retention + augmentation/automation heterogeneity, §6 beliefs, §7 conclusion; Figures 4–8, Tables 4–8.
  • Training novices to think, or giving them LLMs? Evidence from an RCT — Asirvatham, Betti, Brown, Camuffo, Chatterji, Fumagalli, Gambardella, Mariani, Pandey, Ramos, Salvucci & Simic, Training novices to think, or giving them LLMs? Evidence from an RCT (Bocconi University × OpenAI, CEPR DP21882, August 2026), empirical. §4.1 causal-reasoning constructs, §4.2 awareness/usage and expert similarity, §4.3 ideas, §4.4 PDS-lasso decomposition; Figures 1–6 and A.1; Appendix §2 rubric and inter-rater reliability, §4 variable construction, Tables A.3, A.7–A.14. Parse warning — PDF-derived (docling 2.126.0, MLX layout + table stages). Automated checks passed except an author-table weld, but an unflagged standard-error row-shift runs through Tables A.7, A.8, A.10–A.14: parenthesised s.e.'s land on the following row, and in Table A.11 this misattributes the Expert-Alumnus GPT coefficient (0.044***) to the Causal row. Table A.13 is worse — docling dropped the within-solution-diversity coefficients entirely (2.483 / 2.284) along with several s.e.'s, welding the row label into the neighbouring one. Every figure cited above was recovered from pdftotext -layout on the local PDF and reconciled against the prose; no table row is cited unreconciled. Two artefacts of the paper itself, not the parse: the CRT mean is 0.490 in Table A.7 and 0.769 in Table A.8 for the same variable and N (both verified in the reference parse), and the anonymization leaks — the university is withheld "to preserve authors' anonymity" but the appendix names Bocconi twice
  • Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants: Maier, Schwabe, Schneider & Feuerriegel (LMU Munich / MCML), Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants (arXiv 2609.20143, 2026-09-17), empirical. Preregistered on AsPredicted, N=704. Cited from §4 design, §5.1 Table 3, §5.2–5.5, §6.2–6.5, Figure 5 (read from the image), and Appendices H, I.1–I.4. Parse warning, PDF-derived (docling 2.126.0, MLX layout and table stages). Table 2 and Table 3 match pdftotext -layout on the local PDF. Table 8 (Appendix H robustness) is split and shifted in the raw: the H2a estimate (−5.05) and H1's CI upper bound (5.14) land on the section-header rows, and CI brackets weld into the p column. The figures cited above come from pdftotext. Table 5 (the item bank) is cell-collapsed and is not cited.
  • Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28, practitioner-opinion): "AI models are terrible for learning," the homework-up/retention-down claim, the cognitive-offloading mechanism, the six-months-later self-report, the "as most commonly used" qualifier, and the LearnVector announcement. No study cited
§ end
Cited by 23
Related articles