Sources#
- Claude Fable 5 and Claude Mythos 5
- Google AI & Economy ATLAS: AI in Science (September 2026)
- Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods
- Noam Brown – Agent swarms, alignment, & recursive self-improvement
- Normative boundaries of AI in scientific work: Evidence from PhD researchers
- On the Navier–Stokes Millennium Prize Problem
- Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence
- Risk Report: August 2026 (Redacted)
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
Summary#
With Mythos 5 (the bio-safeguards-lifted form of Fable 5), Anthropic reports the first Claude results in which a model conducts novel scientific research largely on its own — choosing experimental moves, running domain tools, recovering from failures, and producing findings that match or beat skilled humans and recent published baselines. This is the wet-lab / life-sciences analogue of AI-Driven Formal Proof Search: where formal proof search has a Lean compiler as an instant verifier, science's verifier is the experiment — slower and more expensive — so the claims here are empirical demonstrations and selected examples, not compiler-checked guarantees. The results are the sharpest evidence yet for the less-conservative reading of recursive self-improvement: that "perspiration is becoming automated" reaches into discovery itself, and that research taste may be "just another capability AI fails at for a time, then gets good at."
The three results#
Drug / protein design — autonomy at human level#
Anthropic's internal protein-design experts accelerated aspects of drug design "by around 10 times" using Mythos 5. In one study, Mythos 5 — equipped with protein-design and bioinformatics tools but no human assistance — matched or beat skilled human operators, executing "all of the tasks normally completed by a scientist: choosing binding sites, selecting and running protein design tools, and recovering from failures along the way." 9 of 14 protein targets yielded strong drug-design candidates now under investigation (immune checkpoints, growth-factor/receptor signaling, neurodegeneration, muscle disease, harder structural targets).
Novel hypotheses — preferred over Opus-class, one corroborated#
Mythos 5 is Anthropic's "first model to consistently produce novel, compelling scientific hypotheses." In blinded head-to-head comparisons against Opus-class models, Anthropic scientists preferred Mythos's molecular-biology hypotheses ~80% of the time, and advanced several to experimental evaluation. One Mythos hypothesis — a novel mechanism for an E. coli protein — was independently corroborated by a study from a lab working on the same problem.
Genomics — a week of autonomy beating a published model at 100× smaller#
Over "more than a week of largely autonomous work," Mythos 5 assembled single-cell data for millions of cells across 138 animal species, then designed and trained a custom machine-learning model to identify cells performing the same role in even distantly related organisms. With only high-level human input, that trained model outperformed a recent model published in Science — despite being 100× smaller. Anthropic intends to publish.
The dual-use shadow#
The same capability is why biology must be safeguarded in the general-access Fable 5. The motivating evaluation: predicting how a genetic modification affects adeno-associated virus (AAV) capsid assembly — a real gene-therapy component whose design capability "in the wrong hands, could enable the design of dangerous viruses." Mythos-class models outperformed dedicated protein-language models on this without being trained for the task, using biological reasoning alone. Autonomous scientific capability and bio-uplift risk are the same capability seen from two sides — the core tension the RSP CB determination and the bio classifier exist to manage.
Why it matters for the trajectory#
- Perspiration automation reaches discovery. When AI builds itself argued most research progress is incremental "scale-it-up-see-what-breaks-fix-it" work that Claude excels at. Autonomous genomics — assemble data, design a model, train it, beat the baseline — is that loop run end-to-end in a science domain, not just engineering.
- It chips at the taste moat. "Consistently produce novel, compelling hypotheses" and "only high-level human input" are exactly the direction-setting functions presumed to stay human. The ~80% blinded preference is a concrete crack — though still human-judged and internally sourced.
- Still jagged, still gated by verification. These are curated demonstrations (Jagged Intelligence (Ghosts, Not Animals)); science's verifier is slow wet-lab confirmation, not a compiler, so unlike AI-Driven Formal Proof Search the results can't be auto-validated — they await experimental and peer review. This keeps it adjacent to, but below, the AI-R&D autonomy threshold Anthropic gates on.
The easy end of the verification spectrum, and what it bounds (August 2026)#
This page is framed by the gap between an instant verifier (Lean's compiler) and a slow one (the wet-lab experiment). Idea Search (arXiv 2608.08958, Caltech / Google Research / Harvard, empirical) is worth recording here precisely because it sits at the near-Lean end: single-cell RNA-seq batch integration is scored by the OpenProblems v2.0.0 benchmark — a fast, automated, deterministic metric over already-collected CZ CELLxGENE data. Nothing in its loop waits on a pipette. The system runs 2,000 search nodes per trial and 5 trials per configuration because each evaluation is cheap, which is exactly the regime the wet-lab results above cannot reach.
That makes it a bound on transfer, not a counterexample to the verification gap. Two things follow, and both cut against reading automated discovery results optimistically:
- This is the friendly case, and the gain is modest. With a fast scorer, a fixed backbone and an explicit bank of expert-derived ideas, the system improves a strong Tree Search baseline's mean from 0.678 ± 0.011 to 0.697 and its best solution from 0.694 to 0.728 — roughly 0.02, which the paper concedes is comparable to its own 0.008–0.018 trial-to-trial spread. Where the verifier is free and instant, plateau-breaking still produced a sub-σ mean shift and a heavier tail. Any extrapolation to domains where each evaluation costs a wet-lab cycle should start from that, not from the headline.
- A fast scorer converts the verification problem into a specification problem. The metric defines what counts as better for the entire run, so "verified" here means "scored well by OpenProblems", not "true." That is the substitution AI-Driven Formal Proof Search escapes only because Lean's compiler is sound rather than merely independent — a benchmark metric is neither. Autonomy scales cleanly against a fast verifier and inherits every gap between that verifier and the thing you actually care about.
Priced against expert time, and an RCT that says novices got little (August 2026)#
Anthropic's August 2026 Risk Report assesses the same dual-use capability from the CB-2 side and supplies the two things this page has lacked: a time-denominated estimate of expert uplift and an independent randomized trial on novices. They point in opposite directions, and both are load-bearing.
Expert uplift, denominated in working days. In a beneficial red-teaming tabletop exercise, PhD biologists paired with dedicated LLM experts designed an end-to-end biological resistance strategy against a hypothetical engineered agricultural pathogen. Some teams held generalist biological expertise; others were world-leading domain specialists.
Two of three generalist teams outperformed all three specialist teams on both scientific quality and feasibility. Expert graders estimated the strategies and implementation protocols would have taken 40–95 working days to produce without AI tools; the two-person teams accomplished this in 16 hours.
That is a 20–50× compression on a real end-to-end scientific design task, and generalists beating specialists — the sharpest single measurement of expertise substitution in the corpus. Expert red-teaming panels separately described Mythos 5 as "the strongest they have evaluated," with two biology experts rating it comparable to or exceeding a knowledgeable specialist and several reporting it supplied work they would otherwise have sought from a specialist consultant.
Novice uplift, measured by a third party: modest. A 2026 RCT (Hong et al.) gave participants with minimal prior laboratory experience access to frontier LLMs from all developers without blocking biological classifiers, and had them attempt complex procedures in a real laboratory over several weeks. Its conclusion: "mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit." Anthropic treats this as "substantially more reflective of the kind of real-world uplift we are concerned with" than its own system-card evaluations — and lists the caveats itself: a post-hoc analysis suggested 1.42× uplift on a typical task with large error bars; AI-assisted participants used the models only moderately (one participant exceeded 1M total tokens); and the study was underpowered, with only 36% power to detect an odds ratio of 2.0. The update Anthropic takes is that Opus 4/4.1-generation models pose somewhat lower risk than feared.
Where Anthropic says the capability stops. The limitations it considers most disqualifying for a substitution threshold are the same two this wiki records elsewhere: "weak open-ended ideation and design" — the model reliably recombines and extends published knowledge but "rarely produced approaches reviewers considered genuinely novel," and tends toward over-engineering — and "poor strategic judgment": extending whatever framing the user supplies rather than challenging it, executing plans containing flaws it had itself detected, presenting timelines reviewers repeatedly forced it to retract, missing how errors compound across a multi-step program. Anthropic's operational conclusion is that this is why uplift arrives through "extended back-and-forth interactions" rather than a single usable plan, which is the whole basis for the argument that classifiers blocking sustained access are an effective mitigation.
And the domain experts say the transformation is not the LLM's. The report's 31 non-AI-domain interviews (AI R&D Autonomy Evaluation (AECI)) repeatedly credit non-LLM ML with the actual breakthroughs — protein structure prediction transforming antibody/enzyme design in biotech and nanotech, superconductor candidate screening in energy, image annotation in connectomics collapsing "tens of thousands of work-hours by hundreds of students per dataset to a few dozen hours" — while LLMs save time on coding, data analysis, literature review and figure-making. Physical bottlenecks (lab robotics, in-vivo validation, clinical trials) are named by nearly every domain as the binding constraint, and one biotech interviewee identified robust laboratory robotics specifically as "the bottleneck to unlocking the kind of dramatic acceleration that the RSP envisions."
The boundary a survey draws through all of this (September 2026)#
The SJTU/Theseus RSI survey (arXiv 2609.11873, §4.1 and Appendix C.1, practitioner-opinion) reviews the AI-for-science self-improvement literature and insists on a distinction this page does not currently make, because its sources are about capability rather than about persistence:
Improving a molecule, equation, or hypothesis alone remains RSI-adjacent rather than evidence that the scientific agent has improved itself.
Every result on this page is on the wrong side of that line by construction — a 20–50× compression on a design task, hypotheses preferred 80% of the time, a corroborated E. coli mechanism. All are outputs. The survey's question is whether anything persists into the next research cycle: HypoForge distilling critiques and execution outcomes into reusable hypothesis-generation skills, Test-Time Tool Evolution synthesizing and validating executable scientific tools for reuse (its SciEvo benchmark reports 1,590 tasks supported by 925 evolved tools), DrugSAGE carrying verified modeling procedures across tasks (experience from 16 tasks improving performance on 17 held-out ones — "direct evidence that inherited experience changes behavior on new problems"), SIA extending the editable surface to scaffolds, tools, search logic and LoRA weights.
Its verdict places science at L2 — "the strongest current frontier, with systems autonomously selecting or applying updates to model weights, tools, skills, workflows, and scaffolds" — with a few L3-like systems doing cross-task experience reuse, and end-to-end L4 environment adaptation and L5 evolution of the improver itself "largely unexplored."
Why science stalls where software engineering does not is a statement about feedback, not about models, and it is the same argument this page's domain-expert interviews make from the other end. Three properties: discovery is open-ended, so neither the hypothesis space nor the required tools are known in advance; feedback is non-identifying — "a failed experiment may indicate an incorrect hypothesis, but it may also result from an inappropriate model, an invalid protocol, or an unreliable instrument"; and validity is conditional on assumptions, domains and measurement procedures, so "an update that succeeds in one setting cannot be inherited safely without recording its provenance, scope of validity, and uncertainty." The requirement that follows — explicit skill preconditions, execution tests, provenance, versioning and rollback, because "without these controls, a locally successful workflow may be inherited outside the conditions under which it was validated" — is a sharper form of the physical-bottleneck point the 31 interviews make, and it binds even where the bottleneck is computational rather than robotic.
A third point on the verification spectrum: mathematics without a compiler (September 2026)#
This page is organized by verifier speed — Lean's compiler at one end, the wet-lab experiment at the other, with Idea Search's automated metric added as the near-Lean bound. Brown's account of OpenAI's Navier-Stokes result (Noam Brown – Agent swarms, alignment, & recursive self-improvement, 2026-09-17, practitioner-opinion) sits in a place none of them occupies: a domain with a sound verifier in principle and no automated one in the loop.
What Brown reports, and every figure is first-party and unverifiable: one of the Millennium Prize Problems solved by a system of 10,000 agents spending 130 billion tokens over 88 hours, on an unreleased internal model. Nothing in the account mentions formalization, Lean, or machine-checked proof; the verification is mathematicians reading the output. (Superseded 2026-09-21 by On the Navier–Stokes Millennium Prize Problem — accurate about Brown's account, wrong about the event. OpenAI's own announcement, published nine days before the interview, claims both a Lean formalization and a verification pass, "an additional 17 hours via GPT‑6 Astra," on top of the 88-hour search. The original sentence is kept because it correctly describes the source it was written from, and because the correction changes where the result sits on this page's spectrum.) So the result's epistemic status at time of reporting is closer to this page's genomics claim ("Anthropic intends to publish") than to a compiler-certified theorem — and, like the genomics claim, its open question is what survives peer review.
Where the correction leaves it, and it does not move as far as it looks. A claimed kernel check is not a kernel check. The announcement is vendor-claim, the formalization claim is one clause with no axiom discipline, no statement review and no library version, the linked Lean repository and proof PDF are not in this corpus, and no independent re-check existed as of 2026-09-21 — so the verification is still, in the only sense this page's spectrum cares about, mathematicians (and now a wiki) reading a vendor's prose. What changes is the kind of uncertainty: the genomics claim awaits an experiment nobody has run, while this awaits someone opening a repository that already exists. That is a materially cheaper falsification path than anything else on this page, which is why The Navier–Stokes AI Claim files it as the top open question rather than a caveat.
Three things it contributes that the wet-lab results cannot.
- The cost is disclosed and enormous, which nothing else here does. 10,000 agents × 88 hours is the corpus's largest single published expenditure on one research question. Against METR's dollar axis this is the extreme end of a curve whose other points are $10,000 trajectories. What it does not come with is the control — Brown says the single-agent counterfactual was never run and the ablation at that scale is unaffordable (Compute-Controlled Benchmarking).
- It is the cleanest case of "perspiration automated" in a discipline with no perspiration to automate. The wet-lab results above are impressive partly because the model ran instruments. Brown's framing of maths is that "you're purely bottlenecked by thinking" — no protocol, no reagent, no robot. If the essay's incremental-work thesis is right, maths is where it should show first and most cleanly, and this is that datum.
- And the same speaker draws the boundary this page keeps testing, harder than Anthropic does. Brown rejects the replacement narrative outright: "there is a narrative going around that these things are replacing mathematicians, that it's just superhuman in mathematics across the board. I think that is the wrong takeaway. … they are weaker than human mathematicians in other ways… They're not very good at posing new problems. They're not really good at understanding what directions, what whole branches of mathematics are worth exploring or developing." That is the ~80%-hypothesis-preference claim above met by the opposite reading from the other lab — problem solving at an extraordinary level, problem posing still absent — and it is the sharpest live disagreement in the corpus about where taste sits. The two are not straightforwardly contradictory (molecular-biology hypotheses within a defined problem area are not "which branch of mathematics is worth developing"), and the distinction between them is exactly the Boden level-2/level-3 line.
His own stated preference is the complement world, which is unusual enough from a capabilities researcher to record: "I would be thrilled to live in a world where AI is a complement to human abilities and is allowing us to discover new knowledge without fully replacing people. That is the best-case scenario." Followed immediately by the concession that he does not expect it to hold — "as they get better, they get better across the board… Over time, it is possible that they're just better across the board. Now, I don't know how long that takes. It depends on how long the long tail is of things that they're bad at."
The adopter side, measured on 637 working scientists (September 2026)#
Everything above is a capability claim from a lab. Google/DeepMind's ATLAS science report (Google AI & Economy ATLAS: AI in Science (September 2026), empirical with vendor COI) is the first population-scale measurement of what AI acceleration is doing to the working scientist's week — not autonomy, but the assisted regime that precedes it, and it lands on exactly the two constraints this page argues from.
The verification gap has a size now, and it is a fraction of the dividend. Of the scientists who report net time savings from AI (just under three quarters of 637, averaging 6.9 hours/week), 89.3% spend more than 10% of the saved time verifying, debugging or fact-checking AI output, 45.7% more than 25%, and 9.3% more than half (Figure 13C) — highest in the Life Sciences. Self-reported, and the report's own footnote 25 says this correlation with AI intensity is the weakest of its three: "less robust, marginally statistically significant … depending on specification," where the downstream-bottleneck shift and the backlog growth are strongly significant.
And the downstream bottleneck is where the scientists put it, independently of Anthropic's 31 domain interviews. 43.5% say their primary rate-limiting bottleneck moved downstream over two years against 13.8% upstream; physical experimentation and data collection is the single largest current bottleneck (24% overall, 30% in physical sciences); 40.5% report a growing backlog of untested hypotheses against 24.8% shrinking. That is the same claim the biotech interviewee makes about lab robotics — "the bottleneck to unlocking the kind of dramatic acceleration that the RSP envisions" — arrived at from a survey panel rather than from expert interviews, on a different vendor's data. Two instruments, two tiers, one answer: generation is not the constraint; execution and validation are.
A third result cuts at the taste question from an unexpected angle. Asked how AI changed the risk profile of the projects they choose, 48.8% say it pushed them toward safer, more incremental questions "where data and AI capabilities are well-established," against 27.5% toward higher-risk ones (Figure 14). The report reads this as a possible Streetlight Effect. If taste is the residual human function, the first measured effect of AI on that function is not that models took it over but that their availability moved where humans point it — toward the tractable. That is a different failure mode from the one this page tracks, and neither the ~80% hypothesis-preference result nor Brown's problem-posing boundary anticipates it.
The obvious limits: it is self-report about perception, from a screened non-probability panel with no population frame, commissioned by a company whose product is being measured, on assisted rather than autonomous use, with no causal identification anywhere.
The next cohort is least willing to concede exactly the tasks autonomous systems are built to take (2026-09-23). A second adopter-side instrument, with no vendor COI and a different population: Angelini & Lyrvall's latent class analysis of 3,785 PhD students in STEM and medical/health sciences (Normative boundaries of AI in scientific work: Evidence from PhD researchers, arXiv 2608.25678, empirical) asks how comfortable they are with AI performing five research activities. Designing experiments — the capability every autonomous-discovery claim on this page turns on — draws 30.5% comfortable against 45.2% uncomfortable, statistically indistinguishable from writing (31.0/50.0) and data analysis (29.5/49.9), while literature tracking and summarisation are the only activities with net acceptance (43.1% and 51.5%). The dominant 44% 'division of labour' profile is organised on exactly that split. Two things follow for this page. First, the constraint on autonomous science is not only verifier speed: the population that would have to hand over hypothesis generation and experimental design currently reports the least willingness to do so, and career stage does not predict it (the first-two-years coefficient is null), so this is not a cohort that ages out. Second, the paper is careful about what it measures — comfort, not legitimacy, and not behaviour — and the same ordering is equally consistent with a plain reliability judgment, which is the reading most favourable to the capability claims above: if resistance tracks current model reliability on those tasks, it moves when the models do.
The outside view: one survey draws the whole map, and it is less impressed than this page (September 2026)#
Everything above is a lab's own report or a measurement of adopters. Jedlička's Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence (Petr O. Jedlička, Institute of Philosophy, Czech Academy of Sciences — arXiv 2608.17970, 2026-08-18, 40pp, practitioner-opinion) is the corpus's first attempt to survey the whole field from outside any of it: a single-author review with no original measurement, covering capability trends, scientometric diffusion, a typology of research-AI systems, a discipline-by-discipline tour, and the limitations. It is secondary on every primary result this page holds, so nothing here supersedes what the vault compiled from the primaries. What it contributes is a map, one category the vault does not have a page for, and two attributed judgments that cut against the acceleration reading.
The typology, and where the vault already holds each category#
The survey's organising move (§4) is four categories ordered by role in scientific practice, not architecture — explicitly "heuristic and not mutually exclusive," and explicitly a ladder of increasing autonomy and scientific participation:
| Category (survey's names) | Survey's exemplars | Where the vault holds it |
|---|---|---|
| Specialized Scientific AI | AlphaFold 1→3 (CNN → Evoformer/Pairformer), MACE-MP, SMI-TED | Transformative Creativity (AlphaFold as the Boden level-2 reference case); no dedicated page |
| Scientific AI Assistants | ChatGPT / Gemini Deep Research, SciSpace, NotebookLM, Cursor — "fundamentally reactive" | Deep Research Agents |
| Scientific AI Agents | Sakana's AI Scientist, Google's Co-Scientist, FutureHouse's Robin; also called "AI scientists," "co-scientists," "hypothesis machines" | this page, Many-Agent Proof Harnesses, Open-Ended Discovery Harnesses |
| Hybrid AI Experimental Systems | ChemAgents, Recursion — AI reasoning wired to robotic platforms, "self-driving" labs | nothing |
The survey itself notes the assistant/agent boundary is porous — ChatGPT Deep Research, Claude Research, Gemini Deep Research and Perplexity Labs "already assume more agentic functions," so the two are "points on a continuum of increasing autonomy" rather than kinds. That matches what Deep Research Agents records from the harness side.
The fourth row is the gap worth naming. This page's spectrum is organised by verifier speed — Lean's compiler, then an automated metric, then the wet-lab experiment. The survey's fourth category is the case where the system runs the slow verifier itself, which is a different axis: not how fast the check is, but whether the loop closes without a human moving a pipette. Its two instances, both secondary here:
- ChemAgents (Song et al., JACS 147, no. 15 (2025): 12534–12545) — a hierarchical multi-agent robotic chemist on Llama-3.1-70B that reviews literature, designs experiments, operates instruments, analyses results and refines the next round. It discovered metal-organic high-entropy catalysts comprising five metallic elements, synthesising and measuring across 100 automated experiments before optimising the composition.
- Recursion — a wet-and-dry platform (BioHive-2 on NVIDIA hardware) running "millions of multi-omic experiments each week," which the survey credits with advancing therapeutic candidates into clinical testing: for familial adenomatous polyposis, a MEK1/2 inhibitor now in phase 1b/2 (Samadder 2026, a company investor-relations presentation — first-party, and the weakest-sourced claim in the section).
The two judgments that cut against this page#
Jumper's own number: AlphaFold made structural biology "approximately 5–10 per cent more efficient." (§6.2, from a Science Friday interview.) The survey's footnote 50 concedes in full: "This estimate is not yet based on any systematic metrics." Take it as the tier it is — one creator's offhand estimate, practitioner-opinion, unmeasured. But it is the corpus's only attempt by a principal to put a field-level number on the flagship specialized-AI result, and it sits at the opposite end from this page's demonstration figures (a ~10× on drug-design tasks) and from ATLAS's self-reported 6.9 hours/week. The reconciliation the survey offers is the one this page already argues: AlphaFold "accelerates specific stages within a much larger and more complex scientific pipeline," and the pipeline's binding constraints — toxicity, solubility, experimental validation, clinical testing, regulatory approval — are untouched. Task-level acceleration, field-level rounding error is the same shape as the complements thesis and the same shape as the downstream-bottleneck finding above, arrived at by a third route.
Derek Lowe's novelty critique of Co-Scientist, and it is the sharpest published challenge to the claim family this page rests on. Lowe (In the Pipeline, Science, 2026-05-27) argues that (a) the human contribution to the Co-Scientist results — framing the problem, selecting among candidate solutions — was substantially greater than the paper acknowledges, and (b) the acute myeloid leukaemia pathway that Co-Scientist presented as a novel suggestion was not new: the line of research had already been investigated, and the prior work was not cited. The survey adds that "similar reservations apply to some of the other putatively de novo results." Two consequences for this page. First, it is a worked instance of the failure mode Sakana's own evaluation reports — that AI Scientist "frequently misidentified generated ideas as novel" — showing up in the more credible system too. Second, it is the obvious stress test for Anthropic's ~80% blinded hypothesis-preference result above: preference by the lab's own scientists is not a novelty check against the literature, and no published account of that result describes one. The one Mythos hypothesis independently corroborated by another lab is corroboration of correctness, not of novelty.
Where the survey lands, and the inversion in its last section#
Its verdict (§8) is the conservative one: "AI systems have not produced paradigm-changing discoveries comparable to the most transformative ones in the history of science." The successes are "highly specific problems, which however already existed within established scientific frameworks," and AI works best "as powerful cognitive tools within existing research programmes designed and directed by human scientists." That is Boden level 2 stated without the vocabulary, and it is the same boundary Brown draws for mathematics.
The interesting move is the one it makes after conceding that. If future systems differ qualitatively, the binding constraint flips from the machine to the reader: AI-generated hypotheses and frameworks "may become increasingly difficult for humans to evaluate, interpret, or even comprehend," and may look "implausible, unintelligible, or hallucinatory from the perspective of contemporary scientific paradigms, despite containing genuine insight." The survey points at the 'alien space of science' hypothesis (Artiles et al., arXiv 2603.01092, 2026) — that scientific communities systematically miss research directions for which they lack the conceptual tools — as the positive version: AI could reach regions of the space that are cognitively unavailable to human researchers. That is a claim about a ceiling in the evaluator rather than in the generator, and it is the one thing in the survey that neither this page nor Research Taste as the Human Bottleneck currently frames. It is also, as posed, unfalsifiable on present evidence.
The mathematics chapter, and the Leiden Declaration (moved from AI-Driven Formal Proof Search, 2026-09-29)#
The survey's §5.1 calls mathematics and computer science "some of the clearest evidence of the rapid improvement of AI capabilities." Its account of AlphaProof Nexus, AlphaEvolve, FunSearch, the OpenAI Erdős work and Sakana's AI Scientist is secondary and less specific than the primaries on AI-Driven Formal Proof Search. Four items it adds, none of them in this corpus:
- Peer-review status, date-stamped. Footnote 21 (accessed 2026-06-11): "these solutions of Erdős open problems have not been yet published in a peer-reviewed journal." A third party outside the labs, covering the OpenAI discrete-geometry disproof and, by the sentence's scope, the Erdős solves generally.
- Gowers on the OpenAI result. Timothy Gowers "described the result as a milestone in AI-assisted mathematics": a combinatorial-geometry conjecture disproved by a general-purpose reasoning model, examined by professional mathematicians and found to involve "unexpected connections between elementary geometry and algebraic number theory." Set beside Tao's more reserved line, two of the discipline's most-cited voices land in different places on the same class of result.
- Knuth on the natural-language branch. Knuth's manuscript Claude Cycles (a PDF on his Stanford page, not a paper) reports that Anthropic's Claude Opus assisted in resolving a graph-theoretic conjecture concerning Hamiltonian cycles, which he "described as evidence of a dramatic advance in automated deduction and creative problem solving." No detail on the conjecture, model version or division of labour, and no verification account; it belongs to the unformalized branch and is an endorsement, not a result.
- A hybrid instance. Ju et al., Automated Conjecture Resolution with Formal Verification (arXiv 2604.03789, 2026), report "the autonomous resolution and Lean 4 verification of an open problem in commutative algebra": an informal reasoning agent paired with a proof assistant, the informal-search/formal-check split in a subdiscipline no compiled result covers.
The Leiden Declaration. §7.1.1 surfaces the Leiden Declaration on Artificial Intelligence and Mathematics (Alper, Barany, Chavarri Villarello, Dahmen, Dean, Ganapathy, Harris, Holmes, Jamnik, Kelk, Kra, Martin, Naskręcki, Ochigame, Portegies & Schmitt, Zenodo, 2026, 10.5281/zenodo.20302944), a multi-author community statement that "warns among other things against the risks that AI poses to correctness, rigour, and standards of proof." The survey's reason: "although such problems are not unique to AI, the scale and speed with which AI systems can produce and disseminate information may greatly amplify their impact." That is the review-capacity argument of Verification as the New Bottleneck in mathematical form, and the institutional counterpart of Logical vs Intelligible Proof: a certificate is not the same thing as rigour, and machine-proof throughput is a hazard to standards whether or not any single proof is correct. The declaration's text, signatory list and asks are not in the vault; everything here is the survey's one-sentence characterization. Treat it as a pointer to fetch.
An executable instrument for the prior-art question, still short of this page's own claims (September 2026)#
The Lowe critique above and this page's own open question ask the same thing in different registers: is an AI-generated result actually novel, or trivially reachable from public information and generic competence? CMU's Discovery Certification Protocol (arXiv 2609.09219, empirical) is the first executable version of that check in the corpus. Its Gate 2 gives a fresh, matched challenger agent the same registered background and every Web byte the target run observed, withholds the target's own research trail, and treats any valid method reaching the score within tolerance as a disqualifying recovery witness. Two complete audits — SQLite query-plan optimization and a virtual chemical-process control task, neither a science-discovery claim of the kind this page tracks — each returned zero recoveries in 96 matched-challenger episodes, with a finite-sample recovery-probability upper bound of 0.0468, certifying both results as Core + Evidence.
What it would take to close the gap, and what it does not yet close. Gate 2 is exactly the "independent, pre-registered prior-art search" the open question below asks for, formalized and made replayable by a deterministic verifier rather than left to a lab's own scientists' preference judgment. But nothing in this source runs it on an open-ended hypothesis-generation claim — both audits are sealed-score engineering optimization under a numerical threshold, not "does this molecular-biology mechanism already exist in the literature." The instrument exists; its application to this page's actual claims (the ~80% hypothesis preference, the corroborated E. coli mechanism) does not yet.
Connections#
-
Discovery Certification Protocol (DCP) — the executable recovery-audit instrument this page's prior-art open question has lacked, demonstrated on engineering-optimization tasks rather than open-ended hypothesis generation
-
Deep Research Agents — the survey's Scientific AI Assistants category, and its concession that the category does not hold: ChatGPT/Claude/Gemini Deep Research and Perplexity Labs "already assume more agentic functions," so assistants and agents are points on one autonomy continuum rather than kinds
-
Open-Ended Discovery Harnesses — the harness-engineering face of the survey's Scientific AI Agents category: what the multi-agent "AI scientist" systems it names (AI Scientist, Co-Scientist, Robin) look like when the question is loop design rather than taxonomy
-
Organizational Complements to AI — where Jumper's "5–10% more efficient" lands: a flagship task-level result that does not reach field-level output because the residual pipeline (toxicity, validation, clinical testing, approval) binds
-
AI Adoption in Scientific Work — the adopter-side complement to this page's capability claims: what AI adoption is measurably doing to scientists' time budgets, bottlenecks, verification load, and choice of question — and, from a second source on the same page, the acceptance constraint beside the capability one: 3,785 PhD students accept AI for literature work and resist it for experiment design, writing and data analysis
-
Many-Agent Proof Harnesses — the Google-side counterpart to the frontier-lab discovery claims, published with a paper rather than a blog post: five mathematics/TCS results from a many-agent harness (all in self-authored companion preprints, none peer-reviewed or formalized) plus the sharpest study in the corpus of rediscovery under information isolation — Colosseum, internet disabled, independently reaching the central architecture of OpenAI's Erdős unit-distance counterexample over 15 exploration rounds
-
The Navier–Stokes AI Claim — the first-party document behind the Navier-Stokes result, and the source of the correction above: OpenAI claims a Lean formalization and verification in 17 hours, which moves the result's position on this page's verifier-speed spectrum if and only if the claim is checked
-
Noam Brown — the second-hand source for the Navier-Stokes result, its cost, its missing control, and the problem-posing boundary he says has not moved
-
Evaluation Horizon Versus Release Cadence — why this result is not available to check: the model that produced it is withheld, and the reason given is that safety evaluation at long horizons no longer fits inside the release interval
-
RSI Autonomy Levels (B0–L5) — the distinction this page's evidence does not satisfy, and should be read against: every result here is a better output, while the RSI question is whether anything persists into the next research cycle. The survey places science at L2 and calls L4/L5 in this domain largely unexplored, for a reason that matches this page's own domain interviews — scientific feedback is non-identifying, so a failure cannot be attributed to the hypothesis, the protocol or the instrument
-
Structured Safety Case (Claim Decomposition) — the CB-2 threshold assessment this evidence feeds, and the substitution framing that decides it
-
AI-Driven Formal Proof Search — the formal-math sibling: AI doing novel research, but with an instant compiler-verifier; science substitutes the (slow, costly) experiment, so verification is the harder bottleneck here
-
Recursive Self-Improvement — the clearest wet-lab evidence for "perspiration is becoming automated," the essay's less-conservative reading
-
Research Taste as the Human Bottleneck — autonomous hypothesis-generation and "only high-level human input" are direct chips at the residual human comparative advantage
-
AI R&D Autonomy Evaluation (AECI) — adjacent autonomy: a model designing+training a model and beating a published baseline is AI-R&D-shaped, though in genomics rather than AI itself
-
Task Time-Horizon Scaling — "over a week of largely autonomous work" is a concrete long-horizon datapoint beyond Mythos Preview's measured 16h
-
Jagged Intelligence (Ghosts, Not Animals) — the caveat: these are selected demonstrations of a still-jagged capability, not uniform competence
-
The Verifiability Thesis — the limiting case: science is less verifiable than Lean proof, so autonomy outruns cheap verification — the experiment, not a compiler, is the reward signal
-
Capability-Gated Model Fallback — the dual-use flip side; the AAV result is the bio classifier's motivating example
-
Responsible Scaling Policy Evaluations — the CB (chemical/biological) risk domain these capabilities advance
-
Claude Mythos 5 — the model (bio safeguards lifted) that produced these results
-
Claude Fable 5 — the general-access sibling on which biology is safeguarded
-
The Abstraction Barrier — the live test of DeepMind's barrier: do these results cross it (novel primitives) or operate within human-defined spaces with the embodied bottleneck still gating wet-lab validation?
-
Transformative Creativity — whether autonomous hypothesis-generation is climbing from Boden-exploratory toward transformative (new-conceptual-space) creativity
-
Automated Conjecturing — the same discovery loop in the domain where the verifier is instant and the hypothesis space is enumerable: machine-generated mathematical conjectures, screened by linear program against a table of known theorems and refutation-tested exhaustively. The contrast that matters here is that its hypotheses can be falsified in milliseconds, so its bottleneck is triage (what is worth proving) rather than the wet-lab cycle this page's bottleneck is
Open Questions#
- Every result is Anthropic-reported and example-selected; the genomics "100× smaller beats Science" claim is "intend to publish" — what survives external peer review?
- Science's verification gap: the formal-proof loop self-validates; here a wrong-but-confident hypothesis costs a wet-lab cycle to falsify. Does autonomy without a fast verifier increase the verification bottleneck rather than relieve it? Bounded, not answered (2026-08-12): Idea Search measures the opposite end of the spectrum — automated discovery where the verifier is free and instant (the OpenProblems v2.0.0 metric over pre-collected data) — and even there the mean gain over a strong baseline is within the trial-to-trial spread. So the friendly case sets a low bar for what the hard case can be expected to deliver, and it relocates the problem rather than removing it: with a fast scorer, "verified" means "scored well by the metric," not "true." The question as posed still needs a source that runs the same method under both verifier regimes. Partially answered (2026-09-22), on the assisted regime rather than the autonomous one: Google AI & Economy ATLAS: AI in Science (September 2026) (
empirical, vendor COI, self-report) measures the bottleneck in a population of 637 working scientists and finds it moving the way this question predicts — of those who save time with AI, 89.3% spend over a tenth of it and 45.7% over a quarter verifying, debugging or fact-checking output (Figure 13C); 43.5% say their rate-limiting bottleneck shifted downstream against 13.8% upstream; 40.5% report a growing backlog of untested hypotheses against 24.8% shrinking. So where the verifier is slow, more generation does buy a larger verification and validation load rather than relieving it, at population scale and outside any single lab's demonstrations. Three things keep it partial: this is AI-assisted, not autonomous science, so the counterfactual the question asks about is never run; it is perception, not measured time; and the verification-tax correlation with AI intensity is the one result in the report its own footnote 25 calls marginal and specification-dependent, while the downstream shift and backlog growth are strongly significant. - The ~80% blinded hypothesis-preference result is a preference judgment by the producing lab's own scientists, not a novelty check against the literature — the exact gap Derek Lowe documents in Co-Scientist's acute-myeloid-leukaemia suggestion and that Sakana's own evaluation reports as AI Scientist "frequently misidentified generated ideas as novel." Does any AI-generated-hypothesis claim survive an independent, pre-registered prior-art search? Partially answered (2026-09-24): CMU's Discovery Certification Protocol (
empirical) supplies exactly this instrument — a fresh matched challenger, given the registered background and observed Web content but withheld the target run's research trail, must fail to recover the result across a registered episode budget. Two complete audits certify it works (0/96 recoveries, upper bound 0.0468, on SQLite query optimization and virtual catalyst control). Still open because neither audited case is an open-ended hypothesis-generation claim like this page's own — the instrument exists, its application to a molecular-biology or genomics claim does not yet. - If hypothesis-generation is genuinely at ~80% preference, how much of "research taste" is left as a distinctively human function — and how would you measure the residue?
Sources#
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus Labs / Tsinghua / ByteDance / ModelBest / Xiaohongshu / Shanghai AI Lab / Humanlaya / Agent-Native Research Lab / Frontis.AI), arXiv 2609.11873, 2026-09-10, 79pp (
practitioner-opinion). Cited here for §4.1 and Appendix C.1 only: the RSI-adjacent boundary, the three properties that make scientific feedback hard to inherit from, the per-system review (HypoForge, Test-Time Tool Evolution, CASCADE, DrugSAGE, STELLA, SIA, CORAL, SAGA), and the L2-frontier verdict. Secondary source — none of those primaries is in the vault. Full treatment on RSI Autonomy Levels (B0–L5) - On the Navier–Stokes Millennium Prize Problem — OpenAI (no byline), "On the Navier–Stokes Millennium Prize Problem", openai.com, 2026-09-08 with a 2026-09-10 update, ~1,900 words,
vendor-claim. Cited here only to supersede the "no formalization" reading above and to bound the correction: the post claims a Lean formalization and verification at 17 hours via GPT‑6 Astra, and publishes none of the discipline that would make it checkable. The proof PDF and Lean repository it links were not fetched at ingest. Disputed on priority, independently unverified as of 2026-09-21, total first-party COI. Full treatment on The Navier–Stokes AI Claim - Noam Brown – Agent swarms, alignment, & recursive self-improvement — Noam Brown (OpenAI) with Dwarkesh Patel, Dwarkesh Podcast, 2026-09-17 (
practitioner-opinion). §00:00:00 for the Navier-Stokes result (10,000 agents, 130B tokens, 88 hours) and the un-run single-agent control; §00:22:02 for the jagged-in-maths reading, the problem-posing and branch-selection boundary, and the complement-world preference with its own hedge. First-party about an unreleased model, with no paper, no formalization, no verification account and no peer review at time of compile — the result is attributed to Brown throughout and is not stated here as an established mathematical fact. The "130B tokens ≈ 4,000 human-years" gloss is the host's arithmetic - Claude Fable 5 and Claude Mythos 5 — §"Evaluating Claude Fable 5 and Claude Mythos 5" (drug design; novel hypotheses; genomics) and §"Biology and chemistry" (AAV dual-use)
- Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods — Wang, Cui, Brenner & Venugopalan (Caltech / Google Research / Harvard), arXiv 2608.08958 (2026-08-09,
empirical): §4.1 the scRNA-seq testbed, CZ CELLxGENE data and the OpenProblems v2.0.0 scoring protocol; §5 and §6 for the results and the authors' own effect-size caveat. Cited here as the fast-verifier bound, not as a discovery result in its own right. Zero tables in the source; numbers quoted from prose, figures read under the two-pass rule. Full treatment on Transformative Creativity - Risk Report: August 2026 (Redacted) — Anthropic, Risk Report: August 2026 (Redacted), RSP v3.4 (
empiricalin method, first-party in provenance; the Hong et al. RCT it cites is independent). §4.4.2.1 (the RCT and its four caveats), §4.4.3 and Table 4.4.3.A (expert red-teaming, the tabletop exercise's 40–95 working days vs 16 hours, the catastrophic-scenario uplift trial, RNA design and AAV capsid results), §4.4.3 prose (the two disqualifying limitations and the extended-interaction conclusion), §3.6 (the 31 non-AI-domain interviews). Table 4.4.3.A is split across four docling blocks; the tabletop and red-teaming figures quoted here are restated in the surrounding prose and were reconciled against it. Parse note: ingest verifywarnontable-collapse(5 cells), all confirmed false positives;table-shiftclean; canary-recall 19/20 - Google AI & Economy ATLAS: AI in Science (September 2026) — AI in Science: Early Insights (Google, Google DeepMind, MIT FutureTech, September 2026, 42pp;
empiricalwith vendor COI — Google measuring Gemini, survey commissioned by Google DeepMind). Cited here for §5.3 and Figures 13–14 only: the verification tax, the downstream bottleneck migration, the untested-hypothesis backlog, and the safer-versus-riskier question split, all self-reported by 637 screened US/UK scientists. Full treatment, sampling caveats and the telemetry half on AI Adoption in Scientific Work - Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence — Petr O. Jedlička (Institute of Philosophy, Czech Academy of Sciences), Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence, arXiv 2608.17970, 2026-08-18, 40pp, ~12k words,
practitioner-opinion. A single-author survey with no original measurement and no primary results of its own — every claim cited here is secondary. Used for §4 (the four-category typology and the assistant/agent-continuum concession), §5.2–5.3 (ChemAgents, Recursion, Co-Scientist, Robin), §6.2 (Jumper's 5–10% estimate and its footnote-50 disclaimer; Lowe's Co-Scientist novelty critique; the AI Scientist weaknesses), and §8 (the no-paradigm-change verdict and the 'alien space of science' inversion). Deliberately not carried: its restatements of AlphaProof Nexus, the OpenAI Erdős results, AlphaEvolve, FunSearch, Navier–Stokes-adjacent maths, HLE and ARC-AGI scores, and METR time-horizons — the vault holds those from their primaries and the survey adds nothing to them. COI/parse: the author discloses using GPT-5.5 for language and literature exploration in a paper about AI in science; funded by a Czech Ministry of Education grant. PDF-derived (docling 2.126.0, MLX layout and table stages, 40pp, 0 tables, 2 pictures, confidenceexcellent) — no table risk, but §3's paragraph order is scrambled in the docling text flow (sentences from adjacent paragraphs interleaved); every figure quoted here and on AI Adoption in Scientific Work was re-read frompdftotext -layouton and matches. Charts 1–2 read as images under the two-pass rule. Full scientometric treatment on AI Adoption in Scientific Work; the mathematics half on AI-Driven Formal Proof Search - Normative boundaries of AI in scientific work: Evidence from PhD researchers — Angelini & Lyrvall, Normative boundaries of AI in scientific work: Evidence from PhD researchers, arXiv 2608.25678, 2026-08-26, 15pp,
empirical. Cited here only for the experiment-design and task-ordering marginals (Table 1), the four-class solution and the null career-stage coefficient (§4–4.1), and §5.4's statement that the items measure comfort rather than legitimacy. Self-selected international sample from Nature's Graduate Survey 2025; full treatment and caveats on AI Adoption in Scientific Work - Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents — Ning, Zhong, Li & Zeng (CMU), arXiv 2609.09219, 2026-09-07,
empirical. Cited here only for Gate 2's recovery-audit design as an executable prior-art instrument. Full treatment on Discovery Certification Protocol (DCP)
Cited by 33
- Open Questions Backlog×4
Autonomous Scientific Discovery (113d) — If hypothesis-generation is genuinely at ~80% preference,…
- Transformative Creativity×4
It is also, read against Anthropic's ~80% blinded hypothesis preference, the corpus's sharpest live…
- AI Adoption in Scientific Work×3
The last finding is the one that earns this page. It is the complements thesis measured inside the…
- AI-Driven Formal Proof Search×3
Jedlička's survey (arXiv 2608.17970, practitioner-opinion) is secondary and vaguer than this page's…
- Discovery Certification Protocol (DCP)×3
The E0/L\ boundary is the whole mechanism: an observation present at the start of the run belongs…
- The Navier–Stokes AI Claim×3
It corrects the corpus on formalization. Autonomous Scientific Discovery recorded, correctly at the…
- Anthropic×2
2026 June — launched Fable 5 and Mythos 5, the first general-access Mythos-class models (the tier…
- Automated Conjecturing×2
Autonomous Scientific Discovery — hypothesis generation without a wet lab; the same discovery
- Capability-Gated Model Fallback×2
Biology and chemistry. Previously Anthropic blocked only a narrow selection of bioweapons queries;…
- Claude Fable 5×2
The announcement leads with Fable 5 (general access); the science-heavy results run under Mythos 5…
- Claude Mythos 5×2
Run with safeguards lifted, Mythos 5 produced the announcement's most striking results — compiled…
- Many-Agent Proof Harnesses×2
Erdős unit-distance rediscovery: OpenAI's internal model produced the counterexample to Erdős's…
- Recursive Self-Improvement×2
And the same exchange supplies the counter, from Brown, immediately: "The main difference is that…
- Responsible Scaling Policy Evaluations×2
Autonomous Scientific Discovery — the CB-domain capability (AAV, autonomous bio) that sharpens the…
- RSI Autonomy Levels (B0–L5)×2
Autonomous Scientific Discovery — the S1 feedback regime, and the boundary the survey draws through…
- Task Time-Horizon Scaling×2
The June 2026 Mythos-class release pushes further still: Fable 5 / Mythos 5 "can work autonomously…
- The Abstraction Barrier
Autonomous Scientific Discovery — June 2026 wet-lab AI results are the empirical test case: do they…
- AI R&D Autonomy Evaluation (AECI)
Autonomous Scientific Discovery — adjacent autonomy in a non-AI science domain (a model…
- Compute-Controlled Benchmarking
Autonomous Scientific Discovery — the corpus's best-disclosed and least-controlled headline at…
- CS329A: Self-Improving AI Agents (Stanford)
The course's own bet on where value lands: repetitive, verifiable work — code migrations, version…
- Deep Research Agents
Autonomous Scientific Discovery — where this pattern sits in a survey typology of research-AI…
- Evaluation Horizon Versus Release Cadence
Autonomous Scientific Discovery — the concrete thing on the internal side of the gap: a Millennium…
- Expenditure Horizon
Autonomous Scientific Discovery — the far end of this page's dollar axis, and a reminder of what it…
- Google AI & Economy ATLAS
AI in Science: Early Insights (September 2026, with Google DeepMind and MIT FutureTech) — → Ai…
- Jagged Intelligence (Ghosts, Not Animals)
Autonomous Scientific Discovery — the Mythos 5 science results are curated demonstrations of a…
- Logical vs Intelligible Proof
Autonomous Scientific Discovery — where the Leiden Declaration now lives: a mathematicians'…
- Superintelligence Trajectory
Autonomous Scientific Discovery — Mythos-class models now conduct novel science with limited human…
- Noam Brown
Autonomous Scientific Discovery). Read as an instrument, that is a useful calibration on this page's
- Open-Ended Discovery Harnesses
Autonomous Scientific Discovery — the same systems named as a category rather than as harness…
- Organizational Complements to AI
Autonomous Scientific Discovery — the complements pattern at its starkest: AlphaFold's own creator…
- Research Taste as the Human Bottleneck
Autonomous Scientific Discovery — Mythos 5's autonomous hypothesis-generation (~80% preferred over…
- RSI Growth Curves: Which Friction Binds First?
3. The fundamental wildcard. The Abstraction Barrier in its strong form — if AI trained on human…
- Structured Safety Case (Claim Decomposition)
Autonomous Scientific Discovery — the CB-2 threshold's capability side; the substitution framing…
Related articles
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
