Sources#
Summary#
The Anthropic Institute is Anthropic's research and policy arm focused on the societal and governance implications of frontier AI. It published When AI builds itself (June 2026) — this wiki's primary source on Recursive Self-Improvement — and has a stated agenda to build, in collaboration with others, the systems that a credible AI slowdown or pause would require (Frontier Pause Verification).
What it does#
- Public-facing trajectory analysis. When AI builds itself combines public benchmarks (Task Time-Horizon Scaling) with previously-unreported internal Anthropic data (AI Accelerating AI Development) to argue AI is already accelerating AI development and to lay out three futures for RSI.
- Coordination infrastructure. It plans to "conduct research — in collaboration with many others — and take actions to help build the systems that a credible slowdown or pause would require": verification that other developers have actually stopped, and that a bad actor cannot exploit a coordinated slowdown to jump ahead in secret (Frontier Pause Verification).
- Convening. In the months after the essay, the Institute plans to organize conversations among policymakers, researchers, civil society, and other AI companies, and to publish the results — explicitly inviting voices outside AI companies into the deliberation.
People#
- Marina Favaro and Jack Clark co-authored When AI builds itself (editorial support from Santi Ruiz; visuals by Shan Carter, Romello Goodman, Nikki Makagiansar from data by Brian Calvert and Jun Shern Chan).
Connections#
-
Anthropic — parent organization
-
Recursive Self-Improvement — the subject of the Institute's flagship essay
-
Frontier Pause Verification — the Institute's concrete governance agenda
-
AI Accelerating AI Development — the internal evidence base the essay draws on
-
Responsible Scaling Policy Evaluations — the Institute's external-coordination work complements Anthropic's internal RSP brake
-
Safety Commitments That Cannot Bind the Actor Who States Them — why the option-to-pause framing costs nothing today: the commitment is conditioned on a verification regime the Institute is itself still building, while the part of Anthropic that can bind now (the RSP) is self-administered
-
Competitor-Indexed Safety Triggers: What the Revision Record Settles About Pause Postures, Acceleration Regret and Benchmark Perimeters — the revision record behind the option-to-pause posture: RSP 3.0's "Anthropic in the lead" delay clause and the Institute's verifiable-multilateral condition index both live pause commitments to competitive position
Open Questions#
- What concrete verification mechanisms will the Institute prototype, and on what timeline relative to the RSI trend it warns about?
Resolved Questions#
- How does the Institute's policy posture (favoring an option to pause) interact with Anthropic's commercial incentive to ship frontier models? The essay acknowledges the competitive/geopolitical pressure but doesn't resolve it. Partially answered (2026-08-19): Safety Commitments That Cannot Bind the Actor Who States Them settles the shape of the interaction without settling motive. The pause commitment is conditioned on a verifiable multilateral regime that does not exist and that the Institute is itself still building, so it imposes no present cost; the part of Anthropic that can bind today is the RSP, which is self-administered, bound once at a moment of its own choosing (Mythos Preview withheld until Fable 5's safeguards existed), bent toward shipping at its two closest calls (the Opus 5 CB-2 determination, the dropped rule-out suite), and has published a forecast of crossing CB-2 before its own recommended security bar exists. The load-bearing front-runner premise is itself disputed in the corpus by Domestic Frontier Pacing's ~1-year catch-up estimate. Still open, and unanswerable from this corpus: whether the conditional shape is chosen because the argument is right or because it is convenient — every observation above is Anthropic assessing Anthropic. Answered 2026-10-01: Competitor-Indexed Safety Triggers: What the Revision Record Settles About Pause Postures, Acceleration Regret and Benchmark Perimeters — the revision record settles the interaction without needing motive: RSP 3.0 removed the unconditional pretraining pause and conditions unilateral delay on "Anthropic in the lead" with "clear evidence that no other competitor will soon develop such a model", the Institute's multilateral pause needs every rival to pause verifiably, and together the two triggers cover every competitive position except the one where pausing would cost Anthropic ground. The commercial incentive is the trigger, not a counterweight to it (Zhu census App. F, ANT-2-046; the earlier "act promptly" substitution was silent in 1.0→2.0). Motive stays unsettleable because the right and the convenient front-runner argument produce the same clause, and it is not what was asked.
Sources#
- When AI builds itself — Anthropic Institute, When AI builds itself (Marina Favaro & Jack Clark, June 2026)
Cited by 14
- Safety Commitments That Cannot Bind the Actor Who States Them×3
The Institute's position is not "we will pause." It is that the world should have the option, and…
- Anthropic×2
Anthropic Institute — Anthropic's policy/governance research arm; published When AI builds itself
- Competitor-Indexed Safety Triggers: What the Revision Record Settles About Pause Postures, Acceleration Regret and Benchmark Perimeters×2
Anthropic Institute: How does the Institute's policy posture (favoring an option to pause) interact…
- Recursive Self-Improvement×2
Recursive self-improvement (RSI) is the point at which an AI system can fully autonomously design…
- AI Accelerating AI Development
The empirical half of the Anthropic Institute's When AI builds itself — the previously-unreported…
- AI R&D Autonomy Evaluation (AECI)
This is the capability-side gate on Recursive Self Improvement: AECI and the substitution threshold…
- Domestic Frontier Pacing
Anthropic Institute — the counterpart agenda: building multilateral verification infrastructure,…
- Frontier Pause Verification
The governance response in When AI builds itself: if the RSI trajectory holds, the world should at…
- LLM-Driven Vulnerability Research
Update (2026-06-07): the Anthropic Institute essay When AI builds itself quantifies Glasswing's…
- METR
METR (Model Evaluation & Threat Research) is an independent organization that evaluates frontier-AI…
- Entities — People, Orgs, Tools & Projects
Anthropic Institute — Anthropic's policy/governance research arm; published When AI builds itself…
- Mythos Model
The Anthropic Institute essay (June 2026) attaches concrete numbers to Mythos Preview as the model…
- Open Questions Backlog
Anthropic Institute (120d) — What concrete verification mechanisms will the Institute prototype,…
- RSI Autonomy Levels (B0–L5)
The corresponding definition: RSI is the capability of a system to "autonomously transform acquired…
Related articles
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- METR
Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Elon Musk
Founder of Tesla, SpaceX and xAI, and the corpus's clearest case of a reversed AI-risk position: 2015 'we'll be pet lab…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
