H
Howardism
Plate IIEntities中文HOWARDISM

Anthropic Institute

Anthropic's policy/governance research arm; published *When AI builds itself* (Favaro & Clark, 2026) on recursive self-improvement; agenda includes building the verification systems a credible multilateral AI slowdown would require

Article metadata
Publication details
Published:June 7, 2026
Filed:Entity
Domain:Entities
Reading:5 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Anthropic Institute

Sources#

Summary#

The Anthropic Institute is Anthropic's research and policy arm focused on the societal and governance implications of frontier AI. It published When AI builds itself (June 2026) — this wiki's primary source on Recursive Self-Improvement — and has a stated agenda to build, in collaboration with others, the systems that a credible AI slowdown or pause would require (Frontier Pause Verification).

What it does#

  • Public-facing trajectory analysis. When AI builds itself combines public benchmarks (Task Time-Horizon Scaling) with previously-unreported internal Anthropic data (AI Accelerating AI Development) to argue AI is already accelerating AI development and to lay out three futures for RSI.
  • Coordination infrastructure. It plans to "conduct research — in collaboration with many others — and take actions to help build the systems that a credible slowdown or pause would require": verification that other developers have actually stopped, and that a bad actor cannot exploit a coordinated slowdown to jump ahead in secret (Frontier Pause Verification).
  • Convening. In the months after the essay, the Institute plans to organize conversations among policymakers, researchers, civil society, and other AI companies, and to publish the results — explicitly inviting voices outside AI companies into the deliberation.

People#

  • Marina Favaro and Jack Clark co-authored When AI builds itself (editorial support from Santi Ruiz; visuals by Shan Carter, Romello Goodman, Nikki Makagiansar from data by Brian Calvert and Jun Shern Chan).

Connections#

Open Questions#

  • What concrete verification mechanisms will the Institute prototype, and on what timeline relative to the RSI trend it warns about?

Resolved Questions#

  • How does the Institute's policy posture (favoring an option to pause) interact with Anthropic's commercial incentive to ship frontier models? The essay acknowledges the competitive/geopolitical pressure but doesn't resolve it. Partially answered (2026-08-19): Safety Commitments That Cannot Bind the Actor Who States Them settles the shape of the interaction without settling motive. The pause commitment is conditioned on a verifiable multilateral regime that does not exist and that the Institute is itself still building, so it imposes no present cost; the part of Anthropic that can bind today is the RSP, which is self-administered, bound once at a moment of its own choosing (Mythos Preview withheld until Fable 5's safeguards existed), bent toward shipping at its two closest calls (the Opus 5 CB-2 determination, the dropped rule-out suite), and has published a forecast of crossing CB-2 before its own recommended security bar exists. The load-bearing front-runner premise is itself disputed in the corpus by Domestic Frontier Pacing's ~1-year catch-up estimate. Still open, and unanswerable from this corpus: whether the conditional shape is chosen because the argument is right or because it is convenient — every observation above is Anthropic assessing Anthropic. Answered 2026-10-01: Competitor-Indexed Safety Triggers: What the Revision Record Settles About Pause Postures, Acceleration Regret and Benchmark Perimeters — the revision record settles the interaction without needing motive: RSP 3.0 removed the unconditional pretraining pause and conditions unilateral delay on "Anthropic in the lead" with "clear evidence that no other competitor will soon develop such a model", the Institute's multilateral pause needs every rival to pause verifiably, and together the two triggers cover every competitive position except the one where pausing would cost Anthropic ground. The commercial incentive is the trigger, not a counterweight to it (Zhu census App. F, ANT-2-046; the earlier "act promptly" substitution was silent in 1.0→2.0). Motive stays unsettleable because the right and the convenient front-runner argument produce the same clause, and it is not what was asked.

Sources#

  • When AI builds itself — Anthropic Institute, When AI builds itself (Marina Favaro & Jack Clark, June 2026)
§ end
Cited by 14
Related articles
  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • METR

    Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…

  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Elon Musk

    Founder of Tesla, SpaceX and xAI, and the corpus's clearest case of a reversed AI-risk position: 2015 'we'll be pet lab…

  • AI R&D Autonomy Evaluation (AECI)

    How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…