資料來源#
摘要#
Anthropic Institute 是 Anthropic 的研究與政策部門,專注於前沿 AI 對社會與治理的影響。該部門發表了 When AI builds itself(2026 年 6 月)——本 wiki 關於 Recursive Self-Improvement 的主要來源——並公布了議程,計畫與他人合作,建立可信的 AI 減速或暫停所需的系統(Frontier Pause Verification)。
工作內容#
- 面向公眾的發展趨勢分析。 When AI builds itself 結合公開基準測試(Task Time-Horizon Scaling)與先前未曾公開的 Anthropic 內部資料(AI Accelerating AI Development),主張 AI 已在加速 AI 開發,並提出 RSI 的三種未來發展情境。
- 協調基礎設施。 該部門計畫「與許多其他人合作進行研究,並採取行動,協助建立可信的減速或暫停所需的系統」:確認其他開發者確實已停止,以及不良行為者無法利用協調減速的機會暗中超前(Frontier Pause Verification)。
- 促進交流。 該篇文章發表後的幾個月內,Institute 計畫促成政策制定者、研究人員、公民社會與其他 AI 公司之間的交流,並發布交流成果;該部門也明確邀請 AI 公司以外的聲音參與討論。
人物#
- Marina Favaro 與 Jack Clark 共同撰寫了 When AI builds itself(Santi Ruiz 提供編輯支援;視覺設計由 Shan Carter、Romello Goodman、Nikki Makagiansar 負責,資料來自 Brian Calvert 與 Jun Shern Chan)。
相關連結#
-
Anthropic — 母組織
-
Recursive Self-Improvement — Institute 旗艦文章的主題
-
Frontier Pause Verification — Institute 具體的治理議程
-
AI Accelerating AI Development — 該篇文章引用的內部證據基礎
-
Responsible Scaling Policy Evaluations — Institute 的外部協調工作,補足 Anthropic 內部的 RSP 煞車機制
-
Safety Commitments That Cannot Bind the Actor Who States Them — 說明「暫停選項」的框架為何現階段毫無成本:承諾以 Institute 自身仍在建構的驗證制度為前提,而 Anthropic 現在能約束自身的部分(RSP)則由其自行管理
待解決的問題#
- Institute 的政策立場(傾向保留暫停的選項)如何與 Anthropic 推出前沿模型的商業誘因互動?該篇文章承認競爭與地緣政治壓力,卻未能解決這個問題。部分解答(2026-08-19):Safety Commitments That Cannot Bind the Actor Who States Them 確立了這種互動的形式,但沒有確定其動機。暫停承諾以一套可驗證的多邊制度為前提,但這套制度目前並不存在,且 Institute 自身仍在建構,因此目前不會帶來任何成本;Anthropic 今天能約束自身的部分是 RSP,由其自行管理,只在自己選擇的時機承諾一次(在 Fable 5 的防護措施就緒之前,保留 Mythos Preview),在最接近的兩次關鍵決策中都朝推出產品的方向調整(Opus 5 的 CB-2 判定,以及遭刪除的排除測試套件),而且還發布預測稱,在自身建議的安全門檻尚未建立前,就會跨越 CB-2。語料中的 Domestic Frontier Pacing 提出約一年的追趕時間估計,對這項關鍵的領先假設本身提出質疑。問題仍然開放,且無法從此語料得出答案:採取這種附帶條件的形式,究竟是因為論點正確,還是因為這樣做方便——上述每一項觀察都是 Anthropic 對 Anthropic 的評估。
- Institute 會試行哪些具體驗證機制?相較於其所警告的 RSI 趨勢,試行時程又會如何安排?
資料來源#
- When AI builds itself — Anthropic Institute, When AI builds itself (Marina Favaro & Jack Clark, June 2026)
Cited by 14
- Safety Commitments That Cannot Bind the Actor Who States Them×3
The Institute's position is not "we will pause." It is that the world should have the option, and…
- Anthropic×2
Anthropic Institute — Anthropic's policy/governance research arm; published When AI builds itself
- Competitor-Indexed Safety Triggers: What the Revision Record Settles About Pause Postures, Acceleration Regret and Benchmark Perimeters×2
Anthropic Institute: How does the Institute's policy posture (favoring an option to pause) interact…
- Recursive Self-Improvement×2
Recursive self-improvement (RSI) is the point at which an AI system can fully autonomously design…
- AI Accelerating AI Development
The empirical half of the Anthropic Institute's When AI builds itself — the previously-unreported…
- AI R&D Autonomy Evaluation (AECI)
This is the capability-side gate on Recursive Self Improvement: AECI and the substitution threshold…
- Domestic Frontier Pacing
Anthropic Institute — the counterpart agenda: building multilateral verification infrastructure,…
- Frontier Pause Verification
The governance response in When AI builds itself: if the RSI trajectory holds, the world should at…
- LLM-Driven Vulnerability Research
Update (2026-06-07): the Anthropic Institute essay When AI builds itself quantifies Glasswing's…
- METR
METR (Model Evaluation & Threat Research) is an independent organization that evaluates frontier-AI…
- Entities — People, Orgs, Tools & Projects
Anthropic Institute — Anthropic's policy/governance research arm; published When AI builds itself…
- Mythos Model
The Anthropic Institute essay (June 2026) attaches concrete numbers to Mythos Preview as the model…
- Open Questions Backlog
Anthropic Institute (120d) — What concrete verification mechanisms will the Institute prototype,…
- RSI Autonomy Levels (B0–L5)
The corresponding definition: RSI is the capability of a system to "autonomously transform acquired…
Related articles
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- METR
Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Elon Musk
Founder of Tesla, SpaceX and xAI, and the corpus's clearest case of a reversed AI-risk position: 2015 'we'll be pet lab…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
