H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

Anthropic Institute

Anthropic 的政策與治理研究部門;發表了探討遞迴自我改進的 *When AI builds itself*(Favaro & Clark, 2026);其議程包括建立可信多邊 AI 減速所需的驗證系統

Article metadata
Publication details
Published:June 7, 2026
Filed:Entity
Domain:Entities
Tags:EntityOrgAI PolicyGovernanceAnthropic
Reading:4 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Anthropic Institute 的插圖

資料來源#

摘要#

Anthropic Institute 是 Anthropic 的研究與政策部門,專注於前沿 AI 對社會與治理的影響。該部門發表了 When AI builds itself(2026 年 6 月)——本 wiki 關於 Recursive Self-Improvement 的主要來源——並公布了議程,計畫與他人合作,建立可信的 AI 減速或暫停所需的系統(Frontier Pause Verification)。

工作內容#

  • 面向公眾的發展趨勢分析。 When AI builds itself 結合公開基準測試(Task Time-Horizon Scaling)與先前未曾公開的 Anthropic 內部資料(AI Accelerating AI Development),主張 AI 已在加速 AI 開發,並提出 RSI 的三種未來發展情境。
  • 協調基礎設施。 該部門計畫「與許多其他人合作進行研究,並採取行動,協助建立可信的減速或暫停所需的系統」:確認其他開發者確實已停止,以及不良行為者無法利用協調減速的機會暗中超前(Frontier Pause Verification)。
  • 促進交流。 該篇文章發表後的幾個月內,Institute 計畫促成政策制定者、研究人員、公民社會與其他 AI 公司之間的交流,並發布交流成果;該部門也明確邀請 AI 公司以外的聲音參與討論。

人物#

  • Marina Favaro 與 Jack Clark 共同撰寫了 When AI builds itself(Santi Ruiz 提供編輯支援;視覺設計由 Shan Carter、Romello Goodman、Nikki Makagiansar 負責,資料來自 Brian Calvert 與 Jun Shern Chan)。

相關連結#

待解決的問題#

  • Institute 的政策立場(傾向保留暫停的選項)如何與 Anthropic 推出前沿模型的商業誘因互動?該篇文章承認競爭與地緣政治壓力,卻未能解決這個問題。部分解答(2026-08-19):Safety Commitments That Cannot Bind the Actor Who States Them 確立了這種互動的形式,但沒有確定其動機。暫停承諾以一套可驗證的多邊制度為前提,但這套制度目前並不存在,且 Institute 自身仍在建構,因此目前不會帶來任何成本;Anthropic 今天能約束自身的部分是 RSP,由其自行管理,只在自己選擇的時機承諾一次(在 Fable 5 的防護措施就緒之前,保留 Mythos Preview),在最接近的兩次關鍵決策中都朝推出產品的方向調整(Opus 5 的 CB-2 判定,以及遭刪除的排除測試套件),而且還發布預測稱,在自身建議的安全門檻尚未建立前,就會跨越 CB-2。語料中的 Domestic Frontier Pacing 提出約一年的追趕時間估計,對這項關鍵的領先假設本身提出質疑。問題仍然開放,且無法從此語料得出答案:採取這種附帶條件的形式,究竟是因為論點正確,還是因為這樣做方便——上述每一項觀察都是 Anthropic 對 Anthropic 的評估。
  • Institute 會試行哪些具體驗證機制?相較於其所警告的 RSI 趨勢,試行時程又會如何安排?

資料來源#

  • When AI builds itself — Anthropic Institute, When AI builds itself (Marina Favaro & Jack Clark, June 2026)
§ end
Cited by 14
Related articles
  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • METR

    Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on i…

  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Elon Musk

    Founder of Tesla, SpaceX and xAI, and the corpus's clearest case of a reversed AI-risk position: 2015 'we'll be pet lab…

  • AI R&D Autonomy Evaluation (AECI)

    How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…