資料來源#
摘要#
人物。 Anthropic 可解釋性團隊的研究員,也是 Verbalizable Representations Form a Global Workspace in Language Models(Transformer Circuits,2026 年 7 月)的通訊作者。他與 Wes Gurnee 共同構思了 Jacobian Lens (J-lens),以及將可言說表徵與意識存取連結起來的猜想。
貢獻#
根據論文的作者貢獻章節:
- 與 Wes Gurnee 共同構思 Jacobian Lens (J-lens) 方法,以及可言說性↔意識存取之間的關聯
- 執行早期實驗,研究模型能否直接調節自己的 J-space(依照指示在心中保持某個概念),以及訓練後對透鏡讀出的影響——這些結果後來成為工作空間中的助理人格
- 與 Nicholas Sofroniew 一同提出將 J-space 與全域工作空間理論連結的實驗——這一步將可解釋性讀出轉化為關於模型認知功能組織的主張
論文也引用了他先前關於 transcoder 與歸因的研究(透過 J-lens 重新檢視的算術特徵來自 Lindsey 等人的研究)。
相關連結#
- Jacobian Lens (J-lens) — 共同發起者
- 語言模型中的全域工作空間(J-space) — 通訊作者;提出工作空間的詮釋框架
- 工作空間中的助理人格 — 執行訓練後差異比較實驗
- Wes Gurnee — 共同發起者與共同第一作者
- Anthropic — 可解釋性團隊
資料來源#
Cited by 4
- Jacobian Lens (J-lens)×2
An interpretability technique from Anthropic's interpretability team (Wes Gurnee, Jack Lindsey et…
- Wes Gurnee×2
Entity. Researcher on Anthropic's interpretability team. Co-first author (with Nicholas Sofroniew)…
- Entities — People, Orgs, Tools & Projects
Jack Lindsey — Anthropic interpretability researcher; corresponding author of the global-workspace…
- Self-Report as a Safety Signal
Jack Lindsey — the intention probe is taken verbatim from Lindsey 2025's introspective-awareness…
Related articles
- Wes Gurnee
Anthropic interpretability researcher; co-first author and co-originator of the Jacobian lens, who conceived the connec…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Internal Signatures of Misalignment
The J-lens reads strategic and deceptive cognition that never reaches the output: `leverage`/`blackmail` while reading…
- The Global Workspace in Language Models (J-space)
Anthropic's July 2026 finding that LLMs maintain a small privileged set of verbalizable representations — the J-space —…
- Access-Consciousness Indicators in AI
The consciousness question the workspace paper deliberately does and doesn't answer: it tests *functional* indicator pr…
