H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Wes Gurnee

Anthropic 可解釋性研究團隊的研究員;共同第一作者,也是 Jacobian lens 的共同構想者。他提出可言說表徵與意識存取之間的連結,並主導此方法的開發

Article metadata
Publication details
Published:July 11, 2026
Filed:Entity
Domain:Entities
Tags:EntityPersonAnthropicInterpretability Researcher
Reading:2 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Wes Gurnee 插圖

資料來源#

摘要#

人物。 Anthropic 可解釋性研究團隊的研究員。與 Nicholas Sofroniew 共同擔任 Verbalizable Representations Form a Global Workspace in Language Models(Transformer Circuits,2026 年 7 月)的第一作者,並與 Jack Lindsey 共同提出 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的猜想。

貢獻#

根據論文的作者貢獻章節:

  • 與 Jack Lindsey 共同構想 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的連結
  • 與 Mateusz Piotrowski 共同開發首個實作
  • 主導後續開發與改良,包括方法變體,以及與 logit lens 和 tuned lens 的比較
  • 執行早期實驗,證明此 lens 能呈現模型在內部推理中使用的概念;論文其餘部分皆建基於這項成果

論文本身也引用了他較早期的研究:workspace 實驗中反覆使用的字元計數任務(模型默默追蹤目前行寬)出自 Gurnee 等人的研究。

相關項目#

資料來源#

§ end
Cited by 4
Related articles
  • Jack Lindsey

    Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobia…

  • The Assistant Persona in the Workspace

    Post-training installs the Assistant's point of view *into* a workspace that already exists in the base model: safety a…

  • Chain-of-Thought Monitorability

    Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…

  • Agentic Misalignment (AM)

    Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…

  • Interference Weights

    Large virtual weights that are irrelevant or actively harmful to a model's behaviour — the residue of weight superposit…