H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Thinking Machines Lab

AI 研究實驗室,推出互動模型(2026 年 5 月)與 Inkling 開放權重系列(2026 年 7 月,從零開始訓練的 975B/41B);Tinker 託管微調平台;主張 harness 會融入模型;使命:透過客製化,讓 AI 擴展人類的意志與判斷力

Article metadata
Publication details
Published:May 13, 2026
Filed:Entity
Domain:Entities
Tags:Type/entityAI Lab
Reading:4 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Thinking Machines Lab 插圖

資料來源#

這是什麼#

一家 AI 研究實驗室(以「Thinking Machines Lab: Connectionism」名義發布內容)。使命是:「打造能擴展人類意志與判斷力的 AI。」在這份 wiki 中,它最早以 Interaction Models 背後組織的身分出現——這是 2026 年 5 月的研究預覽,將即時人機協作重新定位為模型原生能力,而非 harness 的課題。到了 2026 年 7 月,其策略清楚呈現為三部分的堆疊:Tinker(託管微調平台——任何人都能客製化模型)、互動模型(協作介面),以及 Inkling(從零開始訓練、可供客製化的開放權重基礎模型系列)。

已推出的內容/提出的主張(本文所見)#

  • Interaction Models(2026 年 5 月研究預覽)——能原生接收音訊、影片與文字,並即時思考、回應及採取行動的模型。首款模型:TML-Interaction-Small(276B MoE,啟用 12B)。
  • Inkling(2026 年 7 月)——他們首款從零開始訓練、以完整權重發布的模型:975B/啟用 41B 的多模態 MoE、1M context、可控的思考量,並以 GB300 系統進行訓練,完成超過 30M 次 RL rollout。他們明確表示,Inkling 的定位不是最強模型,而是最適合在 Tinker 上微調的基礎模型;同步預覽的 Inkling-Small(276B/12B)具有與 TML-Interaction-Small 完全相同的規模。其設計目標是在互動模型系統中擔任背景推理模型(Interaction / Background Model Split)。
  • Tinker——Inkling 上線使用的託管微調平台(64K/256K context、首日即有 Together/Fireworks/Modal/Databricks/Baseten 等服務合作夥伴,以及 SGLang/vLLM/llama.cpp 支援)。發布展示中,Inkling 自行撰寫並執行 Tinker 微調工作。他們針對微調預測模型的研究(「訓練 LLM 預測世界事件」)也促成了 Trained Calibration 配方。
  • 立場:互動性應隨智慧一同擴展 → 因此必須內建於模型之中;他們引用 The Bitter Lesson,反對以 harness 建構即時系統(VAD、回合偵測)。
  • 工程足跡:已將串流工作階段功能上游貢獻給 SGLang;發表了關於克服 LLM 推論中的非決定性的研究(批次不變核心),該研究被引用於訓練器與取樣器的對齊;先前也發布過 On-Policy Distillation 文章。
  • 正在為互動性/人機協作基準設立研究補助計畫(詳情 TBA);互動模型將於「未來幾個月內」推出有限研究預覽,更廣泛的發布則預計在「今年稍晚」進行;並承諾於 2026 年稍後推出更大型模型。

關聯脈絡#

相關連結#

資料來源#

§ end
Cited by 15
Related articles
  • Interaction Models

    Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…

  • Interaction / Background Model Split

    Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…

  • Time-Aligned Micro-Turns

    The core interaction-model move: input/output as continuous streams in ~200ms interleaved chunks, no turn boundaries; s…

  • Claude Opus 4.7

    GA frontier model from Anthropic; direct upgrade to 4.6 at same price; literal instruction following, 1.0–1.35× tokeniz…

  • Full-Duplex Interaction

    Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…