H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

TML-Interaction-Small

TML 首個互動模型:276B MoE / 12B 啟用參數 輸入音訊+影片+文字 / 輸出文字+音訊、200ms 微回合、非同步背景代理; 所有模型中最佳的輪替發言延遲;2026 年 5 月研究預覽——也是 2026 年 7 月 Inkling-Small 的精確形貌

Article metadata
Publication details
Published:May 13, 2026
Filed:Entity
Domain:Entities
Tags:Type/entityLLM Model
Reading:4 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

TML-Interaction-Small 插圖

資料來源#

模型介紹#

Thinking Machines Lab 的首款**互動模型,於 2026 年 5 月以研究預覽形式發布。它被定位為「第一個兼具強大智慧/指令遵循能力與**互動性的模型」。

  • **架構:**276B 參數 MoE,12B 啟用參數。從頭訓練為互動模型(而非在回合制模型上加裝互動功能)。
  • **模態:**連續輸入音訊、影片與文字;輸出文字與音訊。Encoder-Free Early Fusion(dMel 音訊嵌入、用於影格的 40×40-patch hMLP、用於音訊輸出的 flow head)、單一共享 transformer,所有元件皆從頭共同訓練。
  • 互動機制:Time-Aligned Micro-Turns——以 200ms 交錯輸入/輸出片段運作,不設回合邊界。
  • **推理:**將深度推理/工具使用/長期任務委派給非同步背景模型——參見 Interaction / Background Model Split。即使沒有背景代理,在智慧基準測試上的表現也具競爭力。

主要數據(2026 年 5 月)#

  • 輪替發言延遲:0.40s(FD-bench v1,音訊)——在所有比較模型中最佳。
  • FD-bench v1.5 平均分數:77.8,基準模型約為 39–54,包含 thinking-high 模型。
  • FD-bench v3(音訊+工具):82.8% 回應品質 / 68.0% Pass@1(搭配背景代理)。
  • Audio MultiChallenge APR:43.4%——勝過所有非思考型基準模型;只有 GPT-realtime-2.0 xhigh(48.5%)更高。
  • 比較的基準模型:GPT-realtime-2.0(minimal/xhigh)、GPT-realtime-1.5、Gemini-3.1-flash-live-preview(minimal/high)、Qwen 3.5 Omni-plus-realtime。完整表格見 Interactivity Benchmarks。

限制(已知)#

  • 長時間連續 A/V 工作階段會快速累積上下文——如何謹慎管理上下文仍是未解問題(與 Context Window Smart Zone 的問題相呼應)。
  • 需要穩定的低延遲連線;缺乏連線時效能會大幅下降。
  • 稱為「Small」,是因為較大的預訓練模型目前在此運作模式下推論速度太慢;承諾於 2026 年稍晚推出更大型模型。

提供方式#

研究預覽版將於「未來幾個月內」限量推出,並於「今年稍晚」擴大發布。歡迎透過 interaction@thinkingmachines.ai 提供回饋;研究補助申請已開放。

與 Inkling 的關聯(2026 年 7 月)#

Inkling-Small——TML 開放權重版本的預覽同系列模型——採用276B MoE、12B 啟用參數:尺寸與本模型完全相同,出自同一實驗室,時間相隔兩個月,輸入端也採用相同的 encoder-free dMel/hMLP 堆疊。TML 沒有明言兩者共用權重,但公告指出 Inkling 的設計目標是作為本模型委派任務的背景推理模型(Interaction / Background Model Split)——因此,預覽時承諾「2026 年稍晚推出更大型模型」,後來以開放權重系列的形式實現,而其中的小型成員與互動模型具有相同規模。

延伸關聯#

資料來源#

§ end
Cited by 15
  • Inkling×4

    Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…

  • Full-Duplex Interaction×3

    Proactive interjection — "interrupt when I say something wrong"; the model jumps in mid-turn when…

  • Interaction Models×3

    An interaction model is a model that handles interaction natively — continuously taking in audio,…

  • Thinking Machines Lab×3

    Interaction Models (May 2026 research preview) — models that natively take in audio/video/text and…

  • Encoder-Free Early Fusion×2

    Tml Interaction Small — the model that implements this design (dMel audio, 40×40 hMLP patches, flow…

  • Interaction / Background Model Split×2

    At the split's introduction the background model was an unnamed capability. Inkling fills the slot:…

  • Interactivity Benchmarks×2

    The evaluation surface Thinking Machines Lab uses to argue Tml Interaction Small is "the first…

  • Native Multimodal Modeling: Fusion Depth and I/O Duality×2

    Nothing from the interaction-model line appears at all: no Tml Interaction Small, no Inkling, no…

  • Claude Opus 4.7

    Tml Interaction Small — era-mate (mid-2026 frontier from a different lab); 4.7's xhigh effort tier…

  • Content-Driven Intervention

    Full Duplex Interaction lists proactive interjection — "interrupt when I say something wrong" — as…

  • GPT-Live

    Tml Interaction Small — the research-preview sibling: same architectural conclusions from Thinking…

  • Kimi (Moonshot AI)

    K3 ships a 401M MoonViT-V2 vision encoder at 2.8T scale. Encoder Free Early Fusion documents three…

  • Entities — People, Orgs, Tools & Projects

    Tml Interaction Small — TML's first interaction model: 276B MoE / 12B active, audio+video+text in /…

  • Open Questions Backlog

    Inkling ×3 (oldest 75d) — Inkling-Small's 276B/12B dimensions match Tml Interaction Small exactly.…

  • Time-Aligned Micro-Turns

    Tml Interaction Small — the model built on this mechanism (200ms interleaved input/output chunks)

Related articles