資料來源#
模型介紹#
Thinking Machines Lab 的首款**互動模型,於 2026 年 5 月以研究預覽形式發布。它被定位為「第一個兼具強大智慧/指令遵循能力與**互動性的模型」。
- **架構:**276B 參數 MoE,12B 啟用參數。從頭訓練為互動模型(而非在回合制模型上加裝互動功能)。
- **模態:**連續輸入音訊、影片與文字;輸出文字與音訊。Encoder-Free Early Fusion(dMel 音訊嵌入、用於影格的 40×40-patch hMLP、用於音訊輸出的 flow head)、單一共享 transformer,所有元件皆從頭共同訓練。
- 互動機制:Time-Aligned Micro-Turns——以 200ms 交錯輸入/輸出片段運作,不設回合邊界。
- **推理:**將深度推理/工具使用/長期任務委派給非同步背景模型——參見 Interaction / Background Model Split。即使沒有背景代理,在智慧基準測試上的表現也具競爭力。
主要數據(2026 年 5 月)#
- 輪替發言延遲:0.40s(FD-bench v1,音訊)——在所有比較模型中最佳。
- FD-bench v1.5 平均分數:77.8,基準模型約為 39–54,包含 thinking-high 模型。
- FD-bench v3(音訊+工具):82.8% 回應品質 / 68.0% Pass@1(搭配背景代理)。
- Audio MultiChallenge APR:43.4%——勝過所有非思考型基準模型;只有 GPT-realtime-2.0 xhigh(48.5%)更高。
- 比較的基準模型:GPT-realtime-2.0(minimal/xhigh)、GPT-realtime-1.5、Gemini-3.1-flash-live-preview(minimal/high)、Qwen 3.5 Omni-plus-realtime。完整表格見 Interactivity Benchmarks。
限制(已知)#
- 長時間連續 A/V 工作階段會快速累積上下文——如何謹慎管理上下文仍是未解問題(與 Context Window Smart Zone 的問題相呼應)。
- 需要穩定的低延遲連線;缺乏連線時效能會大幅下降。
- 稱為「Small」,是因為較大的預訓練模型目前在此運作模式下推論速度太慢;承諾於 2026 年稍晚推出更大型模型。
提供方式#
研究預覽版將於「未來幾個月內」限量推出,並於「今年稍晚」擴大發布。歡迎透過 interaction@thinkingmachines.ai 提供回饋;研究補助申請已開放。
與 Inkling 的關聯(2026 年 7 月)#
Inkling-Small——TML 開放權重版本的預覽同系列模型——採用276B MoE、12B 啟用參數:尺寸與本模型完全相同,出自同一實驗室,時間相隔兩個月,輸入端也採用相同的 encoder-free dMel/hMLP 堆疊。TML 沒有明言兩者共用權重,但公告指出 Inkling 的設計目標是作為本模型委派任務的背景推理模型(Interaction / Background Model Split)——因此,預覽時承諾「2026 年稍晚推出更大型模型」,後來以開放權重系列的形式實現,而其中的小型成員與互動模型具有相同規模。
延伸關聯#
- Interaction Models——此模型所屬的模型類別
- Thinking Machines Lab——開發者
- Time-Aligned Micro-Turns / Encoder-Free Early Fusion / Interaction / Background Model Split——三項架構支柱
- Full-Duplex Interaction——此模型展現的互動模式
- Interactivity Benchmarks——完整基準測試表及其勝過的基準模型
- Claude Opus 4.7——兩者同屬 2026 年年中時期的前沿模型;4.7 的
xhigh努力程度層級對應 GPT-realtime 的 minimal/xhigh,此處用作基準設定 - Context Window Smart Zone——長時間工作階段的限制
- Gemma 4——兩個月後採用相同的 encoder-free 設計,出發點是記憶體限制而非延遲;是最接近 TML 架構主張的獨立複現
- Inkling——背景模型系列;其中的小型成員與本模型同為 276B/12B 規模
資料來源#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Inkling: Our Open-Weights Model——Inkling-Small 尺寸相符;Inkling 是為背景模型而設計(
vendor-claim)
Cited by 15
- Inkling×4
Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…
- Full-Duplex Interaction×3
Proactive interjection — "interrupt when I say something wrong"; the model jumps in mid-turn when…
- Interaction Models×3
An interaction model is a model that handles interaction natively — continuously taking in audio,…
- Thinking Machines Lab×3
Interaction Models (May 2026 research preview) — models that natively take in audio/video/text and…
- Encoder-Free Early Fusion×2
Tml Interaction Small — the model that implements this design (dMel audio, 40×40 hMLP patches, flow…
- Interaction / Background Model Split×2
At the split's introduction the background model was an unnamed capability. Inkling fills the slot:…
- Interactivity Benchmarks×2
The evaluation surface Thinking Machines Lab uses to argue Tml Interaction Small is "the first…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×2
Nothing from the interaction-model line appears at all: no Tml Interaction Small, no Inkling, no…
- Claude Opus 4.7
Tml Interaction Small — era-mate (mid-2026 frontier from a different lab); 4.7's xhigh effort tier…
- Content-Driven Intervention
Full Duplex Interaction lists proactive interjection — "interrupt when I say something wrong" — as…
- GPT-Live
Tml Interaction Small — the research-preview sibling: same architectural conclusions from Thinking…
- Kimi (Moonshot AI)
K3 ships a 401M MoonViT-V2 vision encoder at 2.8T scale. Encoder Free Early Fusion documents three…
- Entities — People, Orgs, Tools & Projects
Tml Interaction Small — TML's first interaction model: 276B MoE / 12B active, audio+video+text in /…
- Open Questions Backlog
Inkling ×3 (oldest 75d) — Inkling-Small's 276B/12B dimensions match Tml Interaction Small exactly.…
- Time-Aligned Micro-Turns
Tml Interaction Small — the model built on this mechanism (200ms interleaved input/output chunks)
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
An, Lu, Dong et al. (Tencent Youtu + 5 universities, May 2026) formalize 'native' as two operator definitions — mid-fus…
- Encoder-Free Early Fusion
Multimodal design with minimal pre-processing instead of large standalone encoders: TML co-trains dMel audio + 40×40-pa…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
