H
Howardism
Plate IIInteraction & Multimodal中文HOWARDISM

Interaction Models

Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via harness; interactivity scales with intelligence only if it's in the model — with OpenAI's GPT-Live (July 2026) independently shipping the audio slice of the same conclusions in production

Article metadata
Publication details
Published:May 13, 2026
Filed:Concept
Domain:Interaction & Multimodal
Reading:27 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Interaction Models

Sources#

Summary#

An interaction model is a model that handles interaction natively — continuously taking in audio, video, and text and thinking, responding, and acting in real time — rather than emulating real-time behavior through external scaffolding (VAD, turn-detection, dialog-management harnesses). Announced by Thinking Machines Lab as a research preview in May 2026, with a first model named TML-Interaction-Small.

The central bet: interactivity should scale alongside intelligence. If interaction is part of the model, scaling the model makes it both smarter and a better collaborator. If interaction lives in a hand-crafted harness, The Bitter Lesson says that harness gets outpaced by general capability growth.

The thesis in one line#

"For interactivity to scale with intelligence, it must be part of the model itself."

This is the harness-shrinkage argument (see Harness Shrinkage as Models Improve) applied to the interaction layer: VAD, turn-boundary prediction, dialog state machines — all "meaningfully less intelligent than the model itself" — should dissolve into model behavior. Once they do, capabilities that those harnesses couldn't support (proactive interjection, speak-while-listening, reaction to visual cues) become special cases of what the model does, and improve with scale.

What it unlocks (capabilities, not harness features)#

  • Seamless dialog management — model implicitly tracks whether the speaker is thinking, yielding, self-correcting, or inviting a response. No separate dialog-management component.
  • Verbal and visual interjections — model jumps in when context warrants ("interrupt when I say something wrong", "tell me when I've written a bug"), not only at end-of-turn.
  • Simultaneous speech — user and model speak concurrently (live translation).
  • Time-awareness — direct sense of elapsed time ("how long did it take me to run a mile?").
  • Simultaneous tool calls / search / generative UI — while listening and speaking, the model concurrently searches, browses, generates UI, weaves results back in.

See Full-Duplex Interaction and Interactivity Benchmarks for how these are demonstrated and measured.

Why turn-based interfaces are the bottleneck#

Today's models "experience reality in a single thread": they wait, blind, until the user finishes typing/speaking; then they generate, blind, until done or interrupted. This is a narrow channel for collaboration — it limits how much of a person's knowledge, intent, and judgement reaches the model, and how much of the model's work is legible. Analogy from the post: resolving a crucial disagreement over email instead of in person. Full treatment in Turn-Based Interface Bottleneck.

Architecture (three load-bearing ideas)#

  1. Time-Aligned Micro-Turns — input and output are continuous streams, processed/generated in 200ms chunks; no artificial turn boundaries. Silence, overlap, and interruption stay in context.
  2. Interaction / Background Model Split — a time-aware interaction model maintains real-time presence; an asynchronous background model handles sustained reasoning, tool use, longer-horizon work. The interaction model delegates with a rich context package (the full conversation, not a standalone query) and interleaves results back at a moment appropriate to what the user is doing. Net effect: "planning, tool-use, and agentic workflows of reasoning models at the response latency of non-thinking ones."
  3. Encoder-Free Early Fusion — minimal pre-processing instead of large standalone encoders/decoders: audio in as dMel + light embedding; images as 40×40 patches via hMLP; audio out via a flow head. All components co-trained from scratch with the transformer.

Plus engineering: streaming sessions for low-overhead frequent small prefills/decodes (upstreamed to SGLang); latency-tuned MoE kernels (gather+gemv instead of grouped gemm); bitwise trainer-sampler alignment via batch-invariant kernels (<5% overhead) for stability and debuggability.

Convergence: OpenAI ships the audio slice (July 2026)#

OpenAI's GPT-Live arrives at the same architectural conclusions from the opposite direction — production latency engineering rather than a scaling-research bet — and ships them at ChatGPT scale. The mapping against the three pillars: full-duplex model in control of the conversation, no turn detector (audio-only; the cross-modal generalization stays TML's); background delegation to frontier models (Interaction / Background Model Split, with GPT-5.5 in the background slot); micro-turn internals undisclosed — OpenAI's account is instead the serving side (Live-Path Minimalism): stateful streaming inference, seamless instance handoff, compaction off the live path. Two labs, two months apart, one framing the thesis ("for interactivity to scale with intelligence, it must be part of the model") and one shipping its strongest testable consequence (the turn detector dissolved — see Turn-Based Interface Bottleneck).

Third derivation, from a speech lab (September 2026)#

Hu et al. (NVIDIA, arXiv 2609.19334, 2026-09-16, empirical) arrive at pillar 2 — and only pillar 2 — from a third direction: streaming ASR and full-duplex speech research, with no scaling thesis and no serving rewrite behind it. Their stated motive is a capacity argument rather than a latency one ("audio tokens consume parameters and context budget that text-only LLMs can devote to … tool-call capabilities"), and their contribution is the piece both prior accounts withheld: a published delegation signal — a <tc bos>/<tc eos> control-token pair in the agent-text channel, with results returning by prefill-and-repeat (Interaction / Background Model Split has the mechanism and the three ways it diverges from TML's description).

The paper is also a census of the convergence, which is what makes it useful here beyond its own system. It names, as concurrent designs sharing the shape, KAME, MoshiRAG, this page's own Thinking Machines work, Qwen-audio-agent, GPT-Live, NVIDIA's Nemotron Voice Agent and LiveKit's EXA Deep Researcher — and DuplexSLA as the road not taken, putting tool calls in a dedicated channel inside a Moshi-style duplex model. So the split now has seven-odd instances and one named live alternative, which is the shape a design pattern has before anyone knows whether it is permanent.

What this source does not corroborate is the rest of the thesis. Its frontend is a duplex speech-to-text model plus a separate streaming TTS, audio-only, with no video and no cross-modal generalization; and it delegates rather than staying present, so it is evidence for the architecture and not for "interactivity scales with intelligence."

A fourth lab, and the first that does not describe a split (Google, September 2026)#

Google's Gemini 3.8 Live launch (Gemini Audio Team, 2026-09-15, vendor-claim) is the fourth frontier live-voice system in this corpus and the first whose public account contains no architecture at all — and the absence is structured in a way worth recording, because the post's two framings point in opposite directions.

The framing says one model. Google ships two named models, 3.8 Live and 3.8 Live Extended Thinking, the second "built for high-complexity tasks, with increased intelligence and multi-step reasoning," and says it "reasons and speaks simultaneously" while "maintaining an uninterrupted conversational flow." Every benchmark bar is labelled with a reasoning effort — High, Medium, Minimal — attached to the live model itself. That is the vocabulary of one model with a thinking dial, and it contrasts precisely with OpenAI's cards, which footnote a backend model name and effort (Astra at medium, Terra at low). Two vendors, two grammars: Google's effort setting is a property of the thing you talk to; OpenAI's is a property of the thing behind it.

The vocabulary says two. In the same post: "It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background," plus "early verbal cues like 'Let me check that…'" and "live progress narration to walk users through multi-step background tasks as they progress." Background execution, an acknowledgement filler and step narration are the three observable signatures of a delegating system — the hold phrase is NVIDIA's ~1 s filler by another name, and progress narration is the behaviour OpenAI markets the absence of. A single model that reasons while it speaks has no obvious use for any of them.

What this does and does not establish. It establishes that a fourth vendor presents the internalized branch — one live model, an effort dial on it, no backend named, no backend priced — which is the first time any party in this corpus has publicly framed a frontier voice product that way. It establishes nothing about the implementation: the post names no component, publishes no delegation signal, quotes no latency, and the EVA-Bench runs are footnoted as executed "on the Live API on Gemini Enterprise Agent Platform," a platform rather than a model. The DeepMind model card the post points at is not ingested and, per the scout, carries no architecture beyond "built on Gemini 3 Pro architecture," 128K context and a January 2025 cutoff. So the split question gains a framing and no disclosure, which is exactly the kind of evidence that is easy to over-read. Treated as one vendor's presentation choice, recorded as such in the open question below, and not counted as an instance of either branch.

The one thing it does settle is a smaller question. The delegating shape is not a universal public framing: three labs describe a split and one does not, so "everyone converged on the split" is already too strong as a statement about what vendors say, whatever their systems do.

Safety angle#

Real-time interaction stresses safety differently than turn-based exchange. TML's work focused on two axes:

  • Modality-appropriate refusals — TTS-generated refusal / over-refusal training data so spoken refusals are colloquial but no less firm.
  • Long-horizon robustness — automated red-teaming harness generating multi-turn refusal data, maintaining behavioral parity with the text model's refusals.

Limitations (per the post)#

  • Long sessions — continuous A/V accumulates context fast; streaming-session design handles short/medium well, very long sessions need careful context management (parallels Context Window Smart Zone).
  • Compute & connectivity — low-latency A/V streaming needs reliable connection; degrades badly without one.
  • Scale — TML-Interaction-Small is 276B MoE / 12B active; larger pretrained models too slow to serve in this regime today; larger models promised "later this year".
  • Background agents — agentic intelligence acknowledged as essential but under-explored relative to the real-time work.

A fifth derivation, from a survey with no product to sell (May 2026)#

Every derivation above belongs to a lab shipping a live-voice system. An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities, arXiv 2605.25343, practitioner-opinion) reach the thesis from the opposite direction — a 52-page architecture survey of 43 open multimodal models, with no interaction product in it anywhere — and their §8.4 conclusion is this page's thesis in architecture vocabulary: truly native interactive agents require "streaming by construction, not as a post-hoc wrapper around an autoregressive backbone." Post-hoc wrapper is the harness; by construction is "part of the model itself."

Two things this adds that the vendor derivations cannot.

It supplies the mechanism, not just the verdict. The survey's §5 argues that fusion depth forces a training signature: at mid-fusion differential learning rates become mandatory once gradients reach the encoder; at early-fusion z-loss and QK-Norm become preconditions (Chameleon diverges after ~20% of training without QK-Norm) and modality-mixture scheduling replaces differential LR. If that is right, "interaction in the model" is not a design preference one lab can adopt and another decline — it is a commitment that propagates all the way down into the optimizer. The bitter-lesson argument on this page has always been about outcomes; this is the first source in the corpus to argue it from training mechanics.

It dates the gap from outside the vendor set. In the same section the survey says stable, deployable, low-latency full-duplex systems with consistent quality across modalities are "still an industrial open problem," naming Moshi, ELLSA and FireRedChat as hints rather than solutions. That is a neutral May 2026 statement that the thing was not done, against which the July–September 2026 product assertions catalogued above can be read.

The limit worth stating. The survey's census is restricted to open-source models and technical reports with verified architecture, so TML-Interaction-Small, Inkling, GPT-Live and Gemini 3.8 Live are all absent — it is a derivation of the thesis that has never looked at an interaction model. Its one bridge is Moshi, and its only measured interaction numbers are benchmark targets (Moshi Eval's 200 ms, SoulX-Duplug-Eval's 240 ms), not results.

Connections#

  • Native Multimodal Modeling: Fusion Depth and I/O Duality — the fifth derivation, from an architecture survey with no interaction product: "streaming by construction, not as a post-hoc wrapper," plus the training-mechanics argument for why interaction-in-the-model propagates into the optimizer
  • Build for the Next Model — matching the interaction shape to current capability (Codex-web "too AGI-pilled" vs. Claude Code's question-asking local form) is build-for-the-next-model at the interaction layer
  • Software 3.0 — Karpathy frames interaction models as a step toward the 3.0 neural-computer
  • The Bitter Lesson — the principle the whole approach rests on: scaled general methods beat hand-engineered structure
  • Harness Shrinkage as Models Improve — same move, applied to the interaction harness (VAD/turn-detection dissolve into the model)
  • Turn-Based Interface Bottleneck — the problem being solved
  • Time-Aligned Micro-Turns, Interaction / Background Model Split, Encoder-Free Early Fusion — the three architectural pillars
  • Full-Duplex Interaction — the interaction modes this enables
  • Content-Driven Intervention — the "speak when needed" mode, measured on five open families and absent: interactivity training improves turn-taking precision without adding content proactivity
  • Interactivity Benchmarks — how intelligence + interactivity are measured jointly
  • Thinking Machines Lab, TML-Interaction-Small — who built it, what the model is
  • Context Window Smart Zone — the long-session limitation echoes the smart-zone problem
  • AI Employee Framing / Human-AI Accountability Redesign — both argue against optimizing purely for autonomy; interaction models are the interface-side answer to keeping humans in the loop
  • Design Concept Grilling — collaborative, real-time iteration over spec-and-walk-away; an interaction model is the substrate that would make grilling-style collaboration feel native
  • Agent Harness Engineering — the harness-vs-model division-of-labor question, here resolved firmly toward the model for the interaction layer
  • Claude Opus 4.7 — xhigh effort tier appears as a baseline config (GPT-realtime-2.0 minimal/xhigh)
  • HTML as the New Markdown — a sibling answer to "better human–AI collaboration" by the opposite mechanism: TML dissolves the real-time interface into the model, where Thariq Shihipar enriches the asynchronous artifact; both keep the human in the loop
  • Configurable Human Participation — the task-benchmark companion to this architecture answer: HAS-Bench factors human input into clarification/feedback/control channels with explicit timing, and finds participation must be designed (which channel, when, by whom) rather than defaulted — the outcomes-measured version of "keep the human in the loop"
  • GPT-Live — OpenAI's production convergence on the audio slice of the thesis
  • Live-Path Minimalism — the serving architecture an interaction model demands, from the lab that shipped one
  • NVIDIA — the third lab to derive the interaction/background split, and the first to publish the delegation signal
  • Gemini 3.8 Live — the fourth live-voice product, and the first whose public account describes no split at all

Open Questions#

  • Does the interaction/background split generalize, or is it a transitional artifact until a single model is both fast and deep enough? Partially answered (2026-08-04): it generalizes across labs — GPT-Live ships the same split in production (full-duplex voice model delegating to GPT-5.5), independently derived from latency engineering. Whether the split is permanent or transitional remains open; two implementations are evidence of convergence, not permanence. Extended (2026-09-21) — the generalizes half is now settled and the transitional half is sharper. Hu et al. (NVIDIA, empirical) is a third independent derivation from a third starting point (speech-ASR research, motivated by a parameter-capacity argument rather than latency or serving), and their related-work section names four more concurrent instances plus two cascaded production systems. Three labs, three rationales, one shape: generalization is no longer the open part. Two new facts bear on transitional. Against it: the split behaves as a capability interface — swapping the backend from Qwen2.5-7B to Qwen3-30B-A3B to Qwen3-235B-A22B moves EVA-Bench task completion 40.4 → 57.3 with the interaction half's weights untouched, which is a property a scaffold you intend to absorb would not be designed to have. For it: this implementation buys its agentic ability by giving up the split's own premise — the frontend goes silent during delegation instead of staying present — and the alternative branch is live and named, with DuplexSLA internalising tool calls inside a duplex model. The falsifier is unchanged and now concrete: a single duplex model that matches a delegating system's tool-call accuracy without going quiet. Extended again (2026-09-21) — the capability-interface evidence is now commercial, which is a different kind of durability. OpenAI shipped the split as a product on 2026-09-10 (vendor-claim): the front-end voice layer sells at $0.05/min, the backend is the developer's to choose and pay for, and the post says in as many words that it may be "a third-party model." A vendor that intended the split as a scaffold to absorb would not build a price list around it, publish a delegation call signature (session.commentary.append with a delegation id) for arbitrary developer-side agents, or invite a competitor's model into the back half. Against that, it is still only one more assertion on the transitional question — OpenAI is a party with an interest in selling the front half, and the vendor's own benchmark footnotes show the split doing the work (Astra at medium behind the intelligence cards, Terra at low behind the tool-calling cards) rather than the frontend closing the gap alone. Two facts that would settle it are still missing on both sides. Extended a third time (2026-09-21), and this one subtracts rather than adds. Google's Gemini 3.8 Live launch (vendor-claim) is a fourth frontier live-voice product and the first whose public account describes no split: two named live models distinguished by an effort setting attached to the live model itself, no backend named anywhere, no backend priced, and an Extended Thinking tier that "reasons and speaks simultaneously." Read as evidence for the transitional branch, that is weak and it is the first of its kind — a vendor presenting the internalized alternative as a shipping product rather than as DuplexSLA's research road-not-taken. Read carefully, it is a framing and not a disclosure: the same post says the model "executes tools and API calls in the background while continuing the conversation," sells an acknowledgement filler ("Let me check that…") and live progress narration through multi-step background tasks, and footnotes its agentic benchmark runs as executed on the Gemini Enterprise Agent Platform. Background execution plus a hold phrase plus step narration are the three surface signatures of delegation, so the post is consistent with an undisclosed split and with a single model that thinks while it talks, and discriminates between them nowhere. What it changes for this question is the shape of the falsifier rather than its content: it is now clear that vendor framing cannot answer this — the falsifier still has to be a measurement or a disclosure, and Google supplies neither. Extended a fourth time (2026-09-23), and this one is a measurement rather than a framing. NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities (NVIDIA, empirical) is the internalized alternative built and benchmarked by the same lab that published the delegating system two days earlier, which makes it the nearest thing to the counterfactual this question has ever had. The falsifier stated above — a single duplex model that matches a delegating system's tool-call accuracy without going quiet — is now half-tested and fails on both halves, in instructive directions. On accuracy the internalized model wins the routing column and loses the task: 82.5 tool-selection F1 against the delegating system's 71.7–74.6 tool accuracy, but 42.2 argument accuracy against 52.8–55.2 and 33.0 Pass@1 against 44.0–48.0 on Full-Duplex-Bench 3.0 (the two papers name the routing metric differently, F1 versus accuracy, so read the second and third columns). On going quiet it does not even draw level: during tool execution it speaks a pre-configured per-tool acknowledgement, pads its agent-text channel, and — beyond what the delegating frontend conceded — stops conditioning response generation on incoming audio, so barge-in is unavailable for the duration of the call. What this settles: internalizing tool calls does not recover the live path, and the capability the split is actually buying is argument grounding and multi-step composition, not tool selection. What it leaves open is the same thing as before — nobody has run both against the same backend latency distribution and asked a user which is better — and the internalized branch now carries its own published ceiling (no more than five tools per session, unreliable simultaneous multi-tool invocation). Standing count: three labs describe the split, one declines to describe anything, and the one lab that built both branches shipped the delegating one with the better Pass@1 and the internalized one with the open weights.
  • "Interactivity scales with intelligence" is asserted; the larger-model release later in 2026 is the test. Partially answered (2026-09-21) — one axis measured, and it runs the other way. Peng et al. (empirical) run seven configurations of five open full-duplex speech families and report no intelligence or task-competence score for any of them, so no capability-vs-interactivity correlation across families can be read out of this source; the prediction's actual test, a larger TML release, is untouched. What it does supply is the nearest available proxy, and it is negative. Both interactivity-aligned RL arms — Moshika-RL and PPlex-RL, post-trained on pause handling, turn-taking, backchannelling and interruption — become quieter on content, not more proactive: Moshika-RL drops its neutral onset.07 →.03 while holding direct-question onset.15 and raising silence onset.12 →.18, with false-fact and hazard onset falling to.01/.01, below its own baseline; PPlex-RL drops neutral.10 →.01 and sharpens its question/neutral contrast from ~3.4x to ~21x while false fact and hazard sit.02/.04. The useful distinction this forces on the original claim: interactivity in the turn-taking sense (responsiveness, precision about when the floor is yours) is trainable and demonstrably improves with targeted post-training, while interactivity in the self-selection sense (speaking because the content warrants it) does not come along for the ride and may regress. A single word in the assertion is doing two jobs. Full treatment on the Content-Driven Intervention page. Extended (2026-09-21) with the first within-vendor datum, and it shows the measurement instrument failing before the claim does. Google (vendor-claim) ships two tiers on one live architecture and scores both on one third-party composite: Gemini 3.8 Live 76.0 and Gemini 3.8 Live Extended Thinking 82.6 on the Artificial Analysis Speech to Speech Index. Taken at face value that is the claim's shape — more intelligence, higher interactivity score, same family, no harness change. Three things stop it being evidence. First, the composite compresses the effect it is supposed to show: the one component published separately, the τ-Voice agentic chart, moves 30.1 → 68.6 across the same two tiers, +38.5 points against the composite's +6.6, so the index is dominated by something that barely moves. Second, the composite and that component rank two models in opposite orders — 3.8 Live beats Gemini 3.1 Flash Live at High effort on the index (76.0 vs 71.5) and loses to it on τ-Voice (30.1 vs 37.7). Third, nothing in the corpus states the index's composition or weights, and the methodology footnote printed on the charts points at the vendor's own page rather than the evaluator's. So the datum the question has been waiting for arrives inside an instrument that cannot carry it. The sharpened form: intelligence and interactivity must be scored on separated components to test this at all, and the only vendor to ship the two-tier comparison published a single number instead. Full treatment on the Interactivity Benchmarks page.
  • Research grant announced for interactivity benchmarks — what becomes the FD-bench equivalent for video proactivity? Partially answered (2026-09-21) — the audio half of the question now has an answer; the video half still does not. Peng et al. (empirical, arXiv 2609.19596) is the first third-party, reproducible proactivity instrument in the corpus: context-matched monologues where only the trigger utterance varies, ten conditions derived from turn-allocation theory, inter-word pauses compressed to 0.12 s so opportunity cannot confound reason, a single onset scalar with a stated denominator, and stimuli plus evaluation code released. It is the shape the question is asking for — and it is audio-only, 40 English monologues over two synthetic TTS voices. It also narrows what a video equivalent would have to do, since its two load-bearing design moves are modality-independent: hold the context fixed and vary only the trigger, and control the opportunity to speak separately from the reason. RepCount-A, ProactiveVideoQA and Charades do neither — each supplies a standing instruction and then measures timing, which is the addressed case. Nothing here bears on whether the grant produces a video instrument, so the trigger event is unchanged.

Sources#

§ end
Cited by 30
Related articles
  • Interaction / Background Model Split

    Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…

  • Turn-Based Interface Bottleneck

    Why current AI interfaces limit collaboration: single-thread turn-taking is a bandwidth bottleneck; humans pushed out b…

  • Full-Duplex Interaction

    Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Interactivity Benchmarks

    FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…