Sources#
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
- Toward Native Multimodal Modeling: A Roadmap
Summary#
The core architectural move in Interaction Models: instead of consuming a complete user turn and emitting a complete response, input and output are treated as continuous streams, processed and generated in ~200ms chunks ("micro-turns") that interleave. There are no artificial turn boundaries the model must adhere to — silence, overlap, and interruption all remain part of the model's context.
How it works#
- The model continuously interleaves: process 200ms of input → generate 200ms of output → process next 200ms of input → … across audio, video, and text.
- Human perception preserves concurrent input and output streams; the model sees a single interleaved token sequence (
input 0, output 0, input 1, output 1, …) that encodes the same timing. - Because timing is in the sequence, the model has a direct sense of elapsed time and can act during the user's turn, not only after it.
Contrast: turn-based models see an alternating token sequence with hard turn boundaries; real-time feel is faked by a harness that predicts those boundaries (VAD etc.) — see Turn-Based Interface Bottleneck.
Why 200ms#
200ms chunks are small enough for near-real-time concurrency of multiple input/output modalities. The cost: inference must do frequent small prefills and decodes, each under strict latency constraints — and existing LLM inference libraries aren't built for that (significant per-turn overhead).
Inference: streaming sessions#
TML's fix for the frequent-small-prefill problem:
- Client sends each 200ms chunk as a separate request.
- Inference server appends chunks into a persistent sequence in GPU memory — avoiding repeated memory reallocation and metadata recomputation.
- Upstreamed a version of this to SGLang.
- Plus latency-tuned kernels for bidirectional-serving shapes: e.g. gather+gemv for MoE kernels instead of the standard grouped gemm (citing prior work from PyTorch/
gpt-fastand Cursor's warp-decode).
Trainer-sampler alignment#
Bitwise trainer-sampler alignment is used for training stability and for debugging the system's components. Implemented via batch-invariant kernels with <5% e2e overhead. Two highlighted kernels:
- All-reduce / reduce-scatter — NVLS low-latency comm kernels, deterministic on Blackwell, bitwise-aligned across different parallelism strategies (Sequence Parallelism vs Tensor Parallelism).
- Attention — Split-KV normally causes inconsistent accumulation orders between decode and prefill; fixed by splitting consistently between decode and prefill (e.g. 4096 tokens at a time, left-aligned), keeping efficiency in both.
What it buys you#
Every interaction mode that needs a special-purpose harness today becomes a special case of model behavior — and improves with model size and training data: proactive interjection, simultaneous speech, visual-cue reactions, time estimation. See Full-Duplex Interaction.
(Qualified 2026-09-21 by Peng et al., empirical. The mechanism removes the constraint on acting during the user's turn; it does not supply the reason to. Across seven configurations of five open full-duplex families — all of which can in principle speak mid-turn — a false claim or an imminent hazard moves speech onset by at most.06 from baseline, while being addressed and an actual pause move it by up to +.24 and +.73. Two of the five will not interrupt ongoing speech at all. And the two arms post-trained specifically for interactivity get quieter on content cues than their base checkpoints, so "improves with model size and training data" is, on the one axis anyone has measured, not yet true of proactive interjection. See Content-Driven Intervention.)
The primary-source form of the interleave: Moshi at 12.5 Hz (via the NMM roadmap, May 2026)#
TML describes the interleave as a design (input 0, output 0, input 1, output 1, … at ~200 ms) without publishing the frame rate or the per-frame token layout. An, Lu, Dong et al.'s survey (practitioner-opinion) supplies the one open model in its 43-model census that does: Moshi interleaves text and audio at the 12.5 Hz frame level, combining one text position and eight audio codebook positions per timestep — an 80 ms frame, a 1:8 text:audio token ratio, and a concrete answer to what "a chunk" contains.
Two things make it worth carrying beyond the arithmetic.
The interleave is a pretraining constraint, not only an inference schedule. The survey files this under modality-mixture scheduling, its name for the technique that becomes mandatory at early-fusion once per-modality loss heads disappear and the mixture in each batch directly sets the gradient direction. In the same paragraph: Moshi allocates half of all pretraining batches to text-only data as an explicit anti-forgetting buffer for language capability under the unified loss, and its RVQ semantic codebook carries loss weight α = 100 against the acoustic codebook's α = 1. So a model that streams at micro-turn granularity pays for it in the training mixture — a cost this page had no visibility into.
It sits beside the comparable numbers from other early-fusion models, which is what makes it legible as a choice rather than a constant: Transfusion fixes a 1:1 text-to-image token ratio with captions preceding their image 80% of the time; Chameleon fixes every image at 1,024 tokens from a 512×512 center crop regardless of source resolution, so the per-image gradient is constant. Moshi's 1:8 is the audio-streaming member of that family.
Caveats: this is a survey restating Moshi's own paper, not a measurement, and the survey's evaluation table lists Moshi Eval only as a 200 ms target latency benchmark — a target, not a result.
A second primary source on the same grid, with the turn boundary as a trained target (NVIDIA, September 2026)#
Moshi's 12.5 Hz interleave above arrives via a survey. NemotronLabs VoiceChat (NVIDIA, arXiv 2609.21967, 2026-09-18, empirical) is a primary source on the same frame grid, and it publishes the part Moshi's account does not: what the model is actually trained to predict at each frame.
The grid. A 600M cache-aware FastConformer encoder with causal depthwise-striding subsampling by a factor of eight emits one encoder state every 80 ms — the same 12.5 Hz frame as Moshi, reached from a streaming-ASR lineage rather than an RVQ codec. Self-attention uses a 70-frame left context and no right context; the convolutions are causal; no future audio is required for the current state. On the output side the TTS codec also runs at 12.5 Hz, one 31-level RVQ frame per 80 ms of waveform. The two ends of the system agree on the grid without sharing gradients.
The turn boundary becomes a frame-level target, and it is upweighted an order of magnitude. The agent-text channel is a padded timeline: BOS is placed at response onset, response subword tokens follow in consecutive frames, EOS is the stop target after a brief overlap with the next user turn, and everything else is padding. The paper states the consequence directly — "BOS and EOS therefore serve as frame-level turn-taking targets: BOS teaches when the model should begin responding, EOS teaches when it should stop, and padding teaches it to remain silent." The SFT loss weights make the priority explicit: 12.5 on begin-of-turn, 7.5 on end-of-turn, 5.0 on text content, 1.0 on padding. So the interleave is not only how the model is served — when to speak is a supervised target on the same grid as what to say, weighted above it.
Two augmentations that shape the grid, and one deliberate asymmetry. Turn-based training data does not supervise duplex behaviour, so SFT injects it: early interruption at p = 0.1 truncates a random agent turn mid-utterance and continues for eight further frames (640 ms) before EOS, so the next user turn overlaps agent speech; backchannel injection at p = 0.05 per sample loudness-matches recorded "uh-huh" audio into the user channel during agent speech so acknowledgements are not read as interruptions. And a text-channel delay: agent text targets are shifted two frames (160 ms) later during SFT, buying the model more user audio before it commits to each token — while the function channel is explicitly not shifted, "preserving the true temporal position of each tool call." A system running two output channels on one timeline can afford to delay one and not the other, which is a granularity of control the single-interleave picture above does not have.
(Counterweight, same source. §4 also ships a heuristic endpointer that forcibly injects BOS/EOS when the learned behaviour fails — so the frame-level turn-taking target is backstopped by a speech/silence activity rule, not trusted alone. Detail on Full-Duplex Interaction.)
Connections#
- Native Multimodal Modeling: Fusion Depth and I/O Duality — Moshi's 12.5 Hz / one-text-plus-eight-audio-codebooks interleave as the primary-source form of the micro-turn, and modality-mixture scheduling as the early-fusion training constraint it belongs to
- Interaction Models — parent concept
- Turn-Based Interface Bottleneck — what this replaces
- Encoder-Free Early Fusion — the complementary "minimal pre-processing" choice that makes streaming feasible
- Interaction / Background Model Split — micro-turns keep the interaction model present; deep reasoning is delegated out so it doesn't stall the stream
- Full-Duplex Interaction — the capabilities unlocked
- Content-Driven Intervention — the mechanism supplies the opportunity, not the reason: proactive interjection measured across five open duplex families and absent
- The Bitter Lesson — "no turn boundaries → interaction modes become scalable model behavior" is a direct application
- Context Window Smart Zone — continuous A/V at 200ms granularity accumulates context fast; the open long-session problem
- Interactivity Benchmarks — turn-taking latency (0.40s) is the direct, measured payoff of removing turn boundaries
- TML-Interaction-Small — the model built on this mechanism (200ms interleaved input/output chunks)
- Live-Path Minimalism — the production serving counterpart one level up: GPT-Live's stateful streaming inference (persistent sessions, seamless instance handoff, compaction off the live path) solves the same continuous-inference problem, with chunking internals undisclosed where TML's streaming sessions are public and upstreamed
Sources#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed — Peng, Nuchged, Fu & Yao, arXiv 2609.19596, 2026-09-17 (
empirical): the qualifier on "improves with model size and training data". Full treatment on Content-Driven Intervention - NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA), arXiv 2609.21967, 2026-09-18 (
empirical, 19 pp): §2.1 the 80 ms encoder grid, zero right context and the BOS/EOS frame-level turn-taking targets; §2.3 the 12.5 Hz / 31-level RVQ codec frame; §3.2 the 12.5 / 7.5 / 5.0 / 1.0 agent-text loss weights; §A.3 early-interruption, backchannel injection and the two-frame text-channel delay with the function channel unshifted. A primary source, not a survey restatement; NVIDIA on its own system, single run. Full treatment on Full-Duplex Interaction - Toward Native Multimodal Modeling: A Roadmap — An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities; arXiv 2605.25343, 2026-05-25;
practitioner-opinion): §5.1.3 modality-mixture scheduling — Moshi's 12.5 Hz text/audio interleave (1 text + 8 audio codebook positions per timestep), the half-text-only anti-forgetting buffer, the α=100/α=1 codebook weighting, and the Transfusion/Chameleon mixture comparisons. A survey restating Moshi's own paper, not a measurement. Full treatment on Native Multimodal Modeling: Fusion Depth and I/O Duality
Cited by 14
- Full-Duplex Interaction×3
Turn-taking here is a trained frame-level behaviour (BOS/EOS as per-frame targets on the agent-text…
- Encoder-Free Early Fusion×2
Avoids the latency and complexity of large standalone encoders/decoders — important when you have…
- Interaction / Background Model Split×2
a time-aware interaction model that maintains real-time presence — perceiving and responding in a…
- Interaction Models×2
Time Aligned Micro Turns, Interaction Background Model Split, Encoder Free Early Fusion — the three…
- The Bitter Lesson×2
Time Aligned Micro Turns — remove artificial turn boundaries so interaction modes become scalable…
- TML-Interaction-Small×2
Interaction mechanism: Time Aligned Micro Turns — 200ms interleaved input/output chunks, no turn…
- Turn-Based Interface Bottleneck×2
The Bitter Lesson says these hand-crafted systems get outpaced by general capability growth → the…
- Content-Driven Intervention
Time Aligned Micro Turns — the mechanism that makes acting during the user's turn possible at all;…
- Cursor
Earlier Cursor engineering also shows up obliquely: warp-decode kernels are cited as prior art for…
- Interactivity Benchmarks
Time Aligned Micro Turns — turn-taking latency (0.40s) is the direct payoff of removing turn…
- Live-Path Minimalism
Time Aligned Micro Turns — TML's disclosed counterpart on the inference internals: persistent…
- Interaction & Multimodal
Time Aligned Micro Turns — The core interaction-model move: input/output as continuous streams in…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
Time Aligned Micro Turns — Moshi's 12.5 Hz text/audio interleave (one text position + eight audio…
- NVIDIA
Time Aligned Micro Turns — VoiceChat is the corpus's second primary source on the 80 ms duplex…
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Full-Duplex Interaction
Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
An, Lu, Dong et al. (Tencent Youtu + 5 universities, May 2026) formalize 'native' as two operator definitions — mid-fus…
