H
Howardism
Plate IIInteraction & Multimodal中文HOWARDISM

Time-Aligned Micro-Turns

The core interaction-model move: input/output as continuous streams in ~200ms interleaved chunks, no turn boundaries; streaming-sessions inference (upstreamed to SGLang), latency-tuned MoE kernels, bitwise trainer-sampler alignment; Moshi and NVIDIA VoiceChat both land on an 80ms (12.5Hz) frame, and VoiceChat makes the turn boundary a frame-level BOS/EOS target weighted 12.5/7.5 against padding at 1.0

Article metadata
Publication details
Published:May 13, 2026
Filed:Concept
Domain:Interaction & Multimodal
Reading:11 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Time-Aligned Micro-Turns

Sources#

Summary#

The core architectural move in Interaction Models: instead of consuming a complete user turn and emitting a complete response, input and output are treated as continuous streams, processed and generated in ~200ms chunks ("micro-turns") that interleave. There are no artificial turn boundaries the model must adhere to — silence, overlap, and interruption all remain part of the model's context.

How it works#

  • The model continuously interleaves: process 200ms of input → generate 200ms of output → process next 200ms of input → … across audio, video, and text.
  • Human perception preserves concurrent input and output streams; the model sees a single interleaved token sequence (input 0, output 0, input 1, output 1, …) that encodes the same timing.
  • Because timing is in the sequence, the model has a direct sense of elapsed time and can act during the user's turn, not only after it.

Contrast: turn-based models see an alternating token sequence with hard turn boundaries; real-time feel is faked by a harness that predicts those boundaries (VAD etc.) — see Turn-Based Interface Bottleneck.

Why 200ms#

200ms chunks are small enough for near-real-time concurrency of multiple input/output modalities. The cost: inference must do frequent small prefills and decodes, each under strict latency constraints — and existing LLM inference libraries aren't built for that (significant per-turn overhead).

Inference: streaming sessions#

TML's fix for the frequent-small-prefill problem:

  • Client sends each 200ms chunk as a separate request.
  • Inference server appends chunks into a persistent sequence in GPU memory — avoiding repeated memory reallocation and metadata recomputation.
  • Upstreamed a version of this to SGLang.
  • Plus latency-tuned kernels for bidirectional-serving shapes: e.g. gather+gemv for MoE kernels instead of the standard grouped gemm (citing prior work from PyTorch/gpt-fast and Cursor's warp-decode).

Trainer-sampler alignment#

Bitwise trainer-sampler alignment is used for training stability and for debugging the system's components. Implemented via batch-invariant kernels with <5% e2e overhead. Two highlighted kernels:

  • All-reduce / reduce-scatter — NVLS low-latency comm kernels, deterministic on Blackwell, bitwise-aligned across different parallelism strategies (Sequence Parallelism vs Tensor Parallelism).
  • Attention — Split-KV normally causes inconsistent accumulation orders between decode and prefill; fixed by splitting consistently between decode and prefill (e.g. 4096 tokens at a time, left-aligned), keeping efficiency in both.

What it buys you#

Every interaction mode that needs a special-purpose harness today becomes a special case of model behavior — and improves with model size and training data: proactive interjection, simultaneous speech, visual-cue reactions, time estimation. See Full-Duplex Interaction.

(Qualified 2026-09-21 by Peng et al., empirical. The mechanism removes the constraint on acting during the user's turn; it does not supply the reason to. Across seven configurations of five open full-duplex families — all of which can in principle speak mid-turn — a false claim or an imminent hazard moves speech onset by at most.06 from baseline, while being addressed and an actual pause move it by up to +.24 and +.73. Two of the five will not interrupt ongoing speech at all. And the two arms post-trained specifically for interactivity get quieter on content cues than their base checkpoints, so "improves with model size and training data" is, on the one axis anyone has measured, not yet true of proactive interjection. See Content-Driven Intervention.)

The primary-source form of the interleave: Moshi at 12.5 Hz (via the NMM roadmap, May 2026)#

TML describes the interleave as a design (input 0, output 0, input 1, output 1, … at ~200 ms) without publishing the frame rate or the per-frame token layout. An, Lu, Dong et al.'s survey (practitioner-opinion) supplies the one open model in its 43-model census that does: Moshi interleaves text and audio at the 12.5 Hz frame level, combining one text position and eight audio codebook positions per timestep — an 80 ms frame, a 1:8 text:audio token ratio, and a concrete answer to what "a chunk" contains.

Two things make it worth carrying beyond the arithmetic.

The interleave is a pretraining constraint, not only an inference schedule. The survey files this under modality-mixture scheduling, its name for the technique that becomes mandatory at early-fusion once per-modality loss heads disappear and the mixture in each batch directly sets the gradient direction. In the same paragraph: Moshi allocates half of all pretraining batches to text-only data as an explicit anti-forgetting buffer for language capability under the unified loss, and its RVQ semantic codebook carries loss weight α = 100 against the acoustic codebook's α = 1. So a model that streams at micro-turn granularity pays for it in the training mixture — a cost this page had no visibility into.

It sits beside the comparable numbers from other early-fusion models, which is what makes it legible as a choice rather than a constant: Transfusion fixes a 1:1 text-to-image token ratio with captions preceding their image 80% of the time; Chameleon fixes every image at 1,024 tokens from a 512×512 center crop regardless of source resolution, so the per-image gradient is constant. Moshi's 1:8 is the audio-streaming member of that family.

Caveats: this is a survey restating Moshi's own paper, not a measurement, and the survey's evaluation table lists Moshi Eval only as a 200 ms target latency benchmark — a target, not a result.

A second primary source on the same grid, with the turn boundary as a trained target (NVIDIA, September 2026)#

Moshi's 12.5 Hz interleave above arrives via a survey. NemotronLabs VoiceChat (NVIDIA, arXiv 2609.21967, 2026-09-18, empirical) is a primary source on the same frame grid, and it publishes the part Moshi's account does not: what the model is actually trained to predict at each frame.

The grid. A 600M cache-aware FastConformer encoder with causal depthwise-striding subsampling by a factor of eight emits one encoder state every 80 ms — the same 12.5 Hz frame as Moshi, reached from a streaming-ASR lineage rather than an RVQ codec. Self-attention uses a 70-frame left context and no right context; the convolutions are causal; no future audio is required for the current state. On the output side the TTS codec also runs at 12.5 Hz, one 31-level RVQ frame per 80 ms of waveform. The two ends of the system agree on the grid without sharing gradients.

The turn boundary becomes a frame-level target, and it is upweighted an order of magnitude. The agent-text channel is a padded timeline: BOS is placed at response onset, response subword tokens follow in consecutive frames, EOS is the stop target after a brief overlap with the next user turn, and everything else is padding. The paper states the consequence directly — "BOS and EOS therefore serve as frame-level turn-taking targets: BOS teaches when the model should begin responding, EOS teaches when it should stop, and padding teaches it to remain silent." The SFT loss weights make the priority explicit: 12.5 on begin-of-turn, 7.5 on end-of-turn, 5.0 on text content, 1.0 on padding. So the interleave is not only how the model is served — when to speak is a supervised target on the same grid as what to say, weighted above it.

Two augmentations that shape the grid, and one deliberate asymmetry. Turn-based training data does not supervise duplex behaviour, so SFT injects it: early interruption at p = 0.1 truncates a random agent turn mid-utterance and continues for eight further frames (640 ms) before EOS, so the next user turn overlaps agent speech; backchannel injection at p = 0.05 per sample loudness-matches recorded "uh-huh" audio into the user channel during agent speech so acknowledgements are not read as interruptions. And a text-channel delay: agent text targets are shifted two frames (160 ms) later during SFT, buying the model more user audio before it commits to each token — while the function channel is explicitly not shifted, "preserving the true temporal position of each tool call." A system running two output channels on one timeline can afford to delay one and not the other, which is a granularity of control the single-interleave picture above does not have.

(Counterweight, same source. §4 also ships a heuristic endpointer that forcibly injects BOS/EOS when the learned behaviour fails — so the frame-level turn-taking target is backstopped by a speech/silence activity rule, not trusted alone. Detail on Full-Duplex Interaction.)

Connections#

  • Native Multimodal Modeling: Fusion Depth and I/O Duality — Moshi's 12.5 Hz / one-text-plus-eight-audio-codebooks interleave as the primary-source form of the micro-turn, and modality-mixture scheduling as the early-fusion training constraint it belongs to
  • Interaction Models — parent concept
  • Turn-Based Interface Bottleneck — what this replaces
  • Encoder-Free Early Fusion — the complementary "minimal pre-processing" choice that makes streaming feasible
  • Interaction / Background Model Split — micro-turns keep the interaction model present; deep reasoning is delegated out so it doesn't stall the stream
  • Full-Duplex Interaction — the capabilities unlocked
  • Content-Driven Intervention — the mechanism supplies the opportunity, not the reason: proactive interjection measured across five open duplex families and absent
  • The Bitter Lesson — "no turn boundaries → interaction modes become scalable model behavior" is a direct application
  • Context Window Smart Zone — continuous A/V at 200ms granularity accumulates context fast; the open long-session problem
  • Interactivity Benchmarks — turn-taking latency (0.40s) is the direct, measured payoff of removing turn boundaries
  • TML-Interaction-Small — the model built on this mechanism (200ms interleaved input/output chunks)
  • Live-Path Minimalism — the production serving counterpart one level up: GPT-Live's stateful streaming inference (persistent sessions, seamless instance handoff, compaction off the live path) solves the same continuous-inference problem, with chunking internals undisclosed where TML's streaming sessions are public and upstreamed

Sources#

§ end
Cited by 14
  • Full-Duplex Interaction×3

    Turn-taking here is a trained frame-level behaviour (BOS/EOS as per-frame targets on the agent-text…

  • Encoder-Free Early Fusion×2

    Avoids the latency and complexity of large standalone encoders/decoders — important when you have…

  • Interaction / Background Model Split×2

    a time-aware interaction model that maintains real-time presence — perceiving and responding in a…

  • Interaction Models×2

    Time Aligned Micro Turns, Interaction Background Model Split, Encoder Free Early Fusion — the three…

  • The Bitter Lesson×2

    Time Aligned Micro Turns — remove artificial turn boundaries so interaction modes become scalable…

  • TML-Interaction-Small×2

    Interaction mechanism: Time Aligned Micro Turns — 200ms interleaved input/output chunks, no turn…

  • Turn-Based Interface Bottleneck×2

    The Bitter Lesson says these hand-crafted systems get outpaced by general capability growth → the…

  • Content-Driven Intervention

    Time Aligned Micro Turns — the mechanism that makes acting during the user's turn possible at all;…

  • Cursor

    Earlier Cursor engineering also shows up obliquely: warp-decode kernels are cited as prior art for…

  • Interactivity Benchmarks

    Time Aligned Micro Turns — turn-taking latency (0.40s) is the direct payoff of removing turn…

  • Live-Path Minimalism

    Time Aligned Micro Turns — TML's disclosed counterpart on the inference internals: persistent…

  • Interaction & Multimodal

    Time Aligned Micro Turns — The core interaction-model move: input/output as continuous streams in…

  • Native Multimodal Modeling: Fusion Depth and I/O Duality

    Time Aligned Micro Turns — Moshi's 12.5 Hz text/audio interleave (one text position + eight audio…

  • NVIDIA

    Time Aligned Micro Turns — VoiceChat is the corpus's second primary source on the 80 ms duplex…

Related articles