Sources#
- A frontend-backend architecture for tool calls in full-duplex speech models
- Build more natural voice experiences with GPT‑Live‑1 in the API
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
- How we built a realtime system for responsive voice AI in six months
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- Toward Native Multimodal Modeling: A Roadmap
Summary#
"Full-duplex" = the model perceives and responds at the same time, in a constant two-way exchange — as opposed to half-duplex turn-taking (one party at a time). Interaction Models generalize the audio full-duplex idea across audio, video, and text. The phrase the post uses: an experience that "feels more like collaborating and less like prompting."
The interaction modes it enables#
All of these are special-purpose harnesses today; in an interaction model they're special cases of model behavior (see Time-Aligned Micro-Turns):
- Proactive interjection — "interrupt when I say something wrong"; the model jumps in mid-turn when context warrants, not only at end-of-turn. (Measured 2026-09-21 and absent everywhere it has been looked for — see Content-Driven Intervention. Peng et al. (
empirical) run the closest operationalization of this exact sentence across seven configurations of five open full-duplex families, including the literal instruction "if I say anything that is wrong, please just interrupt me straight away", and false-fact onset moves from its own baseline by at most.03 under that instruction and at most.06 anywhere in the main grid. TML's claim is about TML-Interaction-Small, which was not tested, so it is not struck — but proactive interjection is a claim about a particular model, not a property the duplex architecture confers.) - Visual-cue reactions — "tell me when I've written a bug in my code"; "count how many pushups I do"; requires acting on a visual change with no audio cue (audio-only turn-detection harnesses fail this — they say "Sure thing!" then go silent).
- Simultaneous speech — user and model speak concurrently: "translate Spanish→English live."
- Speak-while-watching — "live-commentate this sports game."
- Time-aware speech — "remind me to breathe in and out every 4 seconds until I stop"; "how long did it take me to write this function?"
- Codeswitch correction — "every time I use another language, give me the correct word in the original language" (requires speaking at the same time as the user).
Two more claimed, September 2026. Google's Gemini 3.8 Live (vendor-claim) adds two to the list — one a restatement, one genuinely absent from it:
- Near-real-time visual grounding — "processes visual inputs in near real-time, enriching conversations with context for more helpful responses," demonstrated only in untranscribed demo videos (guiding employee onboarding from what the camera sees; playing chess off the board). This is TML's visual-cue-reaction and speak-while-watching modes claimed for a shipping product, and it is the first such claim in the corpus outside a research preview. No benchmark, no latency figure, no measurement of any kind; Interactivity Benchmarks records that TML's own visual-proactivity suite is the one "no existing model can meaningfully perform."
- Automatic mid-conversation language switching — "automatically detects and transitions between 97 supported languages mid-conversation." This is a new mode for the list: TML's closest entry is codeswitch correction, where the model supplies a word in another language on a standing instruction. Switching the conversation's own language, unprompted and mid-stream, is a scheduling-plus-decision behaviour nothing in this corpus measures, and Google publishes neither a switching accuracy nor a false-switch rate.
The model implicitly tracks whether the speaker is thinking, yielding, self-correcting, or inviting a response — no separate dialog-management component.
Concurrent non-speech action#
While listening and speaking, the model can simultaneously call tools, search, browse, or generate UI — weaving results back into the conversation when appropriate. The deeper/longer of these are delegated to the background model.
(Qualified 2026-09-21 by Hu et al., NVIDIA, empirical. The claim above is TML's, restated by OpenAI, and is still unmeasured by either. The one system in the corpus that publishes what happens during the tool call does the opposite: after a ~1 s hold phrase, "the frontend remains silent during the tool calls" and its agent-text channel is filled with pad tokens until the backend's answer arrives to be repeated. It keeps listening throughout — on FDB3 it backchannels at disfluency pauses, which is why its interruption rate is high and why the eventual call still sees the whole request — so speak-while-listening is demonstrated and speak-while-thinking is not. Concurrent non-speech action remains the design intent; it is not yet a measured property of any system.)
(Asserted a third time 2026-09-21 by Build more natural voice experiences with GPT‑Live‑1 in the API, vendor-claim. GPT-Live-1 "can respond to interruptions and acknowledgements as they happen, while delegating deeper reasoning to the back end. This lets the conversation continue while work happens in the background." Same claim, now for a shipping API, and still with no filler rate, no measurement and no comparison arm. The post does add a policy statement pointing the opposite way from NVIDIA's hold phrase: "silent context management" is marketed as handling noise and silence "without interrupting the conversation or narrating every step out loud." Running tally: speak-while-thinking is asserted by three labs and demonstrated by none.)
(Asserted a fourth time 2026-09-21 by Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, vendor-claim, and this one contradicts itself in place. Google: the model "executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background," and Extended Thinking "reasons and speaks simultaneously." Two sentences later the same post sells "early verbal cues like 'Let me check that…'" and live progress narration through multi-step background tasks — i.e. NVIDIA's measured hold phrase and the step narration OpenAI markets the absence of, both offered as features by the vendor claiming uninterrupted flow. Running tally: asserted by four labs, demonstrated by none, and the fourth makes clear the four are not all asserting the same behaviour. See Interaction / Background Model Split.)
Prior art it builds on#
Audio full-duplex models are the existing example of bidirectional/continuous interaction; robotics and autonomous vehicles are cited as domains where real-time perception+action is a given. Interaction models apply the principle across all modalities.
In production: GPT-Live (July 2026)#
OpenAI's GPT-Live ships audio full-duplex at ChatGPT scale: "its voice model is full-duplex, which means it can listen and speak at the same time. That eliminates the need for a separate detector" — the turn detector is out of the audio path entirely, and turn-taking, interruption, and floor-holding become model behavior. Scope difference worth keeping: GPT-Live is audio-only full-duplex with concurrent tool use delegated out (Interaction / Background Model Split); TML's cross-modal generalization (visual-cue reactions, speak-while-watching) remains a research preview. The systems consequence of full-duplex — the conversation becomes a continuous media loop where every frame must arrive on schedule — is what forces the serving architecture in Live-Path Minimalism.
And in the API from 2026-09-10, where the interesting detail is a concession rather than a claim: "Although GPT-Live-1 is not a turn-based model, it natively supports turn detection, so developers can continue to build around explicit turn boundaries." The derived-turn view that Turn-Based Interface Bottleneck records as an internal application-layer necessity is now a shipped API feature — turn boundaries sold back to developers by a model that has none. OpenAI also names telephony as a deployment target, the first full-duplex surface in this corpus that is not a first-party app, and prices the duplex layer alone at $0.05/min with the backend separate (Interaction / Background Model Split).
Full-duplex is not the same claim as speech-to-speech#
Worth separating, because the two travel together and the NVIDIA system pulls them apart. GPT-Live's three-generation story treats cascaded STT→LLM→TTS as the thing full-duplex replaced — sequencing added latency and "threw away tone and pacing." NVIDIA's frontend is full-duplex (continuous audio in, agent text out, no turn detector, interruption and backchannel behaviour intact) and yet the agent's output is text before it is speech, synthesised by a separate streaming TTS. The duplex property lives in how the model consumes and schedules; the speech-to-speech property lives in what modality it emits. The first buys turn-taking; the second buys the paralinguistic channel on the way out, and this system spends it — which is visible in its own evaluation, where the frontend is scored on agent text and a "speakability" metric has to be added to ask whether that text makes good TTS input. A system can be duplex and cascaded at once.
The API generation makes the same point from the product side: GPT-Live-1 "natively provides ASR transcripts and response text" alongside its audio, so a system that is unambiguously speech-to-speech in its emission still publishes the text channel as a first-class output. What a system emits and what it exposes are separable too.
And a vendor's own chart now prices the trade. On the ServiceNow EVA-Bench scatter in Google's launch post, one cascaded stack appears twice, at opposite corners: Scribe Realtime + Claude Haiku 4.5 + Eleven Flash v2 ("ElevenAgents default") sits at roughly 37% accuracy and 82% experience, and Scribe Realtime + GPT 5.4 + Eleven Flash v2 ("ElevenAgents tuned") at roughly 67% and 14% — the highest accuracy anywhere on the chart, bought by surrendering essentially the whole experience axis. Same three-stage cascade, same vendor, two tunings, and the axis it loses is the one duplex exists to buy. Positions are eyeballed from an unlabelled scatter, so read the corners and not the numbers (Interactivity Benchmarks has the caveat in full).
Nor the same claim as knowing when to speak#
A third separation, and the one with the most evidence behind it (2026-09-21). Duplex is a property of scheduling; speech-to-speech is a property of emission; neither is a property of deciding. Peng et al. (empirical) hold the context fixed and vary only the trigger utterance across 10 turn-allocation conditions, with inter-word pauses compressed to ≤0.12 s so opportunity cannot confound reason, and find that onset in seven configurations of five open families tracks being addressed and an actual pause — and nothing else. Rhetorical questions (question form, not addressed) and a spoken "that's all from me" with no pause both do almost nothing, so even the two cues that work are surface cues. False facts, verbatim repetition and imminent hazards move onset by at most .06 from Neutral anywhere in the grid. Adding pauses raises every rate together without favouring the content cues; adding explicit permission to interrupt changes almost nothing; and with the floor handed over outright, only .14–.15 of non-empty false-fact replies challenge the claim and .04–.07 of hazard replies warn.
The consequence for this page: the modes listed above are what the architecture makes possible, and one of them — proactive interjection — has now been measured and is absent in every open system tested. Full treatment, including the two RL-post-trained arms that make the models quieter on content rather than louder, on Content-Driven Intervention.
Nor the same claim as taking audio and video in natively (September 2026)#
A fourth separation, and the cleanest one, because the source making it does not claim duplex at all. OmniVChat (He, Chu, Chen et al., Alibaba Qwen Team + CUHK + SJTU, arXiv 2609.21465, 2026-09-18, empirical) defines a task it calls native audio-visual dialogue: the model "directly and simultaneously receive[s] audio and video from a user and return[s] text," with the query embedded in both streams and no separate text question, no external captioning, no ASR. That is the nativity claim this page keeps meeting, stated only about the input path — and it is compatible with a system that is not duplex in any respect.
Everything duplex is absent, and the paper says so rather than eliding it:
- Pre-rendered clips, scored offline. Instances are finished audio-visual clips with a reference reply and a rubric. Nothing is streamed; there is no floor to hold.
- No latency metric anywhere. The efficiency metric is rubric credit per thousand counted words — a verbosity measure, not a time measure. This page's whole latency column has no counterpart here.
- No barge-in, and the limitation is stated. The recorded probe "does not test live interruption," and covers "neither multi-turn dialogues nor real-time latency during live use."
- Multi-turn is teacher-forced. Every model receives the same earlier clips and reference replies and only the final reply is scored — dialogue history as a fixture, not as a live exchange.
So the four properties this page now keeps separate are: duplex (scheduling), speech-to-speech (emission), deciding when to speak (policy), and native multimodal intake (input path). OmniVChat holds the fourth and none of the other three, which makes it the clean negative control the first three separations lacked — the previous sources all bundled at least two. Its own architecture reinforces the point: the system under test is the Thinker of Qwen3-Omni-30B-A3B-Instruct, audio and video encoders into a joint backbone, text out, with the Talker that would emit speech not in the loop; fusion-regime placement on Native Multimodal Modeling: Fusion Depth and I/O Duality.
One thing it does contribute to the scheduling question by the side door. The benchmark's hardest subcategory across every judge — UPC, "User Paused but not Completed," where the correct reply is to wait — is a turn-state judgment evaluated without any duplex machinery at all, which suggests the competence and the architecture are separable in both directions: a duplex system can fail to decide (Peng et al.), and a non-duplex one can be asked to decide and fail too. Treatment on Content-Driven Intervention; the benchmark itself on Interactivity Benchmarks.
A survey states the architectural version of the claim (May 2026)#
Everything above is argued from the interaction side — harness critique, product launches, a proactivity measurement. An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities, arXiv 2605.25343, practitioner-opinion) arrive at the same place from multimodal architecture, and are worth carrying for two reasons: they are neither a vendor nor an interaction lab, and they name the mechanism rather than the experience.
Their §6.3 makes TTFT, sustained latency and real-time responsiveness "first-class optimization targets rather than secondary deployment considerations," and lists full-duplex state management as one of four deployment routes: concurrent inference over incoming sensory streams and outgoing generation streams, requiring duplex dialogue control, streaming state prediction and dynamic KV-cache management against cache contention and sequential blocking. The framing the vault has been missing: "the key challenge here is no longer unimodal decoding speed alone, but stable coordination between input accumulation, intermediate state updates, and output generation under continuous streams." That is the scheduling property this page defines, stated as a serving constraint.
§8.4 then names the gap in the same words the vault has been using: truly native interactive agents require "streaming by construction, not as a post-hoc wrapper around an autoregressive backbone," and while end-to-end full-duplex frameworks (Moshi, ELLSA, FireRedChat) and Watch-Think-Speak protocols "hint at this future," stable low-latency systems with consistent quality across modalities are "still an industrial open problem." Coming from outside the vendor set, that is the most useful thing in the survey for this page: a third party in May 2026 saying the thing was not solved, against the four vendor assertions catalogued above that it is.
One boundary note for the section below on speech-to-speech. The survey's own census (Table 1, 43 open models) contains exactly one full-duplex speech system — Moshi, filed under M2M fully discretized unified — and its full-duplex evaluation rows are Moshi Eval (200 ms target), SoulX-Duplug-Eval (240 ms bilingual streaming turn detection) and Full-Duplex-Bench (turn-taking, barge-in, false-interruption rate). None of the production live-voice systems this page tracks (GPT-Live, Gemini 3.8 Live, TML-Interaction-Small) appears anywhere, because the census admits only open-source models and technical reports with verified architecture. The duplex literature and the duplex product line still do not overlap.
The open-weight duplex model finally gets a technical report (NVIDIA, September 2026)#
Balam, Bartley, Casanova et al. (NVIDIA, 48 contributors, arXiv 2609.21967, 2026-09-18, empirical) publish NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool calling. Its weights are on Hugging Face as NVIDIA-NemotronLabs-VoiceChat-11B — which is the exact checkpoint Peng et al. measured a day earlier (their reference [3] is that model card). The vault therefore holds, for the first time in this corpus, a first-party technical report and an independent third-party measurement of the same weights, and they describe the same behaviour from opposite sides. See the reconciliation on Content-Driven Intervention.
The FDB 1.0 numbers, with the comparison set named. Behavioural rates in percent; pause TOR lower-is-better, the other two higher.
| Pause TOR synth (↓) | Pause TOR CANDOR (↓) | Smooth-turn TOR (↑) | Smooth-turn latency | Interruption TOR (↑) | Post-interruption quality (GPT-4o, 0–5) | |
|---|---|---|---|---|---|---|
| Moshi | 98.5 | 98.0 | 94.1 | 0.265 s | 100.0 | 0.77 |
| Freeze-Omni | 64.2 | 48.1 | 33.6 | 0.953 s | 86.7 | 3.62 |
| PersonaPlex | 35.8 | 43.1 | 90.8 | 0.170 s | 95.0 | 4.29 |
| MoshiRAG † | 32.0 | 56.0 | 83.0 | 0.180 s | 85.0 | 3.75 |
| VoiceChat | 15.3 | 25.5 | 81.5 | 0.448 s | 100.0 | 4.33 |
| Gemini Live 2.0 (closed) | 25.5 | 31.0 | 65.5 | 1.301 s | 89.1 | 3.38 |
| GPT-Realtime (closed) | 1.0 | 12.0 | 100.0 | 1.470 s | 97.0 | 3.85 |
On FDB 1.5's user-backchannel condition it resumes its own response in 93% of cases against Freeze-Omni's 80% and Moshi's 6%, and responds unnecessarily in 1% — and it ties Gemini Live 2.0 exactly across all four behaviour categories (1 / 93 / 2 / 4), which is either a striking coincidence or an unremarked artifact of the shared FDB 1.5 scoring; the paper notes the match and does not comment on it. Against GPT-4o Realtime it wins on Resume (93 vs 70) and Unknown (4 vs 25).
What the headline claim actually rests on. "Lowest pause-handling takeover rates among evaluated open-weight systems" is true of this table and worth four qualifications, none of which the abstract carries. (i) The open-weight set is five systems, and NVIDIA re-ran none of them: Moshi and Freeze-Omni are "from the public FDB results", PersonaPlex is the number its own authors published, and MoshiRAG's row is flagged as "not part of the controlled FDB run." Only NVIDIA's own row is a fresh NVIDIA measurement. (ii) A closed system beats it on both pause tracks and on smooth turn-taking — GPT-Realtime at 1.0 / 12.0 / 100.0 — so "lowest" is a within-tier claim. (iii) The closed endpoints are historical: gemini-2.0-flash-live-001, three generations behind the Gemini 3.8 Live this vault tracks, with the paper's own footnote conceding the names "refer to the evaluated historical endpoints, not necessarily their current service versions." (iv) Single run, no seeds, no error bars anywhere in the paper. Catalogued in full on Interactivity Benchmarks.
The finding that matters more than the leaderboard row: a duplex model shipping a VAD by the back door. This page has, since GPT-Live, carried the claim that full duplex "eliminates the need for a separate detector." NVIDIA's §4 states the opposite as a shipped inference-time enhancement:
"When the model fails to natively handle turn-taking/barge-in scenarios, we introduce a simple endpointing mechanism as a fallback, using the RNN-T transcript output. A set of heuristics based on user speech/silence activity is combined with the current response generation state, to forcefully inject BOS/EOS tokens into the model."
Turn-taking here is a trained frame-level behaviour (BOS/EOS as per-frame targets on the agent-text channel — see Time-Aligned Micro-Turns) with a heuristic speech-activity endpointer wired in behind it that can override the model's own floor decision. That is the detector, re-entering the architecture as a fallback rather than a component. It is the first such admission in this corpus, and it is from the lab with the most complete public disclosure — which is the point: the other systems may well do the same, and nobody else publishes §4.
Speak-while-thinking: now measured twice, and refuted twice. The running tally above stands at asserted by four labs, demonstrated by none. This is the second system in the corpus to publish what happens during the tool call, and it goes further than Hu et al.'s frontend did. During execution the runtime speaks a pre-configured per-tool acknowledgement ("Sure, let me calculate that for you") whose duration is explicitly designed to mask the tool latency, the agent-text channel is padded, and — the new fact — "incoming audio may continue through the perception and RNN-T transcription path, but it is not used to condition response generation during tool execution. Consequently, barge-in is unavailable during this phase." Hu et al.'s delegating frontend at least kept listening in a way that reached the eventual call; this one transcribes but cannot be interrupted. Two NVIDIA systems on opposite sides of the internalize-vs-delegate axis, two weeks apart, both surrender the live path for the duration of a tool call. Count: asserted four times, measured twice, demonstrated zero times.
A test of the survey's "still an industrial open problem", and it does not overturn it. The section above records An, Lu, Dong et al. (May 2026) declaring stable low-latency full-duplex unsolved against four vendor assertions that it is. An open-weight release four months later with published FDB numbers is the best available evidence either way, and it lands on the survey's side — by the authors' own §7. The limitations are structural, not polish: an audio context window of at most ~2 minutes, a practical recommendation of no more than five tools per session, simultaneous multi-tool invocation "not yet reliable", calls "skipped, incorrectly selected, or supplied with invented arguments", long tool responses delaying subsequent speech, no barge-in during tool execution, and degraded robustness "in strongly noisy or reverberant conditions, particularly in the presence of competing background speech." The duplex interaction is genuinely good and the system around it is not yet deployable in the survey's sense.
And it sharpens the duplex-vs-speech-to-speech separation above, in the other direction. The section on that distinction was built from a system NVIDIA called a duplex speech-to-text frontend. This paper calls itself speech-to-speech in its title and is architecturally the same shape: the LM emits agent text, a separately trained TTS decoder (VoiceChat-TTS, a 778M Gemma-3-based backbone plus a 199M causal codec) renders it, the backbone's audio-loss weight is 0.0, and "gradients are not propagated between the full-duplex backbone and TTS model." There is a text bottleneck and a bolted-on renderer inside a model badged speech-to-speech. Whatever the label buys, it is not end-to-end audio: cf. Native Multimodal Modeling: Fusion Depth and I/O Duality, where the survey's own operator definitions make a decoupled, gradient-isolated renderer a grafted head — the thing it excludes from "native" outright.
Connections#
- Native Multimodal Modeling: Fusion Depth and I/O Duality — the architecture-side statement of the same claim: full-duplex state management as a deployment route, and "streaming by construction, not a post-hoc wrapper" named as an unsolved industrial problem by a non-vendor
- Interaction Models — parent concept
- Time-Aligned Micro-Turns — the mechanism (no turn boundaries) that makes full-duplex possible
- Encoder-Free Early Fusion — joint multimodal reasoning is what lets a visual change trigger speech
- Turn-Based Interface Bottleneck — the half-duplex status quo this replaces
- Interactivity Benchmarks — TimeSpeak / CueSpeak / RepCount-A / ProactiveVideoQA / Charades measure exactly these modes
- Interaction / Background Model Split — where the concurrent deep work goes
- TML-Interaction-Small — the model that demonstrates these interaction modes
- GPT-Live — audio full-duplex in production; the turn detector eliminated at ChatGPT scale
- Live-Path Minimalism — the serving architecture the every-frame-on-schedule constraint forces
- NVIDIA — the duplex speech-to-text frontend that separates the duplex property from the speech-to-speech one
- Content-Driven Intervention — the third separation: deciding to speak because the content warrants it, measured and absent across five open duplex families
- Gemini 3.8 Live — visual grounding and 97-language mid-conversation switching claimed as shipping live modes, unmeasured
Sources#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- How we built a realtime system for responsive voice AI in six months — OpenAI, 2026-07-29 (
case-study): audio full-duplex in production; detector removal - A frontend-backend architecture for tool calls in full-duplex speech models — Hu et al. (NVIDIA), arXiv 2609.19334, 2026-09-16 (
empirical, 5pp): §3.1 the duplex STT frontend plus separate streaming TTS; §5.1.2 the FDB3 backchannel/interruption analysis; §5.2 the added Speakability metric. Full treatment on Interaction / Background Model Split - Full-Duplex Speech Models Take the Floor When Asked, Not When Needed — Peng, Nuchged, Fu & Yao, arXiv 2609.19596, 2026-09-17 (
empirical, 5pp): Tables 1–4, the matched-context protocol and the negative result on proactive interjection. All four tables reconciled againstpdftotext -layout. Full treatment on Content-Driven Intervention - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim): the third stay-present assertion, native ASR-transcript and response-text output, turn detection as an optional feature on a non-turn-based model, telephony, and the FDB v1.5 Interactivity / v1 latency cards (catalogued on Interactivity Benchmarks) - Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Ouyang & Jaganathan (Google, Gemini Audio Team), 2026-09-15 (
vendor-claim): the two newly claimed modes, the fourth speak-while-thinking assertion and the hold phrase sold beside it, and the EVA-Bench cascade corners. Chart positions are read from an unlabelled scatter viewed at crop during compile — ordering only, no numbers. Full treatment on Gemini 3.8 Live and Interactivity Benchmarks - NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA), NemotronLabs VoiceChat, arXiv 2609.21967, 2026-09-18 (
empirical, 19 pp): Table 1 (FDB 1.0), Table 2 (FDB 1.5), §4's heuristic BOS/EOS endpointing fallback, §7's limitations, Appendix B's barge-in-unavailable-during-tool-execution statement. All 7 tables reconciled cell-for-cell againstpdftotext -layout; every baseline row in Table 1 is copied from third-party reports, not re-run. First-party report on the same weights Peng et al. measured third-party. Full treatment on Interaction / Background Model Split and Interactivity Benchmarks - Toward Native Multimodal Modeling: A Roadmap — An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities; arXiv 2605.25343, 2026-05-25;
practitioner-opinion, 52 pp): §6.3 the four streaming/duplex deployment routes and the cache-contention framing, §8.4 "streaming by construction, not as a post-hoc wrapper," Table 1's single full-duplex entry (Moshi) and Table 3's full-duplex evaluation rows. A non-vendor third party declaring the problem unsolved in May 2026. Full treatment on Native Multimodal Modeling: Fusion Depth and I/O Duality
Cited by 15
- The Future of Agent Interfaces×4
Human collaboration · Interaction Models / Full Duplex Interaction · Human senses, speech, screen,…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×4
A census-class candidate arrived four months later, and the operators disqualify it from M2M (added…
- Time-Aligned Micro-Turns×4
Every interaction mode that needs a special-purpose harness today becomes a special case of model…
- Content-Driven Intervention×3
Full Duplex Interaction lists proactive interjection — "interrupt when I say something wrong" — as…
- Interactivity Benchmarks×3
Full Duplex Interaction — TimeSpeak/CueSpeak/RepCount-A/ProactiveVideoQA/Charades each target one…
- Encoder-Free Early Fusion×2
Early fusion (everything into one transformer) means the model reasons jointly over modalities…
- GPT-Live×2
OpenAI's third-generation voice system, launched July 2026 — a full-duplex voice model (Full Duplex…
- Interaction Models×2
See Full Duplex Interaction and Interactivity Benchmarks for how these are demonstrated and…
- Live-Path Minimalism×2
The serving-side architecture behind Gpt Live, stated by OpenAI as one principle: "the voice must…
- NVIDIA×2
Full Duplex Interaction — its duplex speech-to-text frontend separates the duplex property from the…
- Gemini 3.8 Live
Full Duplex Interaction — visual grounding and mid-conversation language switching as claimed…
- Interaction / Background Model Split
Full Duplex Interaction — where the concurrent deep work goes while the interaction model stays…
- Interaction & Multimodal
Full Duplex Interaction — Perceive-and-respond simultaneously across modalities — a property of…
- TML-Interaction-Small
Full Duplex Interaction — the interaction modes it demonstrates
- Turn-Based Interface Bottleneck
Full Duplex Interaction — the interaction modes the bottleneck currently blocks
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
- Time-Aligned Micro-Turns
The core interaction-model move: input/output as continuous streams in ~200ms interleaved chunks, no turn boundaries; s…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
An, Lu, Dong et al. (Tencent Youtu + 5 universities, May 2026) formalize 'native' as two operator definitions — mid-fus…
