Sources#
- A frontend-backend architecture for tool calls in full-duplex speech models
- Build more natural voice experiences with GPT‑Live‑1 in the API
- How we built a realtime system for responsive voice AI in six months
- Inkling: Our Open-Weights Model
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Summary#
Interaction Models are architected as two cooperating models:
- a time-aware interaction model that maintains real-time presence — perceiving and responding in a continuous loop (see Time-Aligned Micro-Turns);
- an asynchronous background model that handles sustained reasoning, tool use, and longer-horizon work.
The payoff: the user gets both responsiveness and depth — "the planning, tool-use, and agentic workflows of reasoning models at the response latency of non-thinking ones."
How delegation works#
- When a task needs deeper reasoning than can be produced instantly, the interaction model delegates to the background model, which runs asynchronously.
- The handoff is a rich context package — not a standalone query, but the full conversation.
- The interaction model stays present throughout — answering follow-ups, taking new input, holding the thread.
- Results stream back as the background model produces them; the interaction model interleaves updates into the conversation at a moment appropriate to what the user is currently doing — not as an abrupt context switch.
Both halves are intelligent#
This isn't a "dumb frontend, smart backend" design. The interaction model on its own is "competitive on both interactive and intelligence benchmarks" — see Interactivity Benchmarks (e.g. TML-Interaction-Small beats every non-thinking baseline on Audio MultiChallenge APR even without the background agent; benchmarks marked * use the background agent for reasoning/tool tasks).
Relationship to other multi-model patterns#
This is the latency-vs-depth axis of multi-model orchestration, distinct from:
- the role-based model selection in Client-Side Agent Optimization (assign cheap/expensive models per role in an agent graph) — there the split is cost-driven and static; here it's latency-driven and dynamic-per-turn;
- the three-agent / reviewer-in-fresh-context pattern (Deep Modules for Agents, Agent Harness Engineering) — there the split is for context isolation; here it's for temporal concerns (stay responsive vs. think hard).
The background half gets a name: Inkling (July 2026)#
At the split's introduction the background model was an unnamed capability. Inkling fills the slot: TML states that "a major goal of Inkling's design is to serve as the background reasoning model in the interaction models system" — which is why a 975B open-weights foundation model was trained natively multimodal (encoder-free dMel audio and hMLP vision, the same input stack as the interaction model) rather than as a text reasoner with adapters. The symmetry runs deeper: Inkling-Small is a 276B/12B MoE, TML-Interaction-Small's exact shape, suggesting the two halves of the split share a lineage. Both halves of the architecture are now public artifacts rather than one model and one promise.
The split ships in production: GPT-Live (July 2026)#
OpenAI's GPT-Live is the same two-model architecture arrived at independently and deployed at ChatGPT scale — a full-duplex voice model holds the conversation while deeper reasoning and tool use delegate to frontier models such as GPT-5.5 on an asynchronous path. OpenAI's phrase for the payoff mirrors TML's: "effectively decoupling 'talking' from deeper 'thinking'." Two months after TML's research preview, both halves of the split now exist as production systems at a second lab.
The production account adds the engineering the research framing left open — the delegation loop as a latency budget (case-study, first-party):
- The rich context package becomes a standing prefilled session. At voice-session start, the application server pre-creates the frontier model's inference session and prefills it with the initial conversation context, so the prompt is fully processed before the first delegation is requested — TML's delegate-the-full-conversation handoff with its prefill cost paid in advance.
- Session affinity + prompt caching keep successive delegations cheap for the conversation's duration, with worker failure still cheaply recoverable.
- Every lever on time-to-useful-result is tuned: reasoning effort, output limits, tool schemas, and model↔tool round trips.
- The interaction half can stall, not hide. The voice model "can briefly keep the exchange moving while a frontier model reasons or uses tools, but it cannot hide an arbitrarily slow response" — the empirical bound on how much latency the split's front half can absorb.
Serving-side detail in Live-Path Minimalism.
A third derivation, and the first published delegation signal (NVIDIA, September 2026)#
Hu et al. (NVIDIA, arXiv 2609.19334, 2026-09-16, empirical) build the same split a third time, from a third starting point — a streaming-ASR lineage rather than a scaling bet or a serving rewrite — and they build it because the first two never said how the handoff works. Their survey sentence is the reason this source matters here:
"However, it remains unclear how the delegation signal in [28, 29, 30] is designed and background model interacts with the duplex frontend."
The three bracketed references are Thinking Machines' interaction models, Qwen-audio-agent, and GPT-Live. This page's mechanism section was, until now, assembled from two accounts that describe the delegation's effect and withhold its format. NVIDIA publishes a format.
The system. The interaction half is a duplex speech-to-text model: a 600M-parameter Parakeet streaming speech encoder feeding an NVIDIA Nemotron-Nano-9B-v2-Base backbone, with a streaming-ASR head and an agent-text head sharing that backbone and emitted in a single decoding pass. Agent text goes to a separate streaming TTS (VoiceChat-TTS) for audio. The background half is an off-the-shelf LangGraph ReAct agent — an agent node holding an instruction-following LLM, a tools node, and a conditional edge that loops until the turn produces no more calls.
The delegation signal, in full. Four special tokens in the agent-text channel, and nothing else:
<tc bos>replaces the ordinary<agent bos>roughly 320 ms after the end of the user turn when the frontend judges the query needs a tool.- A short spoken filler (~1 s, e.g. "Give me a moment.") follows, then
<tc eos>. <tc eos>firing is what dispatches: the streaming ASR transcript, endpointed by<user eos>, is sent to the backend graph as a message.- The backend's natural-language answer returns wrapped in
<pf bos>/<pf eos>, is prefilled into the agent-text channel, and the frontend is trained to reproduce it exactly starting from<agent bos>. Prefill spans are masked out of the loss; the reproduction is not. Agent output is suppressed with pad tokens for the duration of the call.
Training makes the token learnable rather than rule-based: the SFT mixture includes ~8.5k hours of multi-turn tool-call conversation synthesised from text dialogues (generated by Nemotron 3 Nano, Gemma-4-31B-IT and Qwen3.5-397B-A17B, LLM-judge filtered, TTS'd and WER/CER-filtered), with When2Call-style variants covering when not to call and when to ask a follow-up instead.
Three ways it diverges from the specification above — worth holding side by side, because this page's mechanism description came from the labs that shipped, and this is the lab that published:
| TML / GPT-Live as described above | NVIDIA, as implemented and measured | |
|---|---|---|
| What crosses the boundary | a rich context package — the full conversation | the current turn's ASR transcript only; multi-turn state is held in the backend by a thread-keyed LangGraph checkpointer |
| Frontend during the call | "stays present throughout — answering follow-ups, taking new input, holding the thread" | silent: "the frontend remains silent during the tool calls", agent output suppressed with pad tokens, after a ~1 s hold phrase |
| Result integration | streamed back, interleaved at a moment appropriate to what the user is doing | prefilled and repeated verbatim; no interleaving policy exists |
The middle row is the substantive contradiction. GPT-Live had already conceded the bound — the voice model "can briefly keep the exchange moving while a frontier model reasons or uses tools, but it cannot hide an arbitrarily slow response." NVIDIA's system is that bound at its floor: a hold phrase, then nothing. The frontend does keep listening throughout (it is still consuming user audio and producing ASR, and on FDB3 it backchannels at disfluency pauses before the call fires) — so speak-while-listening survives and speak-while-thinking does not. Weighting: TML's stays-present claim is vendor-claim, OpenAI's is case-study first-party, and neither publishes a measurement of it; NVIDIA's is empirical but is one implementation whose authors were optimising for minimal frontend change. The honest reading is that staying present during delegation is asserted twice and demonstrated zero times, not that it is false.
A third rationale for the split, and it is not latency. TML argues depth-at-low-latency; OpenAI argues the live path must not stall. NVIDIA argues capacity economics:
"audio-native modeling imposes a fundamental capacity tradeoff: audio tokens consume parameters and context budget that text-only LLMs can devote to factual knowledge, instruction following, and tool-call capabilities. In contrast, a delegation-style backend agent is more modular and consumes little modeling capacity from a frontend speech model."
That is an argument about where parameters go, and it would survive even if delegation were free in latency terms. The motivating gap is measured elsewhere: τ-Voice finds leading commercial duplex voice models complete 31–51% of grounded customer-service tasks under clean conditions against 85% for GPT-5 on the text versions of the same tasks.
The split behaves as a capability interface, not just a convergent shape. The same frontend weights are paired with three different backends and the agentic numbers move while nothing about the interaction half is retrained:
- BFCL-audio AST average: 71.7 (Qwen2.5-7B) → 73.0 (Qwen3-30B-A3B) → 74.6 (Qwen3-30B-A3B with external rather than internal ASR transcripts; the gain concentrates in Parallel-Multiple, 55.1 → 61.1).
- FDB3 response quality: 54.0 (30B) → 67.0 (Qwen3-235B-A22B); Pass@1 44.0 → 48.0.
- EVA-Bench EVA-A: 32.4 (30B) → 46.6 (235B), task completion 40.4 → 57.3 — from below Qwen3-Omni-30B-A3B-Instruct to level with Gemini 3.1 Flash Lite.
This is the first evidence in the corpus that the split is a lever you can pull without retraining the interaction model — which is an argument for it being architectural rather than a transitional scaffold, and simultaneously exactly what a well-designed transitional scaffold looks like. The paper itself frames the underlying question as open: "whether duplex speech models should directly internalize tool-call capabilities or instead delegate such capabilities to a backend text agent," and names DuplexSLA (arXiv 2605.20755) as the internalise-it alternative, which puts tool calls in a dedicated channel inside a Moshi-style duplex model.
What the delegation token costs the frontend it was supposed to leave alone. "Minimal modifications" here means precisely one control token added to the prediction targets and no architecture change — and NVIDIA runs the ablation nobody else has, against an identically trained frontend without tool-call data. The live-path cost is near-zero and the conversational-quality cost is not; figures and caveats on Live-Path Minimalism.
Open source status. No frontend weights or training code are announced — only a Hugging Face Spaces demo. Every other component is public: Parakeet, Nemotron-Nano-9B-v2-Base, VoiceChat-TTS, LangGraph, Triton, Chatterbox, and all three backends are open-weight Qwen checkpoints. So the split is reproducible in shape by anyone, and the trained delegation behaviour is not.
The split gets a price on each side: GPT-Live-1 in the API (September 2026)#
Everything above treats the split as an architecture. OpenAI's API launch (2026-09-10, vendor-claim) makes it a commercial boundary, which is a different kind of evidence about how durable the shape is: you can refactor an internal seam away, but a seam with a price list and two vendors on it is harder to close.
Each half is metered separately, in a different unit. GPT-Live-1 sells at $0.05 per minute for "the front-end voice layer." The back half is not included: "Pair it with the backend model and agent harness that fit your product," billed by whoever supplies it, per token. Per-minute on the interaction half and per-token on the background half is the pricing form of this page's whole thesis — one side is bought by time present, the other by work done. Cost consequences on Cost-per-Task Over Cost-per-Token and Live-Path Minimalism.
The back half is explicitly swappable, including off-vendor. OpenAI names three of its own — GPT-6 Astra for complex reasoning, Luna for high-volume tasks like scheduling and order updates, Terra in the tool-calling benchmark runs — and then says the quiet part: "a backend text model like GPT-6 Astra or a third-party model." The stated reason is per-task fitting: "that flexibility lets developers match reasoning depth, speed, and cost to each task." This is the interface claim of the NVIDIA section above, asserted by a second lab about a production product rather than demonstrated on a bench: the interaction half is backend-agnostic by design, and OpenAI is willing to lose the backend revenue to say so.
And OpenAI's own benchmark footnotes confirm the interface empirically, in its own favour. Four of its seven comparison cards state which backend the GPT-Live configuration used and three do not — Tau3 (86.2% Pass@1) and Tau Banking (32.0%) ran Astra at medium effort, tool calling (87.0% Pass@1) and response quality (90.0%) ran Terra at low, while Conversational Dynamics, Full Duplex Bench v1.5 Interactivity and turn-taking latency carry no backend note. So OpenAI is publishing the same decomposition NVIDIA measured: interaction numbers belong to the frontend, intelligence and tool numbers belong to the pairing — and the vendor chose a different backend at a different effort for the two task families, which is the capability-interface lever being pulled in a marketing table. Figures and the full card set on GPT-Live; catalogue entry on Interactivity Benchmarks.
The delegation call has a published application-facing shape. NVIDIA published the in-model signal (<tc bos>/<tc eos>/<pf bos>/<pf eos>); OpenAI now publishes the other end of the same handoff — a developer-side agent returns its answer with live.send({ type: "session.commentary.append", delegation_id, content }), an async append against a session-issued delegation id, with "connection setup and delegation handling … omitted." The two disclosures are complementary rather than comparable: one is how a duplex model decides to hand off, the other is how an application hands back. Nothing in the API post says what the frontend emits while the delegation is outstanding, so the token-level question the NVIDIA section opens stays open.
Stay-present, asserted a third time and measured a zeroth. The post's claim: GPT-Live-1 "can respond to interruptions and acknowledgements as they happen, while delegating deeper reasoning to the back end. This lets the conversation continue while work happens in the background." That is TML's stays-present and OpenAI's July restatement, now asserted for a shipping API — and again with no measurement, no filler-rate figure, and no comparison against a hold-phrase policy. It does add one policy datum that cuts against NVIDIA's design: OpenAI markets "silent context management" as handling background noise and silence "without interrupting the conversation or narrating every step out loud" — i.e. explicitly selling the absence of the hold phrase that NVIDIA's frontend emits by design at an 83–85% filler rate. So the divergence table above now has a vendor on each side of the middle row stating opposite policies as features, and still nobody has run both against the same backend latency distribution. Count: asserted three times, demonstrated zero times.
A fourth vendor, no backend, no price, and both policies at once (Google, September 2026)#
Google's Gemini 3.8 Live launch (2026-09-15, vendor-claim) is the first frontier live-voice announcement in this corpus that names no background model, publishes no delegation mechanism, and prices nothing — and it is nonetheless the most informative source yet on the middle row of the divergence table above, because Google sells both sides of it in the same paragraph.
The stay-present assertion, fourth time. "It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background," and, of Extended Thinking, it "reasons and speaks simultaneously … while maintaining an uninterrupted conversational flow." One demo caption claims multi-step bookings and asynchronous function calls "all without interrupting natural live conversation." No filler rate, no latency figure, no comparison arm, no protocol. Count: asserted four times, demonstrated zero times.
And the hold phrase, sold as a feature, two sentences later. The same paragraph advertises "early verbal cues like 'Let me check that…' to acknowledge prompts naturally, and live progress narration to walk users through multi-step background tasks as they progress." Set that against the two positions this page already records: NVIDIA's frontend emits a ~1 s hold phrase and goes silent, at an 83–85% FDB3 filler rate, by design; OpenAI markets the absence of exactly this — "without interrupting the conversation or narrating every step out loud" — as "silent context management." Google is now on both sides: it claims the conversation continues and ships the acknowledgement filler and ships the step narration, and treats all three as selling points. A vendor that had genuinely closed the stay-present gap would have no use for an early verbal cue. A vendor whose model goes quiet would not claim uninterrupted flow. The most economical reading is that "stays present" in commercial usage means emits something rather than continues the substantive conversation — which is a semantic drift this page should track, because it makes the open question below harder to falsify rather than easier: three vendors can all be right about different verbs.
The priced boundary has no second instance. The trigger set on 2026-09-10 was whether a second lab would price a live layer separately from a backend. Google's post contains no pricing of any kind — not per minute, not per token, not per audio hour. The only money on the page is a third party's measured cost per hour of input audio ($0.84 for 3.8 Live, $3.50 for Extended Thinking, $5.83 for GPT-Live-1 + Astra), which is a benchmark result, not a price list, and which explicitly does not decompose into a front-half and a back-half charge. So the commercial-durability argument above still rests on one vendor, one price point, one post, and the "priced interaction boundary" page candidate stays declined at pressure 3 — now with a dated negative: the second frontier voice launch of the month declined to expose the seam at all.
A grammatical tell worth recording, since there is no disclosure to record instead. OpenAI's benchmark cards footnote a backend model and its effort (Astra at medium, Terra at low); Google's chart bars label a reasoning effort on the live model itself (Extended Thinking at High, 3.8 Live with none, Gemini 3.1 Flash Live at High and Minimal). The effort dial sits on opposite sides of the seam in the two vendors' own presentation. That is evidence about how each vendor wants the product understood — OpenAI as a pairing you assemble, Google as a model you choose — and it is not evidence about either architecture. Google's agentic runs are footnoted as executed "on the Live API on Gemini Enterprise Agent Platform," a harness, which is the closest the post comes to conceding that something else is in the loop. Argument and caveats on Interaction Models; product capsule on Gemini 3.8 Live.
Open / acknowledged#
TML calls background agents "an essential capability" they've "just scratched the surface" on — both pushing background agentic intelligence to the frontier and exploring how background agents work together with the interaction model.
The other branch, from the same lab, two days later (NVIDIA, September 2026)#
Hu et al. framed the choice and declined to make it: "whether duplex speech models should directly internalize tool-call capabilities or instead delegate such capabilities to a backend text agent." They named DuplexSLA as the internalise-it alternative and left it there. Forty-eight hours later the same lab published the internalised system — NemotronLabs VoiceChat (arXiv 2609.21967, 2026-09-18, empirical) — on overlapping hardware, an overlapping backbone and the same benchmark. This corpus does not usually get the counterfactual; here it does, with the usual caveat that neither paper cites the other and the two runs were not designed as arms of one experiment.
The mechanism: a parallel function channel, not a serialized one. Tool calls live in a dedicated autoregressive channel running alongside the agent-text channel, not inside it. At each 80 ms frame the modality-fusion layer forms a weighted sum of three things — encoded user audio, the preceding agent-text-token embedding, and the preceding function-token embedding — at fusion weights 1, 1, 2, and a separate function head predicts the next function-channel token. The channel emits padding while no action is required, and the paper is explicit that this negative supervision is "important for preventing spurious calls" (the function head is supervised on padding even during CPT, which contains no tool examples at all). A call opens with <SOTC>, closes with <EOTC>, the tool-response span follows and terminates at <EOTR>; internally these are reserved vocabulary slots <SPECIAL_20/21/22>. The <TOOLCALL> payload is a JSON list of {name, arguments} objects, so one or several parallel calls share a representation. Tool-response tokens are fed as context and masked out of the loss. The function-channel loss weights are steeply shaped — 64.0 on <TOOLCALL> content, 6.0 each on <SOTC>/<EOTC>, 3.0 on <EOTR>, 0.3 on padding — against the agent-text channel's 12.5 / 7.5 / 5.0 / 1.0 for begin-of-turn, end-of-turn, content and padding.
This is a third position on the axis, and the paper stakes it deliberately: "Unlike DuplexSLA, which serializes heterogeneous action-related tokens within a shared autoregressive channel, NemotronLabs VoiceChat maintains parallel, specialized streams, preserving the low-latency behavior required for full-duplex interaction." So the choice was never binary. Delegate to a text agent (Hu et al., GPT-Live, TML); serialize actions into the conversational channel (DuplexSLA); or run actions on their own channel on the same timeline. The third costs a head and a fusion term and keeps the conversation channel clean.
The runtime names both halves of the handoff, which is the counterpart of Hu et al.'s <tc bos>/<pf bos> disclosure: "fast-decode" is the asynchronous decode of the complete call once the channel emits <SOTC>, and "fast-inject" is the runtime serializing the tool result and forcibly inserting its tokens into the function channel and decoder context. Between the two, TTS speaks a per-tool acknowledgement string defined in the tool schema itself ("ack_message": "Sure, let me calculate that for you"), and the paper says in as many words that its duration should be chosen to "mask the delay" — skippable for fast tools, long for slow ones.
The head-to-head, same benchmark, same lab, two weeks apart. FDB 3.0, percentages:
| Tool selection | Argument accuracy | Pass@1 | Response quality | |
|---|---|---|---|---|
| Internalized — VoiceChat, parallel function channel | 82.5 (F1) | 42.2 | 33.0 | not reported |
| Delegating — frontend + Qwen3-30B-A3B backend | 74.6 (acc) | 52.8 | 44.0 | 54.0 |
| Delegating — frontend + Qwen3-235B-A22B backend | 71.7 (acc) | 55.2 | 48.0 | 67.0 |
Read carefully, because the two papers report the routing column under different names — F1 for the internalized system, accuracy for the delegating one — and neither states that they are the same quantity. Treat the first column as suggestive and the other two as the real comparison. What falls out is a clean division of labour: the duplex model is at least as good at deciding which tool to call, and materially worse at filling in its arguments and finishing the task. The internalized system's own §5 says the same thing from the inside — "tool routing is substantially stronger than argument extraction and end-to-end execution" — and Pass@1 on FDB 3.0 is scored only when the model picks exactly the right tools and supplies perfect arguments for every call, so 82.5 routing against 33.0 Pass@1 is argument grounding failing, not routing.
That is the sharpest evidence this page has on why the split exists. Delegation is not buying tool selection; it is buying the backend's argument extraction and multi-step composition — precisely the mature instruction-following-and-long-horizon-reasoning capability Hu et al. said they were reaching for. It also means the capacity-economics argument was aimed at the wrong column: audio tokens do not appear to cost the frontend its ability to know that a tool is needed and which one.
Where the internalized branch does not escape the split's own failure. The falsifier standing on Interaction Models is "a single duplex model that matches a delegating system's tool-call accuracy without going quiet." VoiceChat is the single duplex model and it also goes quiet — a pre-configured acknowledgement utterance, a padded agent-text channel, and a stronger concession than Hu et al.'s: incoming audio keeps flowing through perception and RNN-T "but it is not used to condition response generation during tool execution. Consequently, barge-in is unavailable during this phase." Hu et al.'s frontend at least kept listening in a way that reached the eventual call. So on the live-path question the two branches converge on the same behaviour, and the internalised one is slightly worse. The paper's own future work asks for "interruption-aware tool execution," which is the admission in the authors' words.
The practical ceiling is also published, and it is low: no more than five tools per session is the paper's own recommendation, simultaneous multi-tool invocation "is not yet reliable," calls may be "skipped, incorrectly selected, or supplied with invented arguments," the model "may answer from internal knowledge when a tool should instead be invoked," and long tool responses delay subsequent speech. A delegating backend with a 235B instruction-tuned model behind it has none of those limits by construction.
Open source, unlike its sibling. Hu et al. released no frontend weights. This checkpoint is on Hugging Face as NVIDIA-NemotronLabs-VoiceChat-11B — which is what let Peng et al. measure it independently a day before the technical report appeared. So the internalized branch is the reproducible one, and the delegating branch is the one with the better Pass@1 and no public weights.
Connections#
- Native Multimodal Modeling: Fusion Depth and I/O Duality — the survey's operator definitions classify the internalized duplex system above as mid-fusion M2T with a grafted speech renderer rather than M2M, which is the architectural form of the split: agent text is native and the voice is bolted on
- Interaction Models — parent concept
- Time-Aligned Micro-Turns — what keeps the interaction model present while the background model thinks
- Interactivity Benchmarks —
*-marked results use the background agent; also where the tool-calling evals NVIDIA runs (BFCL-audio, FDB3, EVA-Bench, τ-Voice) are catalogued - Client-Side Agent Optimization — a different axis of multi-model design (cost/role, not latency/depth)
- Deep Modules for Agents / Agent Harness Engineering — multi-agent splits for context isolation rather than latency
- Harness Shrinkage as Models Improve — open question whether the split is permanent or a transitional artifact until one model is both fast and deep enough
- Encoder-Free Early Fusion — the other half of the architecture (the perception/generation side)
- Full-Duplex Interaction — where the concurrent deep work goes while the interaction model stays present
- TML-Interaction-Small — the model that implements the split; competitive on intelligence benchmarks even without the background agent
- GPT-Live — OpenAI's production instance of the split: full-duplex voice model delegating to GPT-5.5
- Live-Path Minimalism — the serving architecture that keeps delegation off the live media path
- Cost-per-Task Over Cost-per-Token — the productised split bills its two halves in different units: per-minute presence on the front, per-token work behind it
- NVIDIA — the third lab to derive the split, and the one that published the delegation signal
- Gemini 3.8 Live — the fourth live-voice product: no backend named, no price on either half, and stay-present asserted alongside the hold phrase that contradicts it
Open Questions#
- Does an interaction model that keeps talking during delegation actually beat one that emits a hold phrase and goes silent, on user-perceived quality? Thinking Machines and OpenAI both assert stays-present; NVIDIA measures the silent variant and reports 100% turn-taking with an 83–85% filler rate by design. Nobody has run both policies against the same backend latency distribution and asked users. (2026-09-21: still nobody. Build more natural voice experiences with GPT‑Live‑1 in the API asserts stay-present a third time and markets the absence of a hold phrase — "without narrating every step out loud" — as a product feature, so the two policies are now explicitly opposed vendor positions rather than an unnoticed divergence. No measurement on either side; the question is unmoved and better motivated.) (2026-09-21, second bearing note: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking asserts stay-present a fourth time — "acknowledge requests and keep chatting while tasks finish in the background" — and in the same paragraph sells an early verbal cue ("Let me check that…") and live progress narration through multi-step background tasks. One vendor now holds both policies simultaneously, which retires the reading that the two are competing design schools and replaces it with a harder problem: the phrase "stays present" is not being used to mean the same behaviour by the three parties asserting it. Any experiment that settles this now has to define the dependent variable before it runs — substantive conversational continuation, or any emission at all. Still zero measurements.)
- Is the duplex-quality cost of adding a delegation token intrinsic to the token, or an artifact of synthetic training data? NVIDIA's ablation pays ASR WER 10.80 → 11.47 and FDB-v1 pause TOR 53.2 → 68.2 for the token, and attributes both to synthetic SFT audio and absent natural-pause data. A second frontend trained with natural-pause and real-speech tool-call data settles it. Partially answered (2026-09-23) — the synthetic-audio explanation is weakened, not the token one. NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities (
empirical) trains a sibling duplex model whose speech is also almost entirely TTS-rendered — both CPT and SFT "draw their corpora from the text-only Nemotron backbone rather than from natively recorded conversational speech" — and it nonetheless reaches 15.3% / 25.5% FDB 1.0 pause TOR, the lowest of the five open-weight systems in its table and far below the 53.2 → 68.2 range the ablation moved through. It carries tool calling in a parallel function channel rather than a delegation token, and it adds three conversational augmentations the ablated frontend did not report: early interruption at p = 0.1 (agent turn truncated, continuing 8 frames / 640 ms before EOS), backchannel injection at p = 0.05 per sample, and a two-frame (160 ms) agent-text delay with the function channel deliberately left unshifted. So synthetic audio is not a hard ceiling on pause behaviour, and the augmentation schedule is the better suspect. It is not an ablation — different model, different data, different channel design — so the token half of the question is untouched and the original falsifier still stands.
Sources#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Inkling: Our Open-Weights Model — Inkling as the designed background reasoning model (
vendor-claim) - How we built a realtime system for responsive voice AI in six months — OpenAI, 2026-07-29 (
case-study, first-party): GPT-Live's delegation path as a latency budget (pre-created prefilled sessions, affinity + caching, tuned effort/schemas/round-trips) - A frontend-backend architecture for tool calls in full-duplex speech models — Hu et al. (NVIDIA), arXiv 2609.19334, 2026-09-16 (
empirical, 5pp): §3.1 the frontend and the<tc bos>/<tc eos>/<pf bos>/<pf eos>token scheme, §3.2 the LangGraph backend and its thread-keyed checkpointer, §4 the training recipe and the worked example sequence, §5.1–5.2 the BFCL-audio / FDB3 / EVA-Bench results, §5.3 the no-tool-call ablation. Self-reported on its own system, single run, no error bars; baselines for other models were re-run by these authors. Parse note: Tables 4, 5 and 7 raisedtable-shifton 3 cells — the flag is a multi-level-spanning-header artifact (docling repeats the flattened header text across the spanned columns), and every cell was re-read frompdftotext -f 4 -l 5 -layout; Tables 1–3 were likewise re-read frompdftotext -f 2 -l 3 -layoutand match the docling body digit-for-digit. No value on any wiki page comes from an unreconciled row. - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim, ~1,670 words, unbylined): the split as a priced product boundary — $0.05/min front-end layer, developer-chosen and separately-billed backend including third-party models, the named Astra/Luna/Terra backends, thesession.commentary.appenddelegation call, and the per-card backend footnotes that separate frontend-only from pairing scores. First-party throughout, measured only against OpenAI's own prior models. - NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA, 48 contributors), NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities, arXiv 2609.21967, 2026-09-18 (
empirical, 19 pp): §2.2 the parallel function channel, its<SOTC>/<EOTC>/<EOTR>state machine and the 1/1/2 modality-fusion weights; §3.2 the 64.0/6.0/3.0/0.3 function-channel loss weights; §4 the filler-message and endpointing-fallback enhancements; Table 4 the FDB 3.0 head-to-head; Appendix B the fast-decode / fast-inject runtime and the barge-in-unavailable concession; §7 the five-tool ceiling. All 7 tables reconciled cell-for-cell againstpdftotext -layout; NVIDIA scoring NVIDIA, single run, no error bars, and the FDB 3.0 baseline set differs from the sibling paper's (Gemini Live 2.5/3.1 here, Gemini-3.5 Flash and GPT-realtime there) — see Interactivity Benchmarks - Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Ouyang & Jaganathan (Google, Gemini Audio Team), 2026-09-15 (
vendor-claim, ~1,764 words): the fourth stay-present assertion, the early-verbal-cue and progress-narration features that contradict it, the absence of any pricing, and the effort-label grammar. No backend model, no delegation mechanism and no latency figure appear anywhere in the post; the DeepMind model card it references is not ingested
Cited by 21
- Full-Duplex Interaction×8
While listening and speaking, the model can simultaneously call tools, search, browse, or generate…
- Interaction Models×7
Time Aligned Micro Turns, Interaction Background Model Split, Encoder Free Early Fusion — the three…
- Interactivity Benchmarks×5
The backend footnotes are the most useful part of the table, and OpenAI put them there itself. Four…
- GPT-Live×4
Interaction Background Model Split — the two-model talking/thinking architecture it independently…
- Live-Path Minimalism×4
Interaction Background Model Split — the two-model architecture whose delegation half this page's…
- Cost-per-Task Over Cost-per-Token×3
Interaction Background Model Split — the productised voice split bills its two halves in different…
- Gemini 3.8 Live×3
Background tool execution while the conversation continues. "It executes tools and API calls in the…
- NVIDIA×3
Hu et al. (arXiv 2609.19334, September 2026) is the third independent derivation of the Interaction…
- TML-Interaction-Small×3
Inkling-Small — the preview sibling of TML's open-weights release — is a 276B MoE with 12B active:…
- Google DeepMind×2
Interaction Background Model Split — where the lab becomes the fourth vendor to assert stay-present…
- Inkling×2
Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×2
A census-class candidate arrived four months later, and the operators disqualify it from M2M (added…
- Open Questions Backlog×2
Interaction Background Model Split (13d) — Does an interaction model that keeps talking during…
- Agent Harness Engineering
Interaction Background Model Split — the same multi-agent split, but for temporal concerns (stay…
- Client-Side Agent Optimization
Interaction Background Model Split — another axis of multi-model design: there cost-driven and…
- Content-Driven Intervention
nemotronlabs voicechat — Balam, Bartley, Casanova et al. (NVIDIA), arXiv 2609.21967, 2026-09-18…
- Deep Modules for Agents
Interaction Background Model Split — the async background model is a deep module hiding reasoning…
- Encoder-Free Early Fusion
Interaction Background Model Split — the other half of the architecture
- Interaction & Multimodal
Interaction Background Model Split — Dual-model architecture: a time-aware interaction model stays…
- Thinking Machines Lab
Inkling (July 2026) — their first from-scratch model, released with full weights: 975B/41B-active…
- Time-Aligned Micro-Turns
Interaction Background Model Split — micro-turns keep the interaction model present; deep reasoning…
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Full-Duplex Interaction
Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
- Native Multimodal Modeling: Fusion Depth and I/O Duality
An, Lu, Dong et al. (Tencent Youtu + 5 universities, May 2026) formalize 'native' as two operator definitions — mid-fus…
- Live-Path Minimalism
GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; deleg…
