Sources#
- A frontend-backend architecture for tool calls in full-duplex speech models
- Build more natural voice experiences with GPT‑Live‑1 in the API
- Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- Toward Native Multimodal Modeling: A Roadmap
Summary#
The evaluation surface Thinking Machines Lab uses to argue TML-Interaction-Small is "the first model that has both strong intelligence/instruction following and interactivity." Existing benchmarks barely cover interactivity, so TML uses the few that exist plus several new internal ones — and notes "no existing model can meaningfully perform" the visual-proactivity tasks. A research grant for interactivity benchmarks was announced.
Existing benchmarks used#
- FD-bench (v1 / v1.5 / v3) — one of the few benchmarks intended to measure interactivity. Model is given prerecorded audio and must respond at certain times; scenarios: user interruption, user backchannel, talking to others, background speech. v1 reports turn-taking latency; v1.5 an average score; v3 response quality / Pass@1 with tools.
- Audio MultiChallenge — common benchmark for intelligence + instruction following over audio (APR metric); baselines reported by Scale AI.
- BigBench Audio, IFEval (VoiceBench), IFEval (text), Harmbench (text refusal rate) — standard intelligence/IF/safety checks across modalities.
- QIVD (Qualcomm IVD) — video+audio QA, evaluated in a streaming setting (raw clip from the beginning, grade the transcript; GPT-4o-mini grader, following Qwen 3.5 Omni).
Headline results (TML-Interaction-Small, May 2026)#
- FD-bench v1 turn-taking latency: 0.40s — best responsiveness of any model compared (GPT-realtime-2.0 minimal 1.18s, GPT-realtime-1.5 0.59s, Gemini-3.1-flash-live minimal 0.57s).
- FD-bench v1.5 average: 77.8 vs ~39–54 for all baselines (including thinking-high models).
- FD-bench v3 (audio+tools): 82.8% response quality / 68.0% Pass@1 with background agent enabled — best on both.
- Audio MultiChallenge APR: 43.4% — beats every non-thinking baseline; only GPT-realtime-2.0 xhigh (48.5%) is higher.
- Claim: "dominates interaction quality while being more intelligent than any non-thinking model."
* in the table = result reported with the background agent enabled (used for benchmarks that need reasoning or tool calls).
New internal benchmarks (proactive audio)#
- TimeSpeak — can the model initiate speech at user-specified times with correct content? E.g. "remind me to breathe in and out every 4 seconds until I ask you to stop."
- CueSpeak — does the model speak at the appropriate moment with the semantically correct response? Entries constructed so the model must speak at the same time as the user. E.g. "every time I codeswitch, give me the correct word in the original language."
Both: one expected semantic response + timing window per example; LLM judge; correct only if meaning and timing are right; macro-averaged accuracy.
New internal benchmarks (visual proactivity)#
Adapted from existing video benchmarks; "no existing model can meaningfully perform any of these" — they stay silent or answer wrong, including thinking-high models:
- RepCount-A — repeated-action videos, adapted to online rep-counting; stream video after "count out reps for {action}"; graded by whether the last number said after the penultimate rep is within one rep of ground truth. Measures continuous visual tracking + timely counting.
- ProactiveVideoQA — videos with questions whose answers become available at specific moments; turn-weighted PAUC@ω=0.5 (0–100), averaged across turns/categories; staying silent scores 25.0; correct answers must come at the correct time, wrong answers penalized.
- Charades — temporal action localization; "say 'start' when the person starts {action}, say 'Stop' when they stop"; graded by temporal IoU between predicted and reference intervals.
The tool-calling half of the surface (September 2026)#
Everything above measures interaction. A second cluster of benchmarks measures whether a voice agent can actually do anything, and Hu et al. (NVIDIA, arXiv 2609.19334, 2026-09-16, empirical) run four of them on one system, which is the first time this wiki has them together. Their motivating number comes from τ-Voice (arXiv 2603.13686): leading commercial duplex voice models complete 31–51% of grounded customer-service tasks under clean conditions, against 85% for GPT-5 on the text versions of the same tasks — and the gap widens under noise and accented speech. Spoken interaction and tool-call competence are, as of this evidence, separate strengths.
- BFCL-audio (ServiceNow-AI's audio rendering of the Berkeley Function-Calling Leaderboard v3 single-turn sets, scored by AST accuracy): Simple / Multiple / Parallel / Parallel-Multiple / Irrelevance. NVIDIA's best configuration averages 74.6 against GPT-realtime's 80.8; the gap is concentrated in Parallel-Multiple (61.1 vs 74.0) and Irrelevance (81.2 vs 90.8). The cautionary row is Ultravox-v0.6 Llama-3.1-8B at 0.0 on Irrelevance — it invoked the single offered function on all 240 prompts, so its 43.3 average is an artifact of a model that cannot decline.
- Full-Duplex-Bench v3 (arXiv 2604.04847): 100 scenarios from 12 speakers, real human recordings on everyday microphones with pauses, hesitations and self-corrections, mock APIs, GPT-4o judge. Reports tool accuracy, argument accuracy, Pass@1, response quality, and — uniquely — turn-taking, interruption and filler rates. This is the benchmark that penalises a duplex system for its duplex behaviour: NVIDIA scores 100% turn-taking and a 83–85% filler rate, the latter "by design" because the frontend says a hold phrase before delegating, and a 51–54% interruption rate because it backchannels into the user's disfluencies.
- EVA-Bench (arXiv 2605.13841): 213 grounded multi-turn customer-service scenarios across airline, ITSM and medical-HR domains, with a simulated caller and GPT-5.2 as judge. EVA-A averages Task and Faithfulness; EVA-X averages Progress, Conciseness and Speakability. The domain breakdown is the interesting part — NVIDIA's 235B-backend system beats GPT-realtime-2 on airline task completion (72% vs 54%) and loses badly on ITSM (56.3% vs 75%) and medical HR (49.4% vs 69.9%), and the authors attribute the split to those domains needing chains of 6–8 successful tool calls, i.e. to the frontend's delegation token having to fire reliably many times across one conversation. Depth of tool chain, not task difficulty, is what separates the systems.
- Full-Duplex-Bench v1 reappears here in its Candor-subset form (smooth turn-taking TOR and latency, pause TOR, user-interruption TOR / GPT score / latency), used as an ablation surface rather than a leaderboard — see Live-Path Minimalism.
A vendor scoreboard on the same surface (OpenAI, September 2026)#
OpenAI's GPT-Live-1 API launch (2026-09-10, vendor-claim) publishes seven benchmark cards on exactly this surface, which makes it the first time a shipping commercial voice product has been scored against FD-bench in public. It is catalogued here as claims, at a tier below everything else on this page: OpenAI measuring OpenAI, against only its own prior models, with no third party in the comparison and no protocol published.
| Benchmark (metric) | gpt-live-1 | gpt-realtime-2.1 | gpt-realtime-2 | backend noted |
|---|---|---|---|---|
| Tau3 Voice intelligence (Pass@1) | 86.2% | 45.7% | 42.4% | Astra (medium) |
| Tau Banking Voice knowledge (Pass@1, 97 tasks) | 32.0% | 12.4% | 10.3% | Astra (medium) |
| Artificial Analysis Conversational Dynamics (avg) | 97.3% | 95.7% | 95.3% | — |
| Full Duplex Bench v1.5 Interactivity (avg) | 80.10% | 45.4% | 47.8% | — |
| Full Duplex Bench v1 turn-taking latency ↓ | 0.798 s | 1.41 s | 1.63 s | — |
| Full Duplex Bench v3 tool calling (Pass@1) | 87.0% | 60.0% | 58.0% | Terra (low) |
| Full Duplex Bench v3 response quality | 90.0% | 88.0% | 81.0% | Terra (low) |
Four things worth recording, in descending order of how much they survive the tier discount.
The backend footnotes are the most useful part of the table, and OpenAI put them there itself. Four cards state the backend and reasoning effort behind the GPT-Live configuration; three do not. The starred rows are therefore scores for a pairing, and the vendor picked different backends at different efforts for the two task families — Astra at medium for intelligence, Terra at low for tool calling. That is the *-convention this page already uses for TML, and the capability-interface property NVIDIA measured, now visible in a vendor's own footnotes (Interaction / Background Model Split).
The frontend-only rows move very unevenly, which is the informative pattern. FDB v1.5 Interactivity 45.4 → 80.10 and v1 latency 1.41 → 0.798 s are large; Artificial Analysis Conversational Dynamics 95.7 → 97.3 is not — a benchmark on which the previous generation already scored 95.3 has little room to distinguish anything, and its inclusion illustrates the saturation problem rather than a capability gap.
A rare partial cross-lab consistency check, on the baselines rather than the headline. This page records τ-Voice finding leading commercial duplex voice models at 31–51% on grounded customer-service tasks against 85% for GPT-5 on the text versions. OpenAI's own baselines land inside that band — gpt-realtime-2.1 at 45.7% and gpt-realtime-2 at 42.4% Pass@1 on Tau3's airline/retail/telecom tasks — which is two independent parties placing the GPT-realtime family in the same range on similar task families. Against that, OpenAI claims 86.2% for gpt-live-1 + Astra, i.e. that the voice-agent-versus-text-agent gap τ-Voice measured is now closed. If it holds, it is the single most consequential number in this section; it rests entirely on a vendor's release post, on a benchmark (Tau3) with no independent run in this corpus, and on a configuration that includes a frontier text backend doing the reasoning — so it is not evidence that duplex voice models closed the gap, only that a duplex frontend plus a frontier backend can score like a text agent.
The prose headline does not reconcile with any published card. OpenAI's text claims GPT-Live-1 "improves Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1." No card shows +30: v1.5 Interactivity is +34.7, v3 tool calling +27.0, v3 response quality +2.0, and v1 is a latency metric. The headline is an unstated aggregate over an unstated set; quote the individual cards, not the 30.
Not on this page and not reconcilable here: the customer-relayed figures — Speak's "almost 80%" reduction in interruptions versus unspecified "previous turn-based systems," and Yelp's improved call-handling rates — are business outcomes with no benchmark protocol behind them, recorded on GPT-Live.
A second vendor scoreboard, and this one is drawn on other people's boards (Google, September 2026)#
Google's Gemini 3.8 Live launch (2026-09-15, vendor-claim) arrives five days after OpenAI's and argues itself entirely on charts, five of them — and the rhetorical footing is the opposite of the section above. OpenAI scored itself against only its own prior models; Google reproduces third-party boards that already contain its competitors: three titled Artificial Analysis, one Sierra, one ServiceNow. That is a better shape for an outside reader and it is still a vendor-selected, vendor-transcribed subset of somebody else's leaderboard, with no AA, Sierra or ServiceNow publication in this corpus to check it against. Full product capsule on Gemini 3.8 Live.
Speech to Speech Index (Artificial Analysis) — the first cross-vendor voice index this wiki has seen. Every bar is labelled with a reasoning-effort setting:
| Model | Effort | Score |
|---|---|---|
| Gemini 3.8 Live Extended Thinking | High | 82.6% |
| GPT-Live-1 Astra | Medium | 81.5% |
| Grok Voice Think Fast 2.0 | High | 81.3% |
| Gemini 3.8 Live | — | 76.0% |
| Gemini 3.1 Flash Live | High | 71.5% |
| Gemini 3.1 Flash Live | Minimal | 63.9% |
Agentic Performance, τ-Voice (Artificial Analysis) — ET 68.6 / GPT-Live-1 Astra (Medium) 67.9 / Grok Voice Think Fast 2.0 (High) 56.5 / Gemini 3.1 Flash Live (High) 37.7 / Gemini 3.8 Live 30.1 / Gemini 3.1 Flash Live (Minimal) 26.2.
τ³-Banking Leaderboard (Sierra) — ET 35.1 / GPT-Live-1 Astra (Medium) 32.0 / xAI-Realtime 16.5 / Gemini 3.1 Flash Live (High) 11.3 / GPT-Realtime 2 (High) 10.3. Google's prose calls this "Sierra's τ-Voice-banking benchmark"; the chart title is "τ³-Banking Leaderboard". Use the chart's name.
Cost per hour of input audio, on a Big Bench Audio subset (Artificial Analysis) — Gemini 3.8 Live $0.84 / Gemini 3.1 Flash Live Minimal $1.50 / High $1.75 / Gemini 3.8 Live ET (High) $3.50 / Grok Voice Think Fast 2.0 $4.80 / GPT-Live-1 Astra (Medium) $5.83. Two things about this chart specifically. It is a measured cost on a workload, not a price list, which is the artifact Compute-Controlled Benchmarking keeps asking vendors for and is rarely given. And its denominator is an hour of input audio — conversation duration, not tasks or tokens — which is the third-party echo of the per-minute billing unit OpenAI introduced (Cost-per-Task Over Cost-per-Token). It is also not the Big Bench Audio accuracy chart it looks like: Extended Thinking's 97.7% Big Bench Audio appears in Google's prose and on no chart at all.
EVA-Bench (ServiceNow), and it is the chart that does not flatter. A scatter of Accuracy (EVA-A pass@1) against Experience (EVA-X pass@1) — the same two-axis decomposition NVIDIA's numbers above use — run, per the chart's own footer, "on the Live API on Gemini Enterprise Agent Platform." No data labels are printed, so positions are eyeballed to ~1–2 points and no number from it should be quoted; the ordering is unambiguous and it is the point:
- Google's prose: "our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality."
- The dotted frontier drawn on Google's own chart runs Gemini 3.8 Live (~47, ~89) → Extended Thinking (minimal) (~51, ~79) → GPT Realtime 2.1 (~58, ~74) → Scribe Realtime + GPT 5.4 + Eleven Flash v2, ElevenAgents tuned (~67, ~14). Two of the four points on the frontier belong to competitors, one of them a cascaded third-party stack.
- Extended Thinking (high) (~55, ~73) sits below the line — dominated by GPT Realtime 2.1, which has more accuracy at the same experience. Google's flagship configuration is off the frontier its own chart draws.
The cascade comparators are the other reason to keep this chart. The same ElevenAgents stack appears twice, in opposite corners: default (Scribe Realtime + Claude Haiku 4.5 + Eleven Flash v2) at ~37 accuracy / ~82 experience, and tuned (Scribe Realtime + GPT 5.4 + Eleven Flash v2) at ~67 / ~14 — the highest accuracy anywhere on the chart bought by giving up essentially all of the experience axis. Accuracy and conversational quality are separately purchasable within a single stack, and tuning moves you along the trade rather than off it. That is the strongest version of the duplex-versus-cascaded separation on Full-Duplex Interaction, drawn by a vendor with no interest in drawing it.
Cross-checking Google's charts against OpenAI's cards#
The two vendor scoreboards overlap on two rows, and they behave differently — which is the useful part.
The banking row matches digit for digit, across two adversarial parties. OpenAI's own card reports gpt-live-1 at 32.0% and gpt-realtime-2 at 10.3% on "Tau Banking Voice knowledge (Pass@1, 97 tasks)". Google's Sierra chart reports GPT-Live-1 Astra (Medium) at 32.0% and GPT-Realtime 2 (High) at 10.3%. Identical on both rows. The inference: both vendors are quoting the same Sierra τ³-Banking leaderboard rather than running the benchmark themselves, and OpenAI's "Tau Banking Voice" is that board under another name. This is the first cross-vendor corroboration of any number in this corpus, and it is corroboration of provenance — two parties reading one board — not two independent runs agreeing.
The agentic row does not match, and should not be reconciled. OpenAI reports gpt-live-1 + Astra (medium) at 86.2% on "Tau3 Voice intelligence (Pass@1)". Google reports GPT-Live-1 Astra (Medium) — the same product, the same named backend, the same stated effort — at 67.9% on "Agentic Performance (τ-Voice)". Because the banking rows matched exactly, this is unlikely to be a transcription artifact on either side. The honest reading is that Tau3 and τ-Voice are not the same instrument: different task sets, different versions, or different harnesses under similar-looking names, with neither party publishing a protocol. Do not rank them, do not average them, and do not treat either as a check on the other.
One band this does move. This page records τ-Voice's original finding — leading commercial duplex voice models at 31–51% on grounded customer-service tasks against 85% for a text agent. On Google's chart the top two systems sit at 68.6% and 67.9%, above that band by ~17 points six months later. Both are high-or-medium-effort configurations with a frontier text model named in the bar label, so the band has been exceeded by the pairing, not by a duplex voice model as such — the same caveat this page attached to OpenAI's 86.2%.
What a composite index can and cannot settle#
Google ships two tiers on one live architecture, which is the cleanest within-vendor test yet of "interactivity scales with intelligence" — and the charts answer less than they appear to.
- The composite compresses. 3.8 Live → Extended Thinking moves the S2S Index 76.0 → 82.6, +6.6 points. The one component visible separately moves 30.1 → 68.6 on τ-Voice, +38.5 points and a 2.3× multiple. Nearly all of the Extended Thinking premium lives in the agentic component, and the headline index hides it.
- The composite and its component disagree about the ordering. 3.8 Live beats Gemini 3.1 Flash Live (High) on the index (76.0 vs 71.5) and loses to it on τ-Voice (30.1 vs 37.7). A single number and one of its own inputs rank the same two models in opposite orders.
- The mechanism is unverifiable from here. Neither the raw nor any ingested source states what the Speech to Speech Index is composed of or how its parts are weighted; the methodology footnote on all five charts points at
deepmind.google, the vendor's page, not the evaluator's. A near-saturated interaction component would produce exactly this compression pattern — and this page already records one AA voice component sitting at 95.3–97.3 across three model generations — but that is a hypothesis consistent with two vendors' charts, not something either source establishes.
The transferable rule: a vendor composite can settle that a more expensive tier scores higher; it cannot separate intelligence from interactivity, because the separation lives in the component weights and no vendor publishes them. Ask for the components.
Two disclosure notes#
Effort labels are present and asymmetric. Every bar on all four bar charts carries a reasoning-effort setting, which is more compute disclosure than most cross-vendor grids offer (Compute-Controlled Benchmarking) — and Google's flagship is charted at High while the nearest competitor is charted at Medium, with no note explaining the choice. It may simply be GPT-Live-1's published configuration (OpenAI's own cards also footnote Astra, medium). Either way, a cross-vendor chart in which the home model runs at a higher setting than the comparator is not a controlled comparison, and the reader is told nothing.
No latency number appears anywhere in the Google post — not turn-taking latency, not time-to-first-audio — despite partner quotes citing "impressive latency." So nothing from this source joins the three-way non-reconciliation below, and the voice product with the most third-party chart coverage in the corpus contributes zero responsiveness data.
A proactivity instrument that controls the opportunity (September 2026)#
Everything above measures whether a model responds well. Peng, Nuchged, Fu & Yao (UConn / UT Austin / X Square Robot / Oxford, arXiv 2609.19596, 2026-09-17, empirical) add an instrument for whether it responds at all, unasked — and it is built differently from everything on this page, which is why it is worth cataloguing as a protocol rather than as another leaderboard.
The design move is subtraction, not addition. Every benchmark above enriches the stimulus; this one impoverishes it:
- Context-matched stimuli. 40 first-person English monologues (median 50.3 s), lead-in + trigger + continuation, where within a topic only the trigger utterance changes. Every comparison is therefore within-topic and within-voice, which removes the content confound that an assembled scenario set carries.
- Opportunity removed by construction. Word-level timestamps compress every inter-word gap to ≤0.12 s, and the user keeps talking after the trigger. Taking the floor means interrupting. A natural-pause arm (typical sentence-final pause ~0.88 s) and an inserted-1.5 s-silence arm restore the opportunity as manipulations, so "reason to speak" and "chance to speak" are separately controlled — the thing FD-bench, FLEXI and Instruct-FD all leave entangled.
- Conditions from turn-allocation theory, not from product scenarios. Ten conditions grouped as Addressed (Question, Request, Rhetorical), Self-select (Word search, False fact, Repeat, Hazard) and Yield (Turn end, Silence), plus Neutral — derived from Sacks, Schegloff & Jefferson (1974). Rhetorical and Turn end exist purely as surface-form controls: they carry the form of an invitation without its substance, and a model that responds to them is responding to syntax.
- A single scalar with a stated denominator. Speech-onset rate = share of trials with a ≥4-word segment in [t₀, t₀+4 s); trials already speaking in the preceding 1.5 s stay in the denominator scoring no new onset. Segments are defined per architecture (80 ms text-token frames with no gap >0.64 s for Moshi/PersonaPlex/Raon; turn-boundary markers for VoiceChat; 1 Hz listen/speak outputs for MiniCPM-o), which is what lets one number span five families.
- A second, independent arm for response content. A turn-release experiment hands the floor over (10 s of silence) and scores Relevant and Intervenes with a DeepSeek-V4-Flash judge at temperature 0 — so a positive onset effect can be checked against whether the resulting speech was any use. It is the check word search fails: the only content cue that raises PPlex's onset, and only.36 of its replies actually supply the missing word.
Scale: 7 configurations × 40 topics × 10 conditions × 5 seeds = 2,000 trials per configuration. Stimuli, model outputs and evaluation code are released (github.com/vocaliodmiku/take-the-floor), which makes this the first instrument on this page that a third party can re-run against any duplex system — including the ones whose proactivity is currently a vendor-claim.
Headline numbers and what they mean are on Content-Driven Intervention; the short version is that being addressed and an actual pause move onset, and false facts, hazards and verbatim repetition move it by at most.06.
How it compares to CueSpeak and TimeSpeak. Both of TML's proactive-audio benchmarks test speaking at the right moment in service of a standing user instruction ("every time I codeswitch, give me the correct word"; "remind me to breathe every 4 seconds"). That is the Addressed case in this protocol's taxonomy — the one every model here already handles best. Neither measures self-selection, where no instruction covers the trigger and the model must decide the content warrants a turn. The two instruments are complements, and the gap between them is exactly where the negative result lives.
Evidence note. Third-party, no vendor affiliation among the four authors, on open checkpoints at default decoding — a different footing from the rest of this page, where TML scores TML and NVIDIA scores NVIDIA. The limits are elsewhere: single TTS stimulus pipeline (two synthetic voices, no human speakers, unlike FDB3's real recordings), English only, no inter-rater or human validation of the DeepSeek judge, and no confidence intervals on any rate.
Three cross-source latency numbers that do not reconcile, and should not be forced to. This page records TML-Interaction-Small at 0.40 s FD-bench v1 turn-taking latency, "best responsiveness of any model compared." NVIDIA's frontend reports 92 ms smooth turn-taking latency on Full-Duplex-Bench v1 (221 ms for its no-tool-call baseline). OpenAI now reports 0.798 s for gpt-live-1 on Full Duplex Bench v1, "how quickly the agent starts its reply after the user finishes a turn." All three are self-reported, and they are almost certainly not the same measurement: NVIDIA specifies the Candor subset and reports TOR alongside, TML's table does not say which subset or which latency definition, OpenAI states neither subset nor definition beyond that one sentence, and the three systems are scored on different output modalities (agent text for NVIDIA, audio for OpenAI, unstated for TML). A naive ranking would put the shipping product 8.7× slower than a research frontend and 2× slower than a research preview, which is not a conclusion any of these tables supports. Treat all three as non-comparable until one lab publishes its FD-bench v1 protocol. The comparisons that are legitimate are the within-source ones: OpenAI's 1.41 → 0.798 s against its own prior model, and NVIDIA's 221 → 92 ms against its own ablation baseline.
Evidence weighting across this page. TML's numbers are first-party vendor-claim on a research preview with self-invented benchmarks; NVIDIA's are empirical on public third-party benchmarks with baselines the authors re-ran themselves, still single-run and with no error bars anywhere; OpenAI's are first-party vendor-claim on a shipping product against a baseline set containing nothing but its own two previous models. Peng et al.'s are the only third-party numbers on the page. Google's are a fourth shape again — first-party vendor-claim selection and transcription of third-party boards (Artificial Analysis, Sierra, ServiceNow) that do contain competitors, which is better than self-versus-self and is still the ranked party choosing which boards to show, at which effort setting, with the evaluator's own publication absent from this corpus. None of the four vendor sources is independent evaluation.
An outside inventory of the same surface (May 2026)#
An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities, arXiv 2605.25343, practitioner-opinion) catalogue 39 multimodal benchmarks (Table 3: 18 image, 7 audio, 14 video; rebuilt from pdftotext -layout page 29 during compile and reconciled row-for-row). It measures nothing itself, but it is the first source here to place the interactivity surface inside the general multimodal evaluation landscape, and it names instruments this page did not have.
The full-duplex row confirms the vault's numbers from a neutral source. Moshi Eval at a 200 ms target latency, SoulX-Duplug-Eval at 240 ms bilingual streaming turn detection, and Full-Duplex-Bench scoring turn-taking, barge-in handling and false-interruption rate — the same three dimensions FD-bench covers, from a survey with no product in the race. It adds FireRedChat and ELLSA as cascaded-vs-end-to-end full-duplex evaluations and an LLM-enhanced dialogue-management line testing turn prediction.
Three instruments for the proactivity gap this page records as unmeasured. The survey's video section names, beside the offline suites, a streaming cluster: OVO-Bench (real-time perception plus backward tracing, testing whether a model identifies the appropriate temporal moment to respond), StreamingBench (comprehension under strict latency constraints), OmniMMI (multimodal streaming interaction), and two protocols that are close to what TML's visual-proactivity benchmarks were invented for — ThinkStream's Watch-Think-Speak, which judges "not only on answer accuracy but on response timing — whether sufficient evidence has accumulated before responding," and AURA, which extends streaming evaluation to proactive QA (responding when relevant events occur with no explicit query) and multi-response QA (tracking evolving events over time). Proactive QA with no standing instruction is precisely the construct Content-Driven Intervention shows is missing from RepCount-A / ProactiveVideoQA / Charades, which all supply the instruction and then measure timing. Caveat before anyone treats these as the answer: the survey reports no numbers for any of them and this vault has read none of the underlying papers, so what is established is that the instruments exist, not that they work.
A fourth evaluation axis the vault has no page for: efficiency-aware protocols. The survey's position is that accuracy-only scoring "ignores the dimensions that matter most for native deployment: when a model responds, how much compute it consumes, and how gracefully it handles streaming and interruption," and it wants accuracy reported alongside token budget, latency and energy on an explicit Pareto frontier. Its exhibit is ResAdapt, reportedly eliminating over 90% of visual tokens while processing 16× more frames for >15% relative gains on complex long-video reasoning. That is a third-party restatement, not a measurement made here.
What it does not have. No benchmark in Table 3 scores a symmetric model on aligned understanding-generation pairs — the survey lists that as an open direction ("describe-then-render, listen-then-speak, watch-then-act," penalizing inconsistency across the two directions), alongside robustness under multimodal attack surfaces. And no production live-voice system appears in the census at all, so none of the vendor scoreboards above has a counterpart here.
FDB 1.0 used as a leaderboard for the first time, and what that exposes (NVIDIA, September 2026)#
NemotronLabs VoiceChat (NVIDIA, arXiv 2609.21967, 2026-09-18, empirical) is the first source in this corpus to publish Full-Duplex-Bench 1.0 as a cross-system leaderboard rather than as an ablation surface (Hu et al.) or a vendor card (OpenAI). Three tracks, official scoring protocol: pause handling on synthetic stimuli and on natural CANDOR pauses (Takeover Rate, lower better), smooth turn-taking (TOR and latency), and user interruption (TOR, latency, and GPT-4o-judged post-interruption response quality on 0–5). The full seven-row table is on Full-Duplex Interaction; what belongs here is the instrument.
Four provenance classes in one seven-row table, under one footnote. This is a new shape for this page's catalogue and worth naming, because the table reads as a single controlled run and is not one:
- "From the public FDB results" — Moshi, Freeze-Omni, Gemini Live 2.0, GPT-Realtime. Copied from the benchmark authors' own publication.
- Reported by the system's own authors — PersonaPlex, "the publicly released checkpoint evaluated by its authors."
- Explicitly outside the controlled run — MoshiRAG, daggered, "reported by its authors; this separate evaluation was not part of the controlled FDB run."
- NVIDIA's own fresh measurement — one row, its own.
So NVIDIA re-ran none of its five open-weight comparators, and the "lowest pause-handling TOR among evaluated open-weight systems" claim is a comparison between one new number and five imported ones. The paper's footnote is honest about each piece and the abstract compresses all of it into "evaluated."
Endpoint staleness is stated and then ignored. The closed rows are gemini-2.0-flash-live-001 and GPT-Realtime, with the footnote conceding that closed-system names "refer to the evaluated historical endpoints, not necessarily their current service versions" — while this page's Google section is scoring Gemini 3.8 Live. An FDB 1.0 comparison against a Gemini Live endpoint three generations old is not evidence about the current closed tier, and GPT-Realtime beats VoiceChat on both pause tracks (1.0 / 12.0) and on smooth-turn TOR (100.0) anyway.
FDB 1.5 gains a five-system table. The user-backchannel condition, classifying the model's reaction as Respond / Resume / Uncertain / Unknown, with Resume desired: Moshi 2 / 6 / 0 / 92, Freeze-Omni 7 / 80 / 2 / 11, MoshiRAG † 5 / 61 / 0 / 34, VoiceChat 1 / 93 / 2 / 4, Gemini Live 2.0 1 / 93 / 2 / 4, GPT-4o Realtime 3 / 70 / 1 / 25. The open-weight winner and the closed comparator agree to the digit on all four categories, and neither paper nor footnote remarks on it; with 100 examples per condition a four-way exact tie is possible but unlikely, and it is the kind of coincidence a benchmark table should be asked about. Moshi's 92% Unknown is the useful floor: a model that essentially never handles a backchannel at all.
FDB 3.0 now has two NVIDIA entries that cannot be compared through a shared baseline. Both papers report FDB 3.0, two days apart, both attribute their baseline rows to the benchmark's own paper (arXiv 2604.04847) — and the two baseline sets do not intersect:
| Baselines shown | Metric name for routing | |
|---|---|---|
| Hu et al. (delegating frontend+backend) | GPT-realtime-mini, GPT-realtime, Gemini-3.5 Flash, Ultravox-v0.6 ×2 | Tool-acc |
| VoiceChat (internalized function channel) | Gemini Live 2.5, Gemini Live 3.1 | Tool Sel. F1 |
Neither table contains the other NVIDIA system, no baseline row appears in both, and the first column carries a different metric name in each. The numbers (VoiceChat 82.5 / 42.2 / 33.0; Hu et al.'s best 71.7 / 55.2 / 48.0 / 67.0) are worth comparing and the comparison has to be made by the reader, across two papers from one lab, with the accuracy-vs-F1 mismatch unresolved. Substantive reading on Interaction / Background Model Split. The FDB 3.0 scoring rule that makes the spread legible: Pass@1 counts a sample only when the model selects exactly the expected tools and supplies perfect arguments for every call, so a 82.5 routing score beside a 33.0 Pass@1 localizes the failure in argument extraction rather than tool choice.
VoiceBench enters the page as the intelligence axis. Nine subsets — OpenBookQA, MMSU, BBH (multiple choice), SD-QA (free-form against references), CommonEval / AlpacaEval-Full / WildVoice (open-ended, judged 1–5), IFEval, AdvBench — with everything but the three open-ended subsets on 0–100 and a normalized average across the mixed scales. VoiceChat scores 55.1, effectively tied with Freeze-Omni's 55.2 on a very different profile (OpenBookQA 61.3 vs 31.0, MMSU 46.1 vs 28.1, AdvBench 100.0 vs 97.3, and worse on SD-QA, IFEval and all three open-ended subsets), against Moshi 29.5, PersonaPlex 30.6, the cascaded DuplexCascade 65.4 and the omni-modal MiniCPM-o 4.5 76.1.
That last row carries the caveat this page collects: MiniCPM-o's 76.1 is from an independent third-party evaluation that used GPT-5.4 to judge the three open-ended subsets, and the table's own footnote says its "judge-based scores and aggregate are therefore not directly matched to the official-leaderboard evaluation." A leaderboard row whose aggregate is computed under a different judge, printed in the same column as official-leaderboard rows and flagged only in a footnote, is exactly the composition problem LLM-as-a-Judge tracks — and here the flagged row is the top of the table.
A latency/accuracy dial published as a curve, in the ASR dimension. On the Hugging Face OpenASR Leaderboard's eight sets, the same cache-aware streaming encoder is evaluated at two chunk sizes with no retraining: average WER 9.02% at 80 ms and 8.28% at 160 ms, improving on all eight sets (AMI 15.32 → 13.54, LS Clean 3.58 → 3.19, VoxPopuli 8.85 → 8.17). Both column averages reproduce from their own eight rows to two decimals. This is small but it is the right shape — a deployment-time knob whose cost is published on both axes, which is what Compute-Controlled Benchmarking asks vendors for and rarely gets; the model card ships one checkpoint that can be operated anywhere on the curve.
Where this sits in the page's evidence ordering. NVIDIA's numbers are empirical on public third-party benchmarks, which puts them above the vendor scoreboards above — but the open-weight comparison set is entirely imported, there is one run and no error bars anywhere in the paper, and it is NVIDIA scoring NVIDIA. It sits below Peng et al. (the only fully third-party instrument on this page) and beside Hu et al.
A benchmark whose stimuli are generated, not recorded (Alibaba Qwen, September 2026)#
OmniVChat (He, Chu, Chen et al., 18 authors, CUHK + Alibaba Token Hub / Qwen Team + SJTU + Zhejiang, arXiv 2609.21465, 2026-09-18, empirical) is the first benchmark on this page whose entire stimulus set is synthesized by generative models. It is also the first that is deliberately not live: all instances are pre-rendered clips, scored offline, with no latency metric, no barge-in and — the paper says so in as many words — a recorded probe that "does not test live interruption" and covers "neither multi-turn dialogues nor real-time latency during live use." Read it as the offline audio-visual-comprehension slice of the interactivity surface, not as another duplex leaderboard.
The task definition. OmniVChat = an omni model receives audio and video simultaneously and returns text, with the user's query embedded in the two streams and no separate text question, no external captioning, no ASR. Architecture and fusion placement on Native Multimodal Modeling: Fusion Depth and I/O Duality; what matters here is that the constraint is applied to the whole comparison set — "all systems receive the required audio and video inputs" — with one deliberate exception, JoyAI-VL-Interaction, run with Qwen3-ASR as an explicit ASR-cascade arm. So the native-input rule is not a rule the proposer alone obeys.
Scale and shape. 2,800 synthesized instances: 2,550 single-turn + 250 multi-turn, 17 subcategories (12 single-turn plus 5 multi-turn extensions of them), 5 ability categories, 22 scenario domains, 1,766 English (63.1%) / 1,034 Chinese (36.9%). Per-subcategory counts read off Figure 3 — 200 each, except AR 300 + DR 300 (the paper's "600 combined"), ER 150, and 50 for each multi-turn extension — sum to exactly 2,800 and to the stated 2,550 single-turn, which is what confirms the figure↔count mapping. The five ability categories: DSLP (dialogue state and link perception — connection checks, turn boundaries, camera orientation), MEA (multimodal entity alignment), MSA (model self-awareness — stating identity and physical limits), AH (anti-hallucination), ER (emotion recognition).
The scoring rule is the contribution, more than the task set. Each instance carries a reference reply and a tiered rubric. Tier 0 is a language gate: it earns no points, and failing it makes the whole instance score zero. Outside Tier 0, a met criterion earns one point only if every criterion in every earlier tier was met; an incomplete tier keeps the points it earned and blocks all later tiers. The denominator counts every non-Tier-0 criterion including the ones blocked by an earlier failure, so the score is in [0,1] and a model that nails Tier 3 while missing Tier 1 gets nothing for it. The reported overall is Subcategory Mean (average within each of the 17 subcategories, then average the 17 equally) rather than a pooled per-instance mean — so the 50-instance multi-turn subcategories carry the same weight as the 300-instance AR. The judge is qwen3.7-max, and per Table 4 it is a text-only grader: "Text LLM, Video Unseen." It never sees the clip; it checks the reply against criteria the data engine wrote. Judge-side scrutiny on LLM-as-a-Judge.
Headline table (Table 1 — 12 released systems + the trained model + 2 ablations, one shared 1,296-character system prompt; reconciled cell-for-cell against pdftotext -layout p.9–10, clean).
| Mean (Bench) | Human | RE | Style | |
|---|---|---|---|---|
| Gemini-3.5-Flash | 0.667 | 0.509 | 10.59 | 0.755 |
| Gemini-3.7-Flash | 0.640 | 0.575 | 16.96 | 0.631 |
| Gemini-3.1-Pro | 0.607 | 0.537 | 11.66 | 0.871 |
| Qwen3-Omni-Instruct (base) | 0.465 | 0.402 | 5.75 | 0.710 |
| OmniVChat-RL (theirs) | 0.652 | 0.632 | 18.38 | 0.992 |
| ablation: no efficiency term | 0.697 | 0.691 | 7.12 | 0.987 |
Mean is Subcategory Mean on OmniVChat-Bench; Human is the 360-recording probe; RE is rubric credit per thousand counted words; Style is the all-seven pass rate against the system prompt's own style criteria.
Three things this table says that its abstract does not.
- Synthetic rank does not determine recorded rank, inside one vendor family. Gemini-3.5-Flash leads the synthetic Mean (0.667) and Gemini-3.7-Flash leads the recorded Human (0.575) while sitting third on Mean. The paper names this itself. It is the cleanest datum on this page for the proposition that a synthesized benchmark and a recorded probe of the same subcategories order models differently — and it arrives from the lab that built both.
- The configuration they ship is not the configuration that scores best. The
no efficiency termablation beats the shipped model on Mean (0.697 vs 0.652) and on the recorded probe (0.691 vs 0.632); what it loses is brevity — replies grow 36 → 99 words and RE collapses 18.38 → 7.12. So the headline "0.465 → 0.652" is the score of a model deliberately handicapped for concision, and rubric-only RL gets further on both rubric axes. Honest, and easy to miss. - Style 0.992 is the metric being trained on. The style reward during RL and the Style column at evaluation are the same seven criteria graded by the same
qwen3.7-maxprompt. A 0.710 → 0.992 move on a metric that is the reward is not evidence about style; it is evidence the optimizer works. The rubric score has the same property one step removed (same judge, same rubric form), which is why the recorded probe is load-bearing for this paper and the Style column is not. Adherence-side reading on Harness Activation and Adherence.
The measurement-noise result, which is the most transferable thing in the paper. Appendix E.5 pools the three held-out curves (dev, Bench, Human probe) over 51 checkpoints. Levels correlate strongly — r(dev, Bench) = 0.983, r(dev, probe) = 0.963. Checkpoint-to-checkpoint changes barely correlate at all: r(Δdev, Δbench) = 0.291, r(Δdev, Δprobe) = 0.087, with the signs agreeing on 31, 27 and 28 of 50 intervals — a sign test cannot reject chance. Residual σ̂ is 0.0066 (dev), 0.0044 (Bench), 0.0113 (probe) against a mean 20-step gain of 0.0039–0.0044, i.e. short-term variation is 1.1–2.6× the per-checkpoint improvement. Averaging 5 checkpoints (100-step blocks) restores the correlations to 0.905 / 0.933 with 8-of-9 matching signs; at k = 10 they run 0.966–0.996. The operational rule: on a rubric-judged benchmark of this size, a single-checkpoint delta is noise and only a multi-checkpoint block mean is a measurement — which is the quantitative form of the "one run, no error bars" complaint this page files against every other source on it.
Multi-turn costs less than it looks, and the comparison is not paired. Five subcategories have matched single-turn and multi-turn versions. Gemini-3.7-Flash loses on all five, mean gap −0.114 (t(4) = −4.48, p = 0.011), worst on MRR at −0.202 (recall an object after it leaves frame). OmniVChat-RL's mean gap is −0.026 (t(4) = −0.85, p = 0.45). The paper immediately disarms its own result: the pairs are independently sampled, 50 instances each resolve a gap only to ±0.05, and four degrees of freedom is four degrees of freedom. Also note the multi-turn protocol hands every model the same earlier clips and reference replies and scores only the final reply — a teacher-forced evaluation with no turn-level credit at stake (see Turn-Level Credit Assignment).
The hardest subcategory is knowing when not to speak. All six judges in the sensitivity study rank DSLP-VTT-UPC — "User Paused but not Completed," where the correct reply is to wait or give a brief acknowledgment — as the weakest subcategory. That is the same construct Content-Driven Intervention finds unmeasured on the audio side, arriving here as the floor of a benchmark that did not set out to study it.
COI, stated plainly. One lab supplies the task definition, the data engine that generates the stimuli, the benchmark, the human probe, the judge model, the base model, the RL reward, and the trained model — and no external benchmark is reported for the trained model at all. What partly offsets that: the 360 recordings are not products of the data engine, the ablations are published and unflattering, the ranking disagreements are pointed out by the authors, and Appendix D.3's judge study says outright that it "does not remeasure the OmniVChat-RL gain from 0.465 to 0.652 with other judges." What does not offset it: every number above comes from an instrument the proposer built, and qwen3.7-max sits on both the reward side and the grading side.
Why it matters#
The argument structure: prior interactivity-oriented benchmarks "do not adequately capture the qualitative jumps" — so the case for interaction models rests partly on benchmarks TML had to invent. This is a known pattern (new capability → new eval), and a soft spot (self-defined metrics); TML's response is to open a research grant inviting community benchmarks.
Connections#
- Native Multimodal Modeling: Fusion Depth and I/O Duality — the outside inventory: 39 benchmarks placing the interactivity surface inside general multimodal evaluation, plus ThinkStream/AURA/OVO-Bench as third-party instruments for the proactivity gap and the efficiency-aware Pareto argument
- Interaction Models — what's being measured
- TML-Interaction-Small — the model under test
- Interaction / Background Model Split —
*results use the background agent - Full-Duplex Interaction — TimeSpeak/CueSpeak/RepCount-A/ProactiveVideoQA/Charades each target one of these modes
- Time-Aligned Micro-Turns — turn-taking latency (0.40s) is the direct payoff of removing turn boundaries
- Scale-Dependent Prompt Sensitivity — another case of a paper introducing its own evaluation framing to surface a phenomenon standard benchmarks miss
- Claude Opus 4.7 —
xhigheffort tier shows up here as a baseline config (GPT-realtime-2.0 minimal/xhigh) - Live-Path Minimalism — where FD-bench v1 is used as an ablation surface for what a delegation token costs a frontend
- NVIDIA — the lab that runs the tool-calling cluster end to end on one system
- Content-Driven Intervention — what the September 2026 proactivity instrument measures, and the negative result it produced
- GPT-Live — the first shipping commercial voice product scored against FD-bench in public, by its own vendor
- Gemini 3.8 Live — the second vendor scoreboard, and the first argued on third-party boards that contain its competitors
- Artificial Analysis — the evaluator behind the Conversational Dynamics card, the Speech to Speech Index, the τ-Voice chart and the cost-per-audio-hour chart, and the only reason any two vendors' numbers here line up
- Compute-Controlled Benchmarking — where the effort-label asymmetry and the measured-cost-on-a-workload chart bear
- Production-Sourced Evaluation — the sourcing axis OmniVChat-Bench sits at the far end of: every stimulus generated by a multi-agent engine, with a 360-recording human probe as the only non-generated slice
- Harness Activation and Adherence — the Style column is a system-prompt adherence rate measured across twelve released omni models, and then trained directly against its own judge
- Turn-Level Credit Assignment — OmniVChat's multi-turn protocol hands every model the same history and scores only the final reply, so no turn credit is at stake
Sources#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- A frontend-backend architecture for tool calls in full-duplex speech models — Hu et al. (NVIDIA), arXiv 2609.19334, 2026-09-16 (
empirical, 5pp): §5.1–5.3 and Tables 1, 2, 3, 7. Tables 1–3 were re-read frompdftotext -f 2 -l 3 -layoutand Table 7 frompdftotext -f 4 -l 5 -layout(itstable-shiftwarning is a spanning-header artifact); every figure above comes from a reconciled cell or from the prose. Full treatment on Interaction / Background Model Split - Full-Duplex Speech Models Take the Floor When Asked, Not When Needed — Peng, Nuchged, Fu & Yao, arXiv 2609.19596, 2026-09-17 (
empirical, 5pp): §3 the stimulus design, condition taxonomy and onset definition; §4.4 the turn-release judging protocol. All four tables re-read frompdftotext -layoutand matching digit-for-digit. Full treatment on Content-Driven Intervention - Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Tom Ouyang & Malini Jaganathan (Gemini Audio Team, Google),
blog.google, 2026-09-15 (vendor-claim, ~1,764 words): the five charts above. Parse note: not PDF-derived; the charts are.webpimages transcribed at ingest intoFigure transcriptionblocks, and all five images were viewed during compile. The four bar charts carry printed data labels and match the transcription digit-for-digit. The EVA-Bench scatter carries no data labels — its values are eyeballed against gridlines to ~1–2 points, so it is cited here for ordering and dominance only, never for a number. One transcription correction was made at compile time: the ingest note describes the dotted line as connecting "the three Gemini 3.8 points", but at 3× crop it connects Gemini 3.8 Live → Extended Thinking (minimal) → GPT Realtime 2.1 → Scribe Realtime + GPT 5.4 + Eleven Flash v2, and leaves Extended Thinking (high) below the line. The raw is immutable, so the correction lives here and on Source Notes. - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim, ~1,670 words, unbylined): the seven GPT-Live-1 benchmark cards and their per-card backend footnotes. The cards were lazily-mounted JS chart components recovered as text values from the rendered DOM rather than transcribed from images, so the figures are read, not estimated; the metric definitions are one sentence each and no subset, protocol or run count is published for any of them. Baselines are OpenAI models only. - NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA, 48 contributors), arXiv 2609.21967, 2026-09-18 (
empirical, 19 pp): Table 1 (FDB 1.0, 7 systems, 4 provenance classes), Table 2 (FDB 1.5 behaviour distribution), Table 3 (VoiceBench, 6 systems × 9 subsets), Table 4 (FDB 3.0 vs Gemini Live 2.5/3.1), Table 6 (VoiceChat-TTS on LibriTTS), Table 7 (OpenASR at two chunk sizes). Parse note: PDF-derived (docling: 2.126.0, MLX layout/table stages); ingest verify returnedwarnand the raw carries zero repair blocks, so all seven tables were treated as unrepaired and reconciled cell-for-cell againstpdftotext -layoutat compile — all seven clean, no collapse, no shift, no dropped rows. Tables 5 and 7 additionally self-verify arithmetically (the SFT sampler weights sum to 1.175 and normalize to the 46.8/23.8/26.0/3.4% shares the prose states; both OpenASR column means reproduce 9.02 and 8.28 from their own eight rows). One cosmetic docling artifact: the Table 2 footnote is spliced mid-sentence into the VoiceBench paragraph in the markdown body. Evidence:empiricaland NVIDIA scoring NVIDIA — single run, no error bars anywhere in the paper, and every open-weight baseline row in Table 1 is imported rather than re-run - OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue — He, Chu, Chen et al. (18 authors; CUHK + Alibaba Token Hub / Qwen Team + SJTU + Shanghai Innovation Institute + Zhejiang; arXiv 2609.21465, 2026-09-18;
empirical, 52 pp): §3 the benchmark and the gated-rubric scoring function (Eq. 1–2), §4.3 + Table 1 the model comparison, Appendix D.2 the 360-recording human probe, D.3 + Table 2 the six-judge sensitivity study, D.4 + Table 3 single-vs-multi-turn, E.1 + Table 4 the training configuration, E.5 the checkpoint-noise analysis. Parse note: PDF-derived (docling: 2.126.0, MLX layout/table stages); ingest verify returnedwarnand the raw's only four[!note]repair blocks all sit on Appendix-F rubric examples, so every main-paper table was treated as unrepaired. Tables 1, 2 and 3 were reconciled cell-for-cell againstpdftotext -layout(pp. 9–10, 24, 25) — all three clean, no collapse, no shift, no dropped rows, and Table 1's 15 rows match the prose's "twelve released systems, OmniVChat-RL, and two reward ablations." Table 4 is fully collapsed (every label in one cell, every value in another, both subsystems) and was rebuilt frompdftotext -layoutp.26 before any figure from it was used. Figure 3's per-subcategory counts sum to 2,800 and to the stated 2,550 single-turn, which is what validates the figure↔count mapping. One docling artifact worth knowing: the image emitted immediately under the Figure 2 caption (image_000002) is not Figure 2 — it is the Example 2 trace table rendered as a picture; Figure 2 isimage_000001. Evidence:empiricaland one lab owning the task definition, the data engine, the benchmark, the human probe, the judge, the base model, the reward and the trained model, with no external benchmark reported. Full treatment here and on LLM-as-a-Judge - Toward Native Multimodal Modeling: A Roadmap — An, Lu, Dong et al. (Tencent Youtu Lab + 5 universities; arXiv 2605.25343, 2026-05-25;
practitioner-opinion, 52 pp): §7 + Table 3 (39 benchmarks — 18 image / 7 audio / 14 video), §7.2 full-duplex evaluation, §7.3 streaming video and the ResAdapt efficiency result, §8.5 the four open evaluation directions. Parse note: docling welded and then dropped most of Table 3; the raw carries a[!note]rebuild frompdftotext -layoutpage 29, which this compile re-verified row-for-row (the one discrepancy is cosmetic — the source header reads "Task Group," not "Category"). The survey reports no benchmark numbers of its own. Full treatment on Native Multimodal Modeling: Fusion Depth and I/O Duality
Cited by 22
- Full-Duplex Interaction×8
gpt live 1 in the api — OpenAI, 2026-09-10 (vendor-claim): the third stay-present assertion, native…
- Artificial Analysis×6
Agentic Performance (τ-Voice) · Interactivity Benchmarks, Gemini 3 8 Live · AA's run of the τ-Voice…
- Gemini 3.8 Live×4
ServiceNow's EVA-Bench is the chart that does not flatter. Google's prose says "our models push the…
- Interaction / Background Model Split×4
And OpenAI's own benchmark footnotes confirm the interface empirically, in its own favour. Four of…
- LLM-as-a-Judge×4
OmniVChat (He, Chu, Chen et al., Alibaba Qwen Team + CUHK + SJTU, arXiv 2609.21465, 2026-09-18,…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×4
The full-duplex rows corroborate the vault's own numbers from a neutral source: Moshi Eval at a 200…
- Content-Driven Intervention×3
Turn-taking research asks when the floor is available. Content-driven intervention asks whether the…
- Interaction Models×3
See Full Duplex Interaction and Interactivity Benchmarks for how these are demonstrated and…
- GPT-Live×2
Interactivity Benchmarks — where its seven vendor-reported benchmark cards are catalogued, and why…
- Harness Activation and Adherence×2
Caveats before importing any number: different domain (spoken audio-visual dialogue, not agentic…
- Live-Path Minimalism×2
Interactivity Benchmarks — the benchmark surface the ablation above is run on; FD-bench v1 as a…
- Thinking Machines Lab×2
Benchmarks their model against GPT-realtime-2.0 / 1.5 (OpenAI) and Gemini-3.1-flash-live (Google)…
- TML-Interaction-Small×2
Interactivity Benchmarks — its full benchmark table and the baselines it beats
- Compute-Controlled Benchmarking
Interactivity Benchmarks — the voice-side instance of this page's problem: a cross-vendor chart…
- Cost-per-Task Over Cost-per-Token
Interactivity Benchmarks — where the voice cost chart sits alongside the quality boards it has to…
- Interaction & Multimodal
Interactivity Benchmarks — FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio)…
- NVIDIA
Interactivity Benchmarks — runs the tool-calling evaluation cluster (BFCL-audio, FDB3, EVA-Bench)…
- OpenAI
Realtime voice systems engineering. GPT-Live (July 2026) is its third-generation voice system: a…
- Production-Sourced Evaluation
Interactivity Benchmarks — where the generated-stimulus benchmark above is catalogued as an…
- Scale-Dependent Prompt Sensitivity
Interactivity Benchmarks — another case of a paper inventing its own evaluation framing (FD-bench…
- Time-Aligned Micro-Turns
Interactivity Benchmarks — turn-taking latency (0.40s) is the direct, measured payoff of removing…
- Turn-Level Credit Assignment
Interactivity Benchmarks — the same absence on the measurement side, in a multi-turn dialogue…
Related articles
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Full-Duplex Interaction
Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…
- Content-Driven Intervention
Speaking because the content warrants it — correcting a false claim, warning of a hazard, supplying a searched-for word…
- Live-Path Minimalism
GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; deleg…
