H
Howardism
Plate IIInteraction & Multimodal中文HOWARDISM

Live-Path Minimalism

GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; delegation, compaction, persistence and instance management run asynchronously off it. Stateful-instance handoff (warm a replacement, prefill, run both, cut over) turns compaction and rebalancing into zero-interruption transitions; delegation is a budgeted loop over a pre-warmed prefilled frontier-model session; WARP + Instant Connect collapse WebRTC startup from six round trips to one UDP packet; capacity is concurrent sessions keeping every frame on schedule, not GPU throughput. The boundary is third-party-facing and priced (GPT-Live-1 in the API, 2026-09-10: live path alone at $0.05/min); NVIDIA's with/without-delegation ablation (Sept 2026) prices it: +15 ms turn-taking latency on the live path, paid for with ASR WER 10.80→11.47 and CommonEval 2.87→2.36 off it

Article metadata
Publication details
Published:August 4, 2026
Filed:Concept
Domain:Interaction & Multimodal
Reading:21 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Live-Path Minimalism

Sources#

Summary#

The serving-side architecture behind GPT-Live, stated by OpenAI as one principle: "the voice must flow." A full-duplex model (Full-Duplex Interaction) makes the conversation a continuous media loop, and any delay in transport, processing, or inference becomes an audible pause — a turn-based system could tolerate variation in when an audio blob arrived; a live media system must deliver every audio frame on schedule. The design response is to make the live path as small as possible and move everything else — deeper reasoning, tool use, conversation persistence, context compaction, instance management — onto asynchronous paths that cannot stall it. All figures below are first-party (case-study); the protocol work is externally checkable via public IETF drafts.

Separate the media path from everything else#

  • Audio moves between client and voice model on a dedicated fast path; delegation, tool use, and application work sit behind an asynchronous RPC boundary. A slow tool call or backend service can delay its own result, but cannot stall media.
  • The same boundary is the customization surface: applications change tools, policies, and backend behavior without touching the media frontend — which is what lets ChatGPT Voice grow application behavior (computer control, agent coordination) without risking responsiveness.
  • The media frontend and inference logic were rewritten from Python asyncio to Go; OpenAI reports the new system's p95 frame-delivery smoothness matches the old system's p50.
  • WebRTC is the transport: built for low-latency media, it rides through packet loss, clock drift, and connection changes — subtly stretching audio to cover late packets, then briefly accelerating playback to catch back up.

Long-lived stateful inference: disruptive operations become managed transitions#

Stateful streaming inference means a session's context lives in a model instance — but sessions run long, context grows, and instances spin up and down with demand. The mechanism that squares this is seamless instance handoff: warm a replacement instance alongside the current one, prefill it with the session context, run inference against both in parallel, and cut over when the replacement is fully ready.

The important application is context compaction as a managed transition. Compaction takes time, and because it rewrites past context it invalidates the KV cache, forcing a fresh prefill — dead air if paid on the live path. Instead, the original instance keeps chatting while the system compacts the context and prepares a replacement instance with it; the session cuts over with no media interruption: "even during a handoff, the conversation never misses a beat." Where Context Lifecycle Management's cache-aware commit prices the cache break and can hold a plan pending, this is the serving-side dodge: hide the break's latency behind a parallel instance (its compute cost is still paid — just off the path the user can hear).

The delegation loop is a latency budget#

The Interaction / Background Model Split in production form: the voice model can briefly keep the exchange moving while a frontier model reasons, but "cannot hide an arbitrarily slow response" — so the full delegation loop (routing, prompt processing, inference, tool calls) is treated as part of the responsiveness budget:

  • At voice-session start, the application server pre-creates the frontier model's inference session and prefills it with the initial conversation context — the prompt is fully processed before the first delegation is ever requested.
  • The session persists for the whole conversation with stable session affinity plus prompt caching; a worker failure stays cheaply recoverable.
  • Reasoning effort, output limits, tool schemas, and model↔tool round trips are all tuned as levers on time-to-useful-result.

The principle priced: what delegation costs the live path (NVIDIA, September 2026)#

"The voice must flow" is stated by OpenAI and never measured — GPT-Live has no published counterfactual in which delegation sits on the live path, or in which the voice model was never taught to delegate at all. Hu et al. (NVIDIA, arXiv 2609.19334, 2026-09-16, empirical) supply that ablation on a different system: an identically trained duplex frontend with and without tool-call training, everything else held fixed (details of the architecture on Interaction / Background Model Split).

Their architecture takes live-path minimalism further than OpenAI's, by construction: the backend never touches audio in either direction. It receives an ASR transcript and returns text, so the media loop is not merely shielded from the backend by an async RPC boundary — the backend is not in the media modality at all.

Cost to the live path: essentially nothing. On NVIDIA's internal multi-turn conversation benchmark, adding the delegation capability moves turn-taking latency 423 → 438 ms (+15 ms) and barge-in latency 403 → 395 ms (−8 ms), with barge-in accuracy 100% → 99%. On Full-Duplex-Bench v1 the tool-call-trained frontend is more responsive than its own baseline: smooth turn-taking TOR 93 → 100% and latency 221 → 92 ms, user-interruption latency 590 → 349 ms, GPT interruption score 3.94 → 4.04.

Cost elsewhere: real, and not what the abstract's "minimal modifications" implies. The same ablation pays for the token in conversational quality:

  • turn-taking precision/recall 85/94 → 82/91;
  • streaming ASR word error rate 10.80% → 11.47% on the Open ASR Leaderboard average (the paper attributes this to the added tool-call training data being mostly synthetic);
  • VoiceBench CommonEval 2.87 → 2.36 on a 5-point scale — a ~18% drop, the largest regression in the set, against OpenbookQA 65.0 → 64.7 and AlpacaEval 3.54 → 3.61, which are flat;
  • FDB-v1 pause handling TOR 53.2 → 68.2% (lower is better on this metric — the model interrupts pauses more) and user-interruption TOR 95.5 → 89.5%.

The generalisable form: live-path minimalism buys latency, and the bill is paid in the frontend's behaviour, not its clock. Teaching an interaction model when to hand off is itself a capability that competes with the ones it already had; the cost shows up in where it takes the floor and how it handles disfluency, exactly the properties the live path exists to protect. Nothing in GPT-Live's account rules out the same trade having been paid there — it simply was not measured.

Caveats to carry: all figures are single-run, self-reported, on one system, and the internal turn-taking benchmark is unpublished. The contrast is internal (same recipe, one variable), which is what makes it usable; the absolute numbers are not comparable across labs.

The boundary ships to third parties, and it has a price (September 2026)#

The architecture above was an internal account of an internal system, and the obvious question it left was whether the media/application separation was a customization surface for OpenAI, or a surface anyone could build behind. The API launch (OpenAI, 2026-09-10, vendor-claim) answers it: the separation is the product, and each side of it is priced separately.

  • The live path is the SKU. GPT-Live-1 sells at $0.05 per minute for "the front-end voice layer." The developer does not host, tune or touch it; the meter is on time, not tokens, because the thing being bought is a media loop that must keep schedule.
  • Everything behind the boundary is the developer's. "Developers choose the models, tools, and agent harness behind the conversation" — OpenAI's named backends (GPT-6 Astra for complex reasoning, Luna for high-volume tasks, Terra in the benchmark runs) or a third-party model. Backend inference is billed by whoever supplies it.
  • The async RPC boundary has a published call shape. The post's code sample runs an arbitrary developer-side agent (a Codex SDK thread over a local repo) and returns its answer with live.send({ type: "session.commentary.append", delegation_id, content }). An id issued by the live session, content appended asynchronously against it, "connection setup and delegation handling … omitted." That is the asynchronous RPC boundary of the first section, exposed as an API verb.
  • What is not exposed. Everything in the two sections above that makes the live path survivable — seamless instance handoff, off-path context compaction, the pre-warmed prefilled backend session, session affinity, WARP/Instant Connect startup — is sold as behaviour, not as a control surface. The developer gets tone, pace and style through the system prompt, a choice of twelve voices, optional turn detection, and native ASR transcripts and response text. They do not get to place work on the live path, which is the same statement read from the other side: the boundary is exposed precisely because it is the thing the developer may not cross.

The generalisable form: live-path minimalism turned out to be a business boundary, not only an engineering one. The asset is the part that cannot tolerate a stall, and it is metered per minute; the part that can tolerate a stall is commoditised, interchangeable, and someone else's bill. That also means the split's two halves are now denominated in different units — per-minute on the front, per-token behind — which is a cost structure Cost-per-Task Over Cost-per-Token has not previously had to price.

Everything in this section is a vendor's description of its own product. That is the right evidence tier for the claim being made — a release post is first-party evidence of what a product exposes — and the wrong tier for any claim about how well it works.

Startup: collapsing the handshake#

Session start puts every protocol exchange on the critical path, and vanilla WebRTC predates the round-trip frugality of QUIC-era protocols — its stacked sub-protocols even repeat anti-DoS work. OpenAI's answer, developed with the WebRTC community as open specifications through the IETF's TSVWG:

  • WARP (WebRTC Abridged Roundtrip Protocol) — six network round trips down to one, via backward-compatible pieces: piggybacking the DTLS handshake over ICE (SPED), DTLS 1.3, pre-negotiating the SCTP handshake (SNAP), and pre-negotiated data channels instead of DCEP. Already implemented in libwebrtc and Pion.
  • Instant Connect — pre-negotiates the SDP signaling parameters ahead of time without reserving server capacity; if they're valid the server materializes the session when the first media packet arrives, and if stale the standard signaling flow is already running as fallback, costing nothing extra.

Net effect: a client can start a session with a single UDP packet.

Turns become a derived view#

Removing the turn detector doesn't remove the need for turns — ChatGPT's conversation UI, analytics, and safety systems still consume discrete user/assistant messages. So the application server derives turns from the continuous stream: partial transcripts and timing signals infer who holds the floor; the newest message stays provisional (text, timing, and speaker assignment all revisable) until floor-holding is sustained enough to finalize. Overlap needs policy — a brief assistant "mm hmm" while the user talks should not become its own message, a substantive interjection should. Every segmentation policy trades freshness for certainty, so the system maintains two views: a speculative view feeding the live UI (which can tolerate revisions) and an authoritative record feeding analytics (which needs finality). The turn-based data model survives at the application boundary — maintained by inference over the stream rather than imposed on the audio path (see Turn-Based Interface Bottleneck).

What production testing taught#

Before serving users, a silent shadow test routed a small, growing share of production ChatGPT Voice traffic to both Advanced Voice Mode (still serving) and the new system running read-only — real clients, networks, session lengths, and geography with no user-visible change. Lessons:

  • Capacity is not GPU throughput. Voice sessions hold open and send frames continuously, so CPU-side stream handlers, queues, and network paths must scale alongside inference; a supporting component saturated before load-test estimates predicted, compounding latency. The capacity question became "how many concurrent sessions can the system sustain while keeping every frame on schedule?"
  • Geography is first-order. Distant capacity taxes both startup and streaming; rollouts began validating model, regional capacity, and traffic-steering together, with latency broken down by source geography.
  • Failures that need time and state to exist. Long sessions exposed memory and persistence pressure; reconnects exercised compaction and state restoration; ordinary disconnects revealed races in the shutdown handshake — none visible in short load tests.
  • Observability had to be rebuilt: metrics conflating latency sources, dashboard aggregates hiding individual unhealthy engines, and config drift between tested and deployed systems led to granular telemetry, validation against known-good configurations, staged ramps, and per-path kill switches. The shadow test became a rehearsal for detection, containment, and recovery — not just a throughput check.

The first published capacity number in the corpus' own unit (NVIDIA, September 2026)#

This page's closing claim is that capacity for a live-voice system is concurrent sessions keeping every frame on schedule, not GPU throughput — a unit nobody in the corpus had published. NemotronLabs VoiceChat (NVIDIA, arXiv 2609.21967, 2026-09-18, empirical) publishes it, in Appendix C.4, in one sentence:

"With four concurrent streams, the per-stream p95 inference latency is 118 ms per 160-ms audio chunk. This corresponds to 1.36× real-time processing throughput."

Four streams on one H100 PCIe 80 GB, and 1.36× real time is thin. The stated precision mix is BF16 for the perception encoder and LLM backbone, FP32 for the TTS backbone and the cached Mamba recurrent states, TF32 for eligible FP32 matmuls, and — explicitly — "lower-precision and quantized variants were not evaluated," so this is the unoptimized floor rather than a tuned ceiling. The margin it reports is the interesting part: at 118 ms of compute per 160 ms of audio, a 26% slowdown puts a frame late, and the budget is being spent at only four concurrent sessions on a top-end accelerator. That is the cost of the every-frame-on-schedule constraint stated as a number for the first time, and it is a strong argument for the minimalism this page is named after.

The stack is four runtimes stitched by a fifth, which is itself the page's thesis in assembly form: a PyTorch perception path under CUDA Graph capture; the Nemotron backbone and TTS decoder on a custom vLLM fork extended for encoded-speech tensors appendable through an incremental interface and for multi-codebook generation; a causal PyTorch codec, also CUDA-graphed, incrementally decoding to 22.05 kHz while reusing cached convolutional state; a Triton Inference Server Python backend coordinating scheduling and tensor exchange "because these stages execute in heterogeneous runtimes"; and a FastAPI bidirectional WebSocket at the edge. Clients send 80 ms mono PCM chunks, and audio in both directions moves in integer multiples of 80 ms. Nothing here is a general-purpose serving stack; every component is either patched or wrapped to make one persistent stream work.

A deployment-time latency dial with its price published. The chunk duration is configurable, "allowing the inference cadence to be tuned to the latency budget of a given real-time deployment," and because the encoder is cache-aware the same trained checkpoint serves multiple chunk sizes with no retraining. The cost is published on both axes: average OpenASR WER 9.02% at 80 ms against 8.28% at 160 ms, improving on all eight evaluation sets as the chunk grows. A serving knob whose accuracy price is measured is rare in this corpus; catalogue entry on Interactivity Benchmarks.

What the tool call does to the live path here, and it is worse than the delegating sibling. The runtime keeps the decode of a tool call off the critical path — "fast-decode" runs asynchronously when the function channel emits its start marker, and TTS speaks a pre-configured acknowledgement while the executor works — but the audio path does not survive: incoming audio keeps flowing through perception and RNN-T, "but it is not used to condition response generation during tool execution. Consequently, barge-in is unavailable during this phase." The voice flows; the user cannot stop it. Comparison with Hu et al.'s delegating frontend on Interaction / Background Model Split.

And a persistent-decoder failure mode this page had no data on. Everything above assumes a decoder held open across a whole conversation. VoiceChat-TTS is evaluated at turn 1 and turn 4 of continuous decoding: intelligibility and predicted quality hold (WER 2.00% → 2.20%, SQuIM-MOS 4.380 → 4.376 for unseen speakers), but speaker similarity drifts 0.757 → 0.685 — the zero-shot voice degrades while the content does not. For speakers seen in training the same span moves only 0.785 → 0.778, which localizes the drift in zero-shot speaker conditioning rather than in persistent decoding as such. The thing that breaks first when you keep the live path alive is identity, not quality.

Connections#

  • Native Multimodal Modeling: Fusion Depth and I/O Duality — the architecture-side version of the same constraint: §6.3 makes TTFT and sustained latency first-class targets and names dynamic KV-cache management against cache contention, and §8.4 calls deployable low-latency full-duplex an unsolved industrial problem
  • GPT-Live — the system this architecture serves
  • Full-Duplex Interaction — the model property that creates the every-frame-on-schedule constraint
  • Interaction / Background Model Split — the two-model architecture whose delegation half this page's latency-budget engineering implements
  • Time-Aligned Micro-Turns — TML's disclosed counterpart on the inference internals: persistent GPU-resident sequences for frequent small prefills (upstreamed to SGLang), where OpenAI's stateful stack — instance handoff, off-path compaction — stays proprietary and one level up
  • Turn-Based Interface Bottleneck — the harness this dissolves from the audio path, and the place turns re-enter as a derived view
  • Interaction Models — the research framing whose serving problem this is
  • Context Lifecycle Management — the priced version of the compaction/cache-break trade this architecture instead hides behind a parallel instance
  • Interactivity Benchmarks — the benchmark surface the ablation above is run on; FD-bench v1 as a with/without instrument rather than a leaderboard
  • Cost-per-Task Over Cost-per-Token — the productised split prices its two halves in different units: per-minute on the live path, per-token behind it
  • NVIDIA — the lab whose ablation prices the principle

Open Questions#

  • TML upstreamed streaming-sessions serving into SGLang; GPT-Live's stateful serving (persistent sessions, seamless instance handoff, off-path compaction) is proprietary. Does an open-source inference stack ship instance handoff for full-duplex voice? Trigger: an SGLang/vLLM release with session-handoff support.

Resolved Questions#

  • Does the upcoming GPT-Live API expose the media/application separation to third parties — application logic customizable behind the async RPC boundary without touching the live path — or is the boundary internal-only? Trigger: GPT-Live API launch. Answered (2026-09-21) — the trigger landed and the first branch is correct: GPT-Live-1 shipped in the API on 2026-09-10 (vendor-claim) selling the front-end voice layer alone at $0.05/min, with the developer choosing "the models, tools, and agent harness behind the conversation" — OpenAI's own backends or a third-party model — and returning results across the boundary asynchronously via live.send({ type: "session.commentary.append", delegation_id, content }), demonstrated with an arbitrary developer-side agent (a Codex SDK thread). The boundary is not internal-only; it is the product seam and the price seam. A release post is exactly the right evidence for what a product exposes, so the vendor-claim tier retires this question — it would not retire a question about how well the exposed boundary performs. The bullet was a two-branch question rather than a prediction, but this page's framing leaned, so grading what the lean got right and wrong: right that the same boundary would be the third-party customization surface ("applications change tools, policies, and backend behavior without touching the media frontend" — this is now literally the API contract, down to backends OpenAI did not train); wrong in the implicit scope — the page treated the separation as one boundary, and the product splits it in two, exposing the delegation RPC while keeping the stateful-serving machinery (instance handoff, off-path compaction, the pre-warmed prefilled backend session, WARP startup) entirely behind the curtain as behaviour rather than control surface; right for a reason not yet in evidence that the boundary would hold under third-party load — nothing published measures a developer-built backend's latency distribution against the live path, so the separation is confirmed as exposed and unconfirmed as robust. Full treatment in the section above.

Sources#

  • How we built a realtime system for responsive voice AI in six months — OpenAI engineering blog, 2026-07-29 (case-study, first-party): all architecture and testing detail; performance figures (Go rewrite p95≈old p50, WARP 6→1 round trips) are vendor-reported and unaudited, while WARP/SPED/SNAP exist as public IETF drafts with libwebrtc and Pion implementations. The post's one figure (system-architecture diagram) is fully described by its own caption and the surrounding prose; not separately viewed.
  • NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA), arXiv 2609.21967, 2026-09-18 (empirical, 19 pp): Appendix B the four-runtime serving stack (CUDA-graphed perception, custom vLLM fork, Triton Python backend, FastAPI WebSocket, 80 ms chunks) and the fast-decode / fast-inject tool path with its barge-in concession; Appendix C.4 the capacity measurement (H100 PCIe 80 GB, 4 concurrent streams, p95 118 ms per 160 ms chunk, 1.36× real time, no quantized variants evaluated); Table 7 the 80 ms vs 160 ms chunk trade (9.02% → 8.28% average WER); Table 6 the multi-turn TTS stability and speaker drift. All 7 tables reconciled against pdftotext -layout; single run, no error bars, NVIDIA measuring its own system on its own hardware
  • A frontend-backend architecture for tool calls in full-duplex speech models — Hu et al. (NVIDIA), arXiv 2609.19334, 2026-09-16 (empirical, 5pp): Tables 4–7, the with/without tool-call-training ablation of a duplex speech-to-text frontend. Tables 4, 5 and 7 carried a table-shift ingest warning that reconciliation showed to be a spanning-header artifact; every figure above was re-read from pdftotext -f 4 -l 5 -layout. Single run, no error bars, internal benchmark unpublished.
  • Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (vendor-claim, ~1,670 words, unbylined): the API launch, the $0.05/min front-end price, the developer-chosen backend including third-party models, the session.commentary.append delegation call, and the twelve voices / telephony / turn-detection surface. First-party description of first-party product surface — the right tier for what is exposed, not for how well it works.
§ end
Cited by 14
  • GPT-Live×4

    OpenAI's third-generation voice system, launched July 2026 — a full-duplex voice model (Full Duplex…

  • Interaction / Background Model Split×4

    Each half is metered separately, in a different unit. GPT-Live-1 sells at $0.05 per minute for "the…

  • Cost-per-Task Over Cost-per-Token×2

    It is a clean case for "the expensive component is the one with the hard constraint." OpenAI meters…

  • Full-Duplex Interaction×2

    OpenAI's Gpt Live ships audio full-duplex at ChatGPT scale: "its voice model is full-duplex, which…

  • Interaction Models×2

    Live Path Minimalism — the serving architecture an interaction model demands, from the lab that…

  • Interactivity Benchmarks×2

    Live Path Minimalism — where FD-bench v1 is used as an ablation surface for what a delegation token…

  • NVIDIA×2

    Three things make this NVIDIA's most disclosure-heavy publication in the corpus. It is open-weight…

  • OpenAI×2

    Realtime voice systems engineering. GPT-Live (July 2026) is its third-generation voice system: a…

  • Turn-Based Interface Bottleneck×2

    Live Path Minimalism — where turns re-enter as a derived application-layer view after leaving the…

  • Context Lifecycle Management

    Live Path Minimalism — the serving-side dodge of the cache-aware-commit problem, from a system that…

  • Interaction & Multimodal

    Live Path Minimalism — GPT-Live's serving principle — "the voice must flow": the realtime media…

  • Native Multimodal Modeling: Fusion Depth and I/O Duality

    §6.3 and §8.4 are where the survey meets Full Duplex Interaction and Live Path Minimalism head-on,…

  • Open Questions Backlog

    Live Path Minimalism: TML upstreamed streaming-sessions serving into SGLang; GPT-Live's stateful…

  • Time-Aligned Micro-Turns

    Live Path Minimalism — the production serving counterpart one level up: GPT-Live's stateful…

Related articles
  • Interaction / Background Model Split

    Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…

  • Interactivity Benchmarks

    FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…

  • Content-Driven Intervention

    Speaking because the content warrants it — correcting a false claim, warning of a hazard, supplying a searched-for word…

  • Full-Duplex Interaction

    Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…

  • Interaction Models

    Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…