H
Howardism
Plate IIEntities中文HOWARDISM

GPT-Live

OpenAI's third-generation voice system (July 2026): a full-duplex voice model that listens and speaks simultaneously — no turn detector anywhere in the audio path — and consults a backend text model over an asynchronous delegation path without interrupting the conversation; replaced Advanced Voice Mode after a silent production shadow test; powers ChatGPT Voice including desktop computer control and agent coordination — and since 2026-09-10 ships as **GPT-Live-1 in the API** at $0.05/min for the front-end voice layer alone, with the backend model (GPT-6 Astra / Terra / Luna, or a third party) chosen and billed separately, 12 voices, telephony, and OpenAI-reported gains it attributes to the pairing: Tau3 Pass@1 86.2% and Full Duplex Bench v1.5 Interactivity 80.10% against GPT-Realtime-2.1's 45.7% and 45.4%

Article metadata
Publication details
Published:August 4, 2026
Filed:Entity
Domain:Entities
Reading:13 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for GPT-Live

Sources#

What it is#

OpenAI's third-generation voice system, launched July 2026 — a full-duplex voice model (Full-Duplex Interaction) that listens and speaks at the same time, with the surrounding realtime serving system built over six months (Live-Path Minimalism). The load-bearing design choice: the voice model, not a harness, is in control of the conversation. Audio streams in and out of the model continuously; there is no turn detector anywhere in the audio path.

The three generations it closes out#

OpenAI's own framing of the lineage:

  1. Cascaded — speech-to-text → LLM → text-to-speech in series. Sequencing added latency and threw away tone and pacing.
  2. Speech-to-speech (Advanced Voice Mode) — the model processes audio directly, preserving detail lost in transcription — but inference still waited on a separate turn detector: "guess too soon, and the user gets cut off; guess too late, and the response feels sluggish. Only after the detector made its decision could the much larger LLM get to work." In the post's words: "The model handled more of the interaction, but the interaction remained turn-based."
  3. GPT-Live — full-duplex; the detector is gone; turn-taking is model behavior. The dissolution Turn-Based Interface Bottleneck predicted for exactly this component.

Two models, one conversation#

When deeper reasoning or tool use is needed, GPT-Live delegates to a backend text model — frontier models such as GPT-5.5 (superseded 2026-09-10 by Build more natural voice experiences with GPT‑Live‑1 in the API, which names GPT-6 Astra, Terra and Luna, and allows a third-party model) — on an asynchronous path — OpenAI's phrase is "effectively decoupling 'talking' from deeper 'thinking'". This is the Interaction / Background Model Split arrived at independently and shipped at ChatGPT scale: the voice model briefly keeps the exchange moving while the frontier model reasons, and results are incorporated without interrupting the flow. The delegation loop is engineered as a latency budget (pre-warmed prefilled inference sessions, session affinity, prompt caching — details on Live-Path Minimalism).

Deployment#

  • Validated by a silent shadow test: a gradually increasing share of production ChatGPT Voice sessions was mirrored to both Advanced Voice Mode (still serving users) and GPT-Live running read-only, exposing it to real clients, networks, session lengths, and geography before any user heard it.
  • Powers ChatGPT Voice, including the newly launched ability to control your computer and coordinate agents from the ChatGPT desktop app — voice as a surface over agentic capability rather than a standalone feature.
  • A GPT-Live API is announced as upcoming (superseded 2026-09-10: it launched — see the next section); OpenAI positions the architecture as "a broader platform for realtime interaction" spanning more devices, apps, and modalities.

The API generation: GPT-Live-1, 2026-09-10#

The launch post (Build more natural voice experiences with GPT‑Live‑1 in the API, vendor-claim) makes the July architecture a purchasable product, and the shape of the product is the architecture's own seam.

What is sold, and what is not. OpenAI sells the front-end voice layer at $0.05 per minute; the backend model and agent harness are the developer's, "paired" with it and billed separately. So the media/application separation described on Live-Path Minimalism is the product boundary and the price boundary at once — a per-minute meter on the live path, per-token billing on everything behind it. Backend options named in the post: GPT-6 Astra (complex reasoning), Luna (high-volume tasks like scheduling or order updates), Terra (used for the tool-calling benchmark runs) — "or a third-party model."

What the developer controls. The models, tools and agent harness behind the conversation; tone, pace and conversational style through the system prompt; turn detection, which is natively supported "although GPT-Live-1 is not a turn-based model," so applications can keep building against explicit turn boundaries. The model natively emits ASR transcripts and response text alongside audio, and supports alphanumeric understanding and keyword biasing. Twelve new voices ship (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder); custom voices are gated behind sales. Telephony is supported, which is the first named deployment target outside an OpenAI surface.

The delegation boundary has a published wire shape. The post's one code sample returns a backend answer to the live session with live.send({ type: "session.commentary.append", delegation_id, content }) — an async append against a delegation id, with "connection setup and delegation handling … omitted." It is a Codex SDK thread in the example, i.e. an arbitrary developer-side agent. This is the first time any lab has shown the call signature of the interaction/background handoff; NVIDIA published the in-model token scheme, OpenAI publishes the application-facing RPC.

What OpenAI claims for it (all first-party, measured against its own prior models). Seven benchmark cards compare gpt-live-1 to gpt-realtime-2.1 and gpt-realtime-2; four of the seven state which backend the GPT-Live configuration used, and the other three do not (which is itself informative — see below). Headline figures, all OpenAI-reported:

Benchmark (metric)gpt-live-1gpt-realtime-2.1gpt-realtime-2backend used
Tau3 Voice intelligence (Pass@1)86.2%45.7%42.4%Astra (medium)
Tau Banking Voice knowledge (Pass@1)32.0%12.4%10.3%Astra (medium)
Artificial Analysis Conversational Dynamics (avg)97.3%95.7%95.3%—
Full Duplex Bench v1.5 Interactivity (avg)80.10%45.4%47.8%—
Full Duplex Bench v1 turn-taking latency (lower better)0.798 s1.41 s1.63 s—
Full Duplex Bench v3 tool calling (Pass@1)87.0%60.0%58.0%Terra (low)
Full Duplex Bench v3 response quality90.0%88.0%81.0%Terra (low)

Two readings the table supports and the prose does not state. First, the starred rows are the delegating configuration's score, not the voice model's — every intelligence and tool-calling number is gpt-live-1 plus a named backend at a named reasoning effort, which makes this the second corpus instance of the split behaving as a capability interface (see Interaction / Background Model Split). Second, the unstarred rows are where the frontend alone moves, and they move very unevenly: Full Duplex Bench v1.5 Interactivity nearly doubles (45.4 → 80.10) and turn-taking latency roughly halves (1.41 → 0.798 s), while Artificial Analysis Conversational Dynamics moves 95.7 → 97.3 — a benchmark already saturated at the top of its range.

Claims that do not fully reconcile. The prose headline is a "30 percentage point" Full Duplex Bench gain over GPT-Realtime-2.1; no single published card shows +30 (v1.5 Interactivity is +34.7, v3 tool calling +27.0, v3 response quality +2.0), so the headline is an unstated aggregate. The customer figures are relayed, not measured by OpenAI: Speak reports "cutting interruptions by almost 80% versus previous turn-based systems" in early evaluations, with the baseline unspecified; an unnamed company's co-founder reports GPT-Live-1 "simplified our code base by 80% and removed 23K lines of code" against a cascaded build (see Harness Shrinkage as Models Improve); Yelp reports improved call-handling rates for Yelp Host and Hatch. Three of the four customer testimonials (Speak, Fin, Cognition) did not render in the captured DOM and are not reproduced anywhere in this wiki.

What the post claims in words and nowhere measures: "silent context management" (handling background noise and silence "without interrupting the conversation or narrating every step out loud"), "long-session reliability" (context retention across extended interactions), and the stay-present property — "this lets the conversation continue while work happens in the background." The last is the third independent assertion of speak-while-thinking in this corpus and the third with no measurement attached (Interaction / Background Model Split tracks the count).

Scored by a competitor, on boards OpenAI did not pick (September 2026)#

Five days after the API launch, Google's Gemini 3.8 Live post (Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, 2026-09-15, vendor-claim) puts GPT-Live-1 Astra on four third-party charts — the first time any number for this product comes from outside OpenAI, even if it still reaches the wiki through a vendor. Attributed to Google throughout; neither party's board has been read at source.

  • Artificial Analysis Speech to Speech Index: GPT-Live-1 Astra (medium) 81.5, second to Gemini 3.8 Live Extended Thinking (high) at 82.6, ahead of Grok Voice Think Fast 2.0 at 81.3. A 1.1-point gap on an index whose composition is published nowhere in this corpus.
  • Artificial Analysis τ-Voice (agentic): 67.9, again second to 82.6's sibling at 68.6.
  • Sierra τ³-Banking: 32.0 — identical to OpenAI's own "Tau Banking Voice" card, as is gpt-realtime-2 at 10.3. Two adversarial parties printing the same two rows means both are quoting Sierra's board rather than running it, which is the corpus's first cross-vendor corroboration of a number and also its explanation.
  • Artificial Analysis cost per hour of input audio: $5.83, the most expensive system on the chart — against Gemini 3.8 Live's $0.84 and Extended Thinking's $3.50. That is a measured cost for the pairing, not OpenAI's $0.05/min front-half price; the arithmetic relating the two, and why it is an inference rather than a finding, is on Cost-per-Task Over Cost-per-Token.

The row that does not reconcile is the interesting one. OpenAI's own card puts gpt-live-1 + Astra (medium) at 86.2% on "Tau3 Voice intelligence (Pass@1)"; Google's chart puts the same product with the same named backend at the same stated effort at 67.9% on "Agentic Performance (τ-Voice)". Since the banking rows matched exactly, this is not a transcription artifact — the two names are almost certainly two different instruments, and neither vendor publishes a protocol. It is not evidence that either number is wrong, and it is a standing warning against reading any voice leaderboard row without its benchmark's version. Full reconciliation attempt on Interactivity Benchmarks.

One further observation, offered as a caveat and not an accusation: on Google's charts GPT-Live-1 is shown at Medium effort while Google's flagship is shown at High. Medium is the setting OpenAI's own cards footnote for Astra, so it may simply be the published configuration — but a cross-vendor chart in which the home model runs hotter than the comparator is not a controlled comparison.

Connections#

  • OpenAI — builder; the post is OpenAI's first-party build account
  • Live-Path Minimalism — the serving architecture built for it: "the voice must flow"
  • Full-Duplex Interaction — the interaction property it ships in production (audio-only; TML's generalization to video/text remains a research preview)
  • Interaction / Background Model Split — the two-model talking/thinking architecture it independently converges on
  • Turn-Based Interface Bottleneck — the harness component it dissolved, and the nuance that turns survive as a derived application-layer view
  • Interactivity Benchmarks — where its seven vendor-reported benchmark cards are catalogued, and why the 0.798 s latency figure must not be raced against other labs' self-reported ones
  • TML-Interaction-Small — the research-preview sibling: same architectural conclusions from Thinking Machines Lab two months earlier, generalized past audio, with disclosed internals
  • Gemini 3.8 Live — the competitor that arrived five days later and scored it on four third-party boards
  • Artificial Analysis — the evaluator behind its Conversational Dynamics card and behind three of the four charts Google scores it on

Sources#

  • How we built a realtime system for responsive voice AI in six months — OpenAI engineering blog, 2026-07-29 (case-study, first-party): the system-architecture account. The separate "Introducing GPT-Live" launch post is not in the corpus; capability and product claims here are limited to what the engineering post states.
  • Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, "Build more natural voice experiences with GPT-Live-1 in the API", 2026-09-10 (vendor-claim, ~1,670 words, no individual byline): the API launch, the $0.05/min front-end price and separate backend billing, the named backends, the twelve voices, telephony, and all seven benchmark cards. Every number above is OpenAI measuring its own model against its own prior models, with no third-party or independent run; four of the seven cards name the backend used and three do not. Capture note: the chart cards were lazily-mounted JS components recovered as text rather than images, so the figures are read values, not transcriptions; three of four customer testimonials did not render and are absent.
  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google, 2026-09-15 (vendor-claim): a competitor's reproduction of four third-party charts on which GPT-Live-1 appears. Cited only for rows where GPT-Live-1 is the subject, always attributed to Google, and never used to rank the two products — the two vendors' agentic rows differ by 18 points under two benchmark names
§ end
Cited by 13
  • Interaction Models×6

    The framing says one model. Google ships two named models, 3.8 Live and 3.8 Live Extended Thinking,…

  • Full-Duplex Interaction×5

    OpenAI's Gpt Live ships audio full-duplex at ChatGPT scale: "its voice model is full-duplex, which…

  • Gemini 3.8 Live×4

    The launch is argued almost entirely on third-party boards rather than internal evals — Artificial…

  • Interaction / Background Model Split×4

    The three bracketed references are Thinking Machines' interaction models, Qwen-audio-agent, and Gpt…

  • OpenAI×3

    gpt live 1 in the api — OpenAI, "Build more natural voice experiences with GPT-Live-1 in the API",…

  • Artificial Analysis×2

    Conversational Dynamics · Interactivity Benchmarks, Gpt Live · near-saturated: 95.3 → 95.7 → 97.3…

  • Interactivity Benchmarks×2

    Gpt Live — the first shipping commercial voice product scored against FD-bench in public, by its…

  • Live-Path Minimalism×2

    The serving-side architecture behind Gpt Live, stated by OpenAI as one principle: "the voice must…

  • Turn-Based Interface Bottleneck×2

    Two months after TML's argument, OpenAI shipped its conclusion: Gpt Live "removes the turn detector…

  • Content-Driven Intervention

    Gpt Live — production audio full-duplex, untested by this protocol; and, from September 2026, a…

  • Google DeepMind

    A sixth posture, and the lab's first appearance in the interaction/multimodal domain: Gemini 3.8…

  • Entities — People, Orgs, Tools & Projects

    Gpt Live — Entity. OpenAI's third-generation voice system (July 2026): a full-duplex voice model…

  • Native Multimodal Modeling: Fusion Depth and I/O Duality

    Nothing from the interaction-model line appears at all: no Tml Interaction Small, no Inkling, no…

Related articles
  • Interaction Models

    Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…

  • Interaction / Background Model Split

    Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…

  • Interactivity Benchmarks

    FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…

  • Full-Duplex Interaction

    Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…

  • Gemini 3.8 Live

    Google's September 2026 live-dialogue pair — 3.8 Live (scale/cost) and 3.8 Live Extended Thinking (high-complexity), sh…