Sources#
- Build more natural voice experiences with GPT‑Live‑1 in the API
- How we built a realtime system for responsive voice AI in six months
- Interaction Models: A Scalable Approach to Human-AI Collaboration
Summary#
Thinking Machines Lab's framing of why current AI interfaces limit collaboration: the turn-based interface is a bandwidth bottleneck between human and model. It is the problem Interaction Models are built to dissolve.
The two-claim argument#
-
AI labs over-optimize for autonomy. Labs treat autonomous capability as the model's most important property; as a result, today's models and interfaces "aren't optimized for humans to remain in the loop." But in most real work users can't fully specify requirements upfront and walk away — good results come from a collaborative loop of clarification and feedback.
-
Humans get pushed out by the interface, not the work. "Humans increasingly get pushed out not because the work doesn't need them, but because the interface has no room for them." The fix is to let people collaborate with AI the way they collaborate with other people: messaging, talking, listening, seeing, showing, interjecting — and the model doing the same.
The mechanism: a single thread#
Today's models "experience reality in a single thread":
- Until the user finishes typing/speaking, the model waits with no perception of what the user is doing or how.
- Until the model finishes generating, its perception is frozen — no new information arrives until it finishes or is interrupted.
This narrow channel limits how much of a person's knowledge, intent, and judgement can reach the model, and how much of the model's work is legible to the human. Analogy: "trying to resolve a crucial disagreement over email rather than in person."
Why harnesses don't fix it#
Existing real-time systems bolt interactivity on with a harness — VAD (voice-activity detection), turn-boundary prediction, dialog state machines — components "meaningfully less intelligent than the model itself." That harness precludes whole interaction modes:
- proactive interjection ("interrupt when I say something wrong")
- reaction to visual cues ("tell me when I've written a bug in my code")
- speak-while-listening ("translate Spanish→English live")
- speak-while-watching ("live-commentate this sports game")
The Bitter Lesson says these hand-crafted systems get outpaced by general capability growth → the resolution is to make interactivity model-native (see Time-Aligned Micro-Turns).
The harness dissolves in production: GPT-Live (July 2026)#
Two months after TML's argument, OpenAI shipped its conclusion: GPT-Live "removes the turn detector from the audio path" of production ChatGPT Voice. OpenAI's own retrospective states the bottleneck exactly as this page does — the turn detector "faced an unenviable task: guess too soon, and the user gets cut off; guess too late, and the response feels sluggish. Only after the detector made its decision could the much larger LLM get to work." Even the intermediate speech-to-speech generation, which absorbed transcription into the model, kept the gate: "the model handled more of the interaction, but the interaction remained turn-based." The less-intelligent harness component named by this page's thesis is now gone from a deployed system — Harness Shrinkage as Models Improve landing at the interaction layer, from a second lab, as production engineering rather than research argument.
The nuance the production system adds: turns don't disappear — they move. ChatGPT's conversation UI, analytics, and safety systems still consume discrete user/assistant messages, so GPT-Live's application server derives turns from the continuous stream after the fact — speculative and authoritative views, a freshness-for-certainty trade (details in Live-Path Minimalism). The turn-based structure was a bottleneck as a gate on inference; as a data model for the surrounding product it survives, maintained by inference over the stream rather than imposed on it.
The derived view becomes a product feature, and the harness shrinkage gets a number (September 2026)#
GPT-Live-1's API launch (OpenAI, 2026-09-10, vendor-claim) closes the loop on both halves of the section above.
Turns are now an API affordance. "Although GPT-Live-1 is not a turn-based model, it natively supports turn detection, so developers can continue to build around explicit turn boundaries." What was an internal inference-over-the-stream necessity for ChatGPT's own UI is now a shipped feature for third parties. The bottleneck-versus-data-model distinction survives intact and becomes commercial: OpenAI removed turn detection from the inference gate and sells it back as an interface convention. Note what this implies about the bitter-lesson argument on this page — the less-intelligent harness component did dissolve into the model, and then reappeared as an opt-in output of it.
The first number on what dissolving the harness deletes. This page and Harness Shrinkage as Models Improve have argued harness dissolution mostly from system prompts. The launch post carries a customer claim about a codebase: replacing a cascaded STT→LLM→TTS build with GPT-Live-1 "simplified our code base by 80% and removed 23K lines of code" (Tony Stoyanov, Co-Founder & CTO; the company is not named in the captured page). Heavy discounts apply — it is a testimonial in a vendor's release post, the baseline architecture is undescribed, and "code base" is undefined (the whole product or the voice layer). But it is the first quantified instance in this corpus of the interaction harness actually being deleted rather than the model merely being able to replace it, and the direction is the one this page predicts: the coordination code for turn-taking, interruption, handoffs and dialog state goes away when the model owns the conversation.
Connections#
- Interaction Models — the proposed resolution
- The Bitter Lesson — why the harness-based status quo loses
- Time-Aligned Micro-Turns — the architectural move that removes turn boundaries
- Full-Duplex Interaction — the interaction modes the bottleneck currently blocks
- Harness Shrinkage as Models Improve — the general version of "the less-intelligent harness should dissolve into the model"
- AI Employee Framing / Human-AI Accountability Redesign — the org-side mirror: both warn against treating autonomy as the goal and pushing the human to the margin; this page is the interface-side version of the same critique
- Design Concept Grilling — argues the value is in collaborative iteration; this page argues the interface is what blocks it
- Context Window Smart Zone — orthogonal limitation that also makes "fully autonomous, walk away" brittle
- Configurable Human Participation — HAS-Framework's typed human/agent-initiated channels are what an interface with room for the human would route; the measured variation in agent-initiated clarification quality (and in whether the agent asks at all) is this bottleneck scored as an outcome
- GPT-Live — the production system that removed the turn detector
- Live-Path Minimalism — where turns re-enter as a derived application-layer view after leaving the audio path
Sources#
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- How we built a realtime system for responsive voice AI in six months — OpenAI, 2026-07-29 (
case-study): the turn detector removed in production; turns re-derived at the application boundary - Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (
vendor-claim): turn detection shipped as an optional API feature on a non-turn-based model, and the unnamed customer's "80% of our code base / 23K lines" claim against a cascaded build. A testimonial inside a release post — no baseline described, no definition of "code base", and three of the four customer quotes on the page did not render at ingest
Cited by 15
- Interaction Models×3
Today's models "experience reality in a single thread": they wait, blind, until the user finishes…
- Full-Duplex Interaction×2
And in the API from 2026-09-10, where the interesting detail is a concession rather than a claim:…
- GPT-Live×2
GPT-Live — full-duplex; the detector is gone; turn-taking is model behavior. The dissolution Turn…
- Live-Path Minimalism×2
Turn Based Interface Bottleneck — the harness this dissolves from the audio path, and the place…
- The Bitter Lesson×2
Interaction Models — TML cites "the bitter lesson" directly: hand-crafted interactivity systems…
- Thinking Machines Lab×2
Stakes out a different priority than the labs critiqued in Turn Based Interface Bottleneck ("AI…
- Time-Aligned Micro-Turns×2
Contrast: turn-based models see an alternating token sequence with hard turn boundaries; real-time…
- AI Employee Framing
Interface-side mirror: Turn Based Interface Bottleneck — argues humans get pushed out of the loop…
- Opinions on Using AI Tools & the Future of the Software Engineering Role
Interface, not just code. Interaction Models / Turn Based Interface Bottleneck: today's turn-taking…
- Configurable Human Participation
Turn Based Interface Bottleneck — HAS-Framework's typed edges and agent/human-initiated channels…
- Design Concept Grilling
Interaction Models — grilling is collaborative real-time iteration; turn-based interfaces are…
- The Future of Agent Interfaces
Human collaboration · Interaction Models / Full Duplex Interaction · Human senses, speech, screen,…
- Harness Shrinkage as Models Improve
The first counter-datum to the counter-datum, and it is on a different axis entirely (2026-09-21).…
- Human-AI Accountability Redesign
Interface-side mirror: Turn Based Interface Bottleneck — the interface-level version of the same…
- Interaction & Multimodal
Turn Based Interface Bottleneck — Why current AI interfaces limit collaboration: single-thread…
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Interaction / Background Model Split
Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reas…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual…
- Full-Duplex Interaction
Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive…
