Howardism · Vol. 03Plate II · No. 02
Interaction & Multimodal, in order.
Notes11DomainInteraction & MultimodalOpen Qs17Newest23 Sept 2026Oldest13 May 2026
Real-time, multimodal, full-duplex, and human-AI collaboration.
Map of Content for the interaction-multimodal domain — 11 concepts. Curated entry point; see Home for all domains.
- Content-Driven Intervention — Speaking because the content warrants it — correcting a false claim, warning of a hazard, supplying a searched-for word — as opposed to speaking because the conversational structure offered the floor; Peng et al. (2026) separate the two with context-matched monologues and find that across seven configurations of five full-duplex speech families, being addressed and silence move onset while false facts and hazards move it by at most.06, that extra pauses and an explicit 'interrupt me if I'm wrong' instruction do not close the gap, and that even with the floor handed over only.14-.15 of non-empty false-fact replies challenge the claim and.04-.07 of hazard replies warn
- Encoder-Free Early Fusion — Multimodal design with minimal pre-processing instead of large standalone encoders: TML co-trains dMel audio + 40×40-patch hMLP + flow head in one transformer for 200ms latency; Gemma 4's 12B independently discards a 305M audio conformer for on-device memory; Inkling carries the design to 975B open-weight scale; Kimi K3 keeps a 401M MoonViT-V2 encoder at 2.8T and tops the corpus on dense-text-in-image — and Tuna-2 finally supplies the matched-size arm, where a patch-embedding-only 7B beats its SigLIP-carrying twin on 10 of 12 understanding benchmarks (and on OCRBench, killing the dense-text worry) while losing generation, but only above a 1.5B backbone and only after ~2T of 3T pretraining tokens
- Full-Duplex Interaction — Perceive-and-respond simultaneously across modalities — a property of scheduling, not of emitting in speech; proactive interjection, visual-cue reactions, simultaneous speech, live translation and time-aware speech are all special cases of model behaviour; audio-only in production since July 2026 as GPT-Live, and in the API from 2026-09-10 where OpenAI sells duplex as a priced layer and ships turn detection as an optional feature on a non-turn-based model; NVIDIA (Sept 2026) shows a system can be full-duplex and cascaded at once, and its open-weight VoiceChat-11B ships a heuristic BOS/EOS endpointer as a fallback — the detector back by the side door; Peng et al. add a third separation — deciding to speak is not conferred by duplex either; and Google claims visual grounding and 97-language mid-conversation switching as live modes; OmniVChat: native audio+video in, nothing duplex
- Interaction / Background Model Split — Dual-model architecture: a time-aware interaction model stays present while an async background model handles deep reasoning and tools; rich-context-package delegation; 'reasoning-model planning at non-thinking latency'; Inkling is the named background half and OpenAI's GPT-Live the production one, while NVIDIA (Sept 2026) publishes the delegation signal the other two withheld — a
<tc bos>/<tc eos>token pair, transcript-only handoff, prefill-and-repeat return, and a frontend that goes silent rather than staying present; on 2026-09-10 the split becomes a priced product boundary, $0.05/min for the front half with a developer-chosen back half; on 2026-09-15 Google ships two live tiers with no backend named and no price on either half, asserting stay-present a fourth time while selling the hold phrase and step narration that contradict it - Interaction Models — Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via harness; interactivity scales with intelligence only if it's in the model — with OpenAI's GPT-Live (July 2026) independently shipping the audio slice of the same conclusions in production
- Interactivity Benchmarks — FD-bench, Audio MultiChallenge + TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (visual proactivity); TML-Interaction-Small at 0.40s turn-taking latency; the tool-calling half (BFCL-audio, FDB3, EVA-Bench, τ-Voice), where voice agents complete 31–51% of grounded customer-service tasks against 85% for text agents; Peng et al.'s context-matched monologues, on which content cues move speech onset by at most.06; and two vendor scoreboards a week apart — OpenAI's seven self-versus-self GPT-Live-1 cards, and Google's five charts drawn on third-party boards, whose Sierra banking row matches OpenAI's digit for digit while the same configuration scores 67.9% and 86.2% under two benchmark names; and OmniVChat-Bench, 2,800 fully synthesized audio-visual dialogues on a gated tiered rubric, where single-checkpoint deltas are noise
- Live-Path Minimalism — GPT-Live's serving principle — "the voice must flow": the realtime media loop is the only thing on the live path; delegation, compaction, persistence and instance management run asynchronously off it. Stateful-instance handoff (warm a replacement, prefill, run both, cut over) turns compaction and rebalancing into zero-interruption transitions; delegation is a budgeted loop over a pre-warmed prefilled frontier-model session; WARP + Instant Connect collapse WebRTC startup from six round trips to one UDP packet; capacity is concurrent sessions keeping every frame on schedule, not GPU throughput. The boundary is third-party-facing and priced (GPT-Live-1 in the API, 2026-09-10: live path alone at $0.05/min); NVIDIA's with/without-delegation ablation (Sept 2026) prices it: +15 ms turn-taking latency on the live path, paid for with ASR WER 10.80→11.47 and CommonEval 2.87→2.36 off it
- Native Multimodal Modeling: Fusion Depth and I/O Duality — An, Lu, Dong et al. (Tencent Youtu + 5 universities, May 2026) formalize 'native' as two operator definitions — mid-fusion injects encoder features into a joint backbone, early-fusion maps every modality through one unified tokenizer into a shared space, and late-fusion (frozen LLM + grafted head) is excluded outright — then cross that axis with an orthogonal input/output duality (Multi-to-Text, Multi-to-Target, Multi-to-Multi) over a 43-model census; the payoff is the claim that each fusion regime forces its own training signature, with differential learning rates mandatory at mid-fusion (CogVLM's 1/10 encoder rate), z-loss + QK-Norm as preconditions at early-fusion (Chameleon diverges after ~20% of training without them), modality-mixture scheduling replacing differential LR, and RL pathway-locality collapsing entirely once the softmax unifies
- Time-Aligned Micro-Turns — The core interaction-model move: input/output as continuous streams in ~200ms interleaved chunks, no turn boundaries; streaming-sessions inference (upstreamed to SGLang), latency-tuned MoE kernels, bitwise trainer-sampler alignment; Moshi and NVIDIA VoiceChat both land on an 80ms (12.5Hz) frame, and VoiceChat makes the turn boundary a frame-level BOS/EOS target weighted 12.5/7.5 against padding at 1.0
- Turn-Based Interface Bottleneck — Why current AI interfaces limit collaboration: single-thread turn-taking is a bandwidth bottleneck; humans pushed out by the interface, not the work; less-intelligent harness (VAD/turn-detection) should dissolve — and did: GPT-Live removed the turn detector from the production audio path (July 2026), with turns surviving only as a derived application-layer view — and, from the September 2026 API, sold back to developers as an optional turn-detection feature on a model that has no turns
- Why AI Lags at Design — Andrew Ambrosino's four reasons frontier models are worse at visual/product design than at code: design is hard to grade (no clean reward like 'does it compile'), it sat outside the AI-research flywheel labs optimized for, it rewards novelty where code rewards known patterns, and it hides a design↔code abstraction layer (a rebrand is 263 components on the surface, semantic relationships underneath)
Derived#
(none)
Open questions 17 open
- SourceDo the models fail to notice the false claim or the hazard, or notice and stay silent? The text-token probe reads the emission distribution, not comprehension, so it cannot separate them — and the two failures have opposite fixes (better grounding vs. a changed speaking policy). Falsifiable on the released stimuli: probe the internal representation, or hand the same transcript to the same model as a text question and check whether it flags the claim it declined to challenge in speech.
- SourceIs the absence a pretraining-data property or a post-training one? Other-correction is dispreferred in human conversation, so dialogue corpora are thin in exactly this behaviour; but both interactivity-aligned RL arms here lower content-cue onset below their own baselines, which is a post-training effect pointing the same way. Settled by an RL arm that rewards content-triggered onset on matched stimuli and reports whether turn-taking quality survives it.
- SourceDo delegating architectures — a duplex frontend with a text backend holding the content understanding — intervene on content better than duplex-native models? None of the seven configurations delegates, and the frontend-backend systems in this corpus are evaluated on tool calls, never on uninvited correction. The experiment is cheap: run this protocol against a frontend-backend system and against a cascaded production voice stack.
- SourceDoes an encoder-free model at matched size still match? TML co-trains everything from scratch; Gemma 4 freezes encoders on four models and drops them on one, at a different scale; Kimi K3 retains a 401M encoder at 2.8T and leads on dense-text vision. Partially answered (2026-09-21) by tuna 2 pixel embeddings beat vision encoders: for vision, at 7B, on a matched backbone (Qwen2.5-7B-Instruct) and matched data, the encoder-free arm does better than match — it wins 10 of 12 understanding benchmarks against its encoder-carrying twin (MMVet +5.0, CountBench +3.9, OCRBench +1.4) while being ~400M parameters smaller and training one stage fewer, and loses only RealWorldQA (−0.2) and MMMU (−0.4). Three reasons it stays open. The ablation is image-only — Tuna-2 has no audio path, and the audio arm is where the original evidence (a 305M conformer deleted for 0.002 WER) actually lives. It is at 7B, an order of magnitude below the scale where the two labs disagree. And the encoder arm still wins generation (GenEval 0.88 vs 0.87, ImgEdit 4.18 vs 4.09, LLM-judge quality under both judges), so "match" holds for understanding and not for synthesis.
- SourceIs the dense-text degradation intrinsic to a projection-only vision path, or an artifact of the 12B's particular training run? The prediction is falsifiable: an encoder-free 31B should show the same InfographicVQA cliff at 280 tokens. Partially answered (2026-09-21) by tuna 2 pixel embeddings beat vision encoders, graded in three parts. Right: the question was right to doubt that the cliff is a property of the 12B's run alone and to treat the vision path's representation as the variable — the matched 7B comparison isolates exactly that variable and produces a clean signal. Wrong: the "intrinsic to projection-only" horn. A patch-16 projection-only path is best or tied on every dense-text row against its encoder twin (OCRBench 79.7 vs 78.3, AI2D 79.6 vs 79.4, ChartQA 85.6 vs 85.6), so nothing about reading dense text requires an encoder. Right for the wrong reason: this page's own proposed mechanism — a projection performs no feature compression, so resolution must substitute for parameters, making dense text token-budget-bound — survives untouched, because Tuna-2 runs at a single token budget and never sweeps it, and its suite contains no InfographicVQA, DocVQA or OmniDocBench. The falsifier that would settle it is now narrower and cheaper: a vision-token sweep on any encoder-free model, not a 31B.
- ResolvedTML deletes encoders and co-trains from scratch. Gemma 4 deletes encoders and trains the 12B from scratch, but keeps frozen encoders elsewhere. Which half of "encoder-free + from-scratch" does the work? Answered (2026-09-21): the encoder-free half, at least for the vision path — and the two halves are separable. tuna 2 pixel embeddings beat vision encoders initializes both arms from a pretrained text LLM (Qwen2.5-7B-Instruct), so nothing in it is trained from scratch, and the encoder-free arm still beats its encoder-carrying twin on 10 of 12 understanding benchmarks. From-scratch co-training is therefore not a precondition for the encoder-free advantage. Two riders that do not reopen the question but bound it: the advantage is conditional on scale, reversing completely on a 1.5B backbone at 100k steps (OCRBench 56.8 vs 59.2, GenEval 48.2 vs 56.0) and appearing only after roughly 2T of 3T pretraining tokens; and removing the encoder is paid for with an added masking objective, which is worth +1.4/+3.4/+4.2 to the encoder-free arm against +0.9/+1.3/+1.0 to the encoder arm.
- SourceDoes an interaction model that keeps talking during delegation actually beat one that emits a hold phrase and goes silent, on user-perceived quality? Thinking Machines and OpenAI both assert stays-present; NVIDIA measures the silent variant and reports 100% turn-taking with an 83–85% filler rate by design. Nobody has run both policies against the same backend latency distribution and asked users. (2026-09-21: still nobody. gpt live 1 in the api asserts stay-present a third time and markets the absence of a hold phrase — "without narrating every step out loud" — as a product feature, so the two policies are now explicitly opposed vendor positions rather than an unnoticed divergence. No measurement on either side; the question is unmoved and better motivated.) (2026-09-21, second bearing note: gemini 3 8 live announcement asserts stay-present a fourth time — "acknowledge requests and keep chatting while tasks finish in the background" — and in the same paragraph sells an early verbal cue ("Let me check that…") and live progress narration through multi-step background tasks. One vendor now holds both policies simultaneously, which retires the reading that the two are competing design schools and replaces it with a harder problem: the phrase "stays present" is not being used to mean the same behaviour by the three parties asserting it. Any experiment that settles this now has to define the dependent variable before it runs — substantive conversational continuation, or any emission at all. Still zero measurements.)
- SourceIs the duplex-quality cost of adding a delegation token intrinsic to the token, or an artifact of synthetic training data? NVIDIA's ablation pays ASR WER 10.80 → 11.47 and FDB-v1 pause TOR 53.2 → 68.2 for the token, and attributes both to synthetic SFT audio and absent natural-pause data. A second frontend trained with natural-pause and real-speech tool-call data settles it. Partially answered (2026-09-23) — the synthetic-audio explanation is weakened, not the token one. nemotronlabs voicechat (
empirical) trains a sibling duplex model whose speech is also almost entirely TTS-rendered — both CPT and SFT "draw their corpora from the text-only Nemotron backbone rather than from natively recorded conversational speech" — and it nonetheless reaches 15.3% / 25.5% FDB 1.0 pause TOR, the lowest of the five open-weight systems in its table and far below the 53.2 → 68.2 range the ablation moved through. It carries tool calling in a parallel function channel rather than a delegation token, and it adds three conversational augmentations the ablated frontend did not report: early interruption at p = 0.1 (agent turn truncated, continuing 8 frames / 640 ms before EOS), backchannel injection at p = 0.05 per sample, and a two-frame (160 ms) agent-text delay with the function channel deliberately left unshifted. So synthetic audio is not a hard ceiling on pause behaviour, and the augmentation schedule is the better suspect. It is not an ablation — different model, different data, different channel design — so the token half of the question is untouched and the original falsifier still stands.
- Interaction Models3 open
- WaitDoes the interaction/background split generalize, or is it a transitional artifact until a single model is both fast and deep enough? Partially answered (2026-08-04): it generalizes across labs — GPT-Live ships the same split in production (full-duplex voice model delegating to GPT-5.5), independently derived from latency engineering. Whether the split is permanent or transitional remains open; two implementations are evidence of convergence, not permanence. Extended (2026-09-21) — the generalizes half is now settled and the transitional half is sharper. Hu et al. (NVIDIA,
empirical) is a third independent derivation from a third starting point (speech-ASR research, motivated by a parameter-capacity argument rather than latency or serving), and their related-work section names four more concurrent instances plus two cascaded production systems. Three labs, three rationales, one shape: generalization is no longer the open part. Two new facts bear on transitional. Against it: the split behaves as a capability interface — swapping the backend from Qwen2.5-7B to Qwen3-30B-A3B to Qwen3-235B-A22B moves EVA-Bench task completion 40.4 → 57.3 with the interaction half's weights untouched, which is a property a scaffold you intend to absorb would not be designed to have. For it: this implementation buys its agentic ability by giving up the split's own premise — the frontend goes silent during delegation instead of staying present — and the alternative branch is live and named, with DuplexSLA internalising tool calls inside a duplex model. The falsifier is unchanged and now concrete: a single duplex model that matches a delegating system's tool-call accuracy without going quiet. Extended again (2026-09-21) — the capability-interface evidence is now commercial, which is a different kind of durability. OpenAI shipped the split as a product on 2026-09-10 (vendor-claim): the front-end voice layer sells at $0.05/min, the backend is the developer's to choose and pay for, and the post says in as many words that it may be "a third-party model." A vendor that intended the split as a scaffold to absorb would not build a price list around it, publish a delegation call signature (session.commentary.appendwith a delegation id) for arbitrary developer-side agents, or invite a competitor's model into the back half. Against that, it is still only one more assertion on the transitional question — OpenAI is a party with an interest in selling the front half, and the vendor's own benchmark footnotes show the split doing the work (Astra at medium behind the intelligence cards, Terra at low behind the tool-calling cards) rather than the frontend closing the gap alone. Two facts that would settle it are still missing on both sides. Extended a third time (2026-09-21), and this one subtracts rather than adds. Google's Gemini 3.8 Live launch (vendor-claim) is a fourth frontier live-voice product and the first whose public account describes no split: two named live models distinguished by an effort setting attached to the live model itself, no backend named anywhere, no backend priced, and an Extended Thinking tier that "reasons and speaks simultaneously." Read as evidence for the transitional branch, that is weak and it is the first of its kind — a vendor presenting the internalized alternative as a shipping product rather than as DuplexSLA's research road-not-taken. Read carefully, it is a framing and not a disclosure: the same post says the model "executes tools and API calls in the background while continuing the conversation," sells an acknowledgement filler ("Let me check that…") and live progress narration through multi-step background tasks, and footnotes its agentic benchmark runs as executed on the Gemini Enterprise Agent Platform. Background execution plus a hold phrase plus step narration are the three surface signatures of delegation, so the post is consistent with an undisclosed split and with a single model that thinks while it talks, and discriminates between them nowhere. What it changes for this question is the shape of the falsifier rather than its content: it is now clear that vendor framing cannot answer this — the falsifier still has to be a measurement or a disclosure, and Google supplies neither. Extended a fourth time (2026-09-23), and this one is a measurement rather than a framing. nemotronlabs voicechat (NVIDIA,empirical) is the internalized alternative built and benchmarked by the same lab that published the delegating system two days earlier, which makes it the nearest thing to the counterfactual this question has ever had. The falsifier stated above — a single duplex model that matches a delegating system's tool-call accuracy without going quiet — is now half-tested and fails on both halves, in instructive directions. On accuracy the internalized model wins the routing column and loses the task: 82.5 tool-selection F1 against the delegating system's 71.7–74.6 tool accuracy, but 42.2 argument accuracy against 52.8–55.2 and 33.0 Pass@1 against 44.0–48.0 on Full-Duplex-Bench 3.0 (the two papers name the routing metric differently, F1 versus accuracy, so read the second and third columns). On going quiet it does not even draw level: during tool execution it speaks a pre-configured per-tool acknowledgement, pads its agent-text channel, and — beyond what the delegating frontend conceded — stops conditioning response generation on incoming audio, so barge-in is unavailable for the duration of the call. What this settles: internalizing tool calls does not recover the live path, and the capability the split is actually buying is argument grounding and multi-step composition, not tool selection. What it leaves open is the same thing as before — nobody has run both against the same backend latency distribution and asked a user which is better — and the internalized branch now carries its own published ceiling (no more than five tools per session, unreliable simultaneous multi-tool invocation). Standing count: three labs describe the split, one declines to describe anything, and the one lab that built both branches shipped the delegating one with the better Pass@1 and the internalized one with the open weights. - Wait"Interactivity scales with intelligence" is asserted; the larger-model release later in 2026 is the test. Partially answered (2026-09-21) — one axis measured, and it runs the other way. Peng et al. (
empirical) run seven configurations of five open full-duplex speech families and report no intelligence or task-competence score for any of them, so no capability-vs-interactivity correlation across families can be read out of this source; the prediction's actual test, a larger TML release, is untouched. What it does supply is the nearest available proxy, and it is negative. Both interactivity-aligned RL arms — Moshika-RL and PPlex-RL, post-trained on pause handling, turn-taking, backchannelling and interruption — become quieter on content, not more proactive: Moshika-RL drops its neutral onset .07 → .03 while holding direct-question onset at .15 and raising silence onset .12 → .18, with false-fact and hazard onset falling to .01/.01, below its own baseline; PPlex-RL drops neutral .10 → .01 and sharpens its question/neutral contrast from ~3.4x to ~21x while false fact and hazard sit at .02/.04. The useful distinction this forces on the original claim: interactivity in the turn-taking sense (responsiveness, precision about when the floor is yours) is trainable and demonstrably improves with targeted post-training, while interactivity in the self-selection sense (speaking because the content warrants it) does not come along for the ride and may regress. A single word in the assertion is doing two jobs. Full treatment on the Content-Driven Intervention page. Extended (2026-09-21) with the first within-vendor datum, and it shows the measurement instrument failing before the claim does. Google (vendor-claim) ships two tiers on one live architecture and scores both on one third-party composite: Gemini 3.8 Live 76.0 and Gemini 3.8 Live Extended Thinking 82.6 on the Artificial Analysis Speech to Speech Index. Taken at face value that is the claim's shape — more intelligence, higher interactivity score, same family, no harness change. Three things stop it being evidence. First, the composite compresses the effect it is supposed to show: the one component published separately, the τ-Voice agentic chart, moves 30.1 → 68.6 across the same two tiers, +38.5 points against the composite's +6.6, so the index is dominated by something that barely moves. Second, the composite and that component rank two models in opposite orders — 3.8 Live beats Gemini 3.1 Flash Live at High effort on the index (76.0 vs 71.5) and loses to it on τ-Voice (30.1 vs 37.7). Third, nothing in the corpus states the index's composition or weights, and the methodology footnote printed on the charts points at the vendor's own page rather than the evaluator's. So the datum the question has been waiting for arrives inside an instrument that cannot carry it. The sharpened form: intelligence and interactivity must be scored on separated components to test this at all, and the only vendor to ship the two-tier comparison published a single number instead. Full treatment on the Interactivity Benchmarks page. - WaitResearch grant announced for interactivity benchmarks — what becomes the FD-bench equivalent for video proactivity? Partially answered (2026-09-21) — the audio half of the question now has an answer; the video half still does not. Peng et al. (
empirical, arXiv 2609.19596) is the first third-party, reproducible proactivity instrument in the corpus: context-matched monologues where only the trigger utterance varies, ten conditions derived from turn-allocation theory, inter-word pauses compressed to 0.12 s so opportunity cannot confound reason, a single onset scalar with a stated denominator, and stimuli plus evaluation code released. It is the shape the question is asking for — and it is audio-only, 40 English monologues over two synthetic TTS voices. It also narrows what a video equivalent would have to do, since its two load-bearing design moves are modality-independent: hold the context fixed and vary only the trigger, and control the opportunity to speak separately from the reason. RepCount-A, ProactiveVideoQA and Charades do neither — each supplies a standing instruction and then measures timing, which is the addressed case. Nothing here bears on whether the grant produces a video instrument, so the trigger event is unchanged.
- WaitDoes the interaction/background split generalize, or is it a transitional artifact until a single model is both fast and deep enough? Partially answered (2026-08-04): it generalizes across labs — GPT-Live ships the same split in production (full-duplex voice model delegating to GPT-5.5), independently derived from latency engineering. Whether the split is permanent or transitional remains open; two implementations are evidence of convergence, not permanence. Extended (2026-09-21) — the generalizes half is now settled and the transitional half is sharper. Hu et al. (NVIDIA,
- Live-Path Minimalism1 open
- WaitTML upstreamed streaming-sessions serving into SGLang; GPT-Live's stateful serving (persistent sessions, seamless instance handoff, off-path compaction) is proprietary. Does an open-source inference stack ship instance handoff for full-duplex voice? Trigger: an SGLang/vLLM release with session-handoff support.
- ResolvedDoes the upcoming GPT-Live API expose the media/application separation to third parties — application logic customizable behind the async RPC boundary without touching the live path — or is the boundary internal-only? Trigger: GPT-Live API launch. Answered (2026-09-21) — the trigger landed and the first branch is correct: GPT-Live-1 shipped in the API on 2026-09-10 (
vendor-claim) selling the front-end voice layer alone at $0.05/min, with the developer choosing "the models, tools, and agent harness behind the conversation" — OpenAI's own backends or a third-party model — and returning results across the boundary asynchronously vialive.send({ type: "session.commentary.append", delegation_id, content }), demonstrated with an arbitrary developer-side agent (a Codex SDK thread). The boundary is not internal-only; it is the product seam and the price seam. A release post is exactly the right evidence for what a product exposes, so thevendor-claimtier retires this question — it would not retire a question about how well the exposed boundary performs.
- SourceCan a single probabilistic objective (or one unified tokenization scheme, or a continuous latent grammar) carry both understanding and generation without regression on either side? Today's "unified" models are hybrids — NTP for text plus diffusion/flow heads for image and audio (Transfusion, Show-o2, BAGEL) — and the discrete-unified path (Chameleon, AnyGPT, Janus-Pro) versus the continuous-latent path (TUNA-2, Mamoda2.5) is unresolved. Falsifier: a single-objective model that matches the best hybrid on both a generation and an understanding suite at matched scale.
- SourceIs "expert nativity" — how far MoE experts are jointly trained across modalities versus specialized per modality — a real axis with measurable consequences, as the survey proposes it be formalized alongside architectural nativity? Every flagship in the census above ~100B is sparse (1TA32B, 744BA40B, 310BA15B), yet no source here reports modality-aware routing ablations or how sparsity interacts with cross-modal attention.
- SourceCan symmetric M2M behaviour be distilled into a compact model under streaming and full-duplex constraints? MOPD is reported once, on MiMo-V2.5, and only for the M2T projection; the survey calls distilling M2M behaviour "largely uncharted," which is precisely the capability the vault's interaction/background split currently buys with a second model instead.
- Why AI Lags at Design3 open
- WaitAre reasons 3–4 (novelty, the abstraction layer) genuine ceilings, or — like reasons 1–2 — just under-invested capabilities that fall once a lab builds the grader?
- WaitCan design be made gradable without a human in the loop (learned taste models, preference data at scale), or does the "human aspect of taste" resist automation the way research taste might?
- WaitDoes the design↔code abstraction layer improve with better code-understanding models even if pure visual design stalls — i.e. is reason 4 a coding-capability problem in disguise?