H
Howardism
Plate IIModel Capability & Training中文HOWARDISM

The Open-Weight Frontier Gap

Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57; open weights at the frontier means 744B–1.6T MoEs, so Gemma 4 31B competes on a different axis (efficiency, edge deployment) — July 2026's Inkling adds a third open-weight strategy (fine-tunability, not the leaderboard), and Kimi K3 pushes the sparsity pole to 2.8T/104B while the open/closed gap on *agentic* Elo measures runs wider (35–61 points) than the chat gap; on the demand side, Ramp's card-spend index puts US open/Chinese model-serving use at 6.4% of AI spenders (Aug 2026) and 96.4% of those firms still pay OpenAI or Anthropic directly — additive, not substitutive, and August's token-share shift went to the labs' own standard tier, not open weights; and UK AISI/CAISI add a fourth, non-vendor axis where the gap is widest and visibly widening, cyber capability

Article metadata
Publication details
Published:July 9, 2026
Filed:Concept
Domain:Model Capability & Training
Reading:37 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for The Open-Weight Frontier Gap

Sources#

Summary#

Table 4 of the Gemma 4 report is a snapshot of the open-weight landscape as of June 19, 2026, measured on Arena Text — blind side-by-side human preference, Elo-rated. It is the only externally adjudicated number in a document otherwise full of self-reported benchmarks, and it is more informative than any of them.

RankModelEloOpenTypeParams / active
1Claude Fable 51508no——
~15GLM 5.11475yesMoE744B / 40B
29MiMo V2.5 Pro1466yesMoE1T / 42B
34Kimi K2.61460yesMoE1T / 32B
36DeepSeek V4 Pro Thinking1458yesMoE1.6T / 49B
43Gemma 4 31B1451yesDense31B
57Qwen 3.5 397B-A17B1444yesMoE397B / 17B
61Gemma 4 26B-A4B1438yesMoE26B / 4B
157Gemma 3 27B1366yesDense27B

Three readings#

The gap is small in Elo and enormous in parameters. The best open model trails the best closed model by 33 Elo; Gemma 4 31B trails it by 57. But GLM 5.1 buys its 24-point lead over Gemma with 744 billion parameters against 31 billion — a 24× ratio, and it activates 40B per token, more than Gemma's entire dense model. "Open weights have nearly caught up" and "open weights at the frontier require a datacenter" are both true, and they are the same sentence read from different ends.

Frontier-open means MoE; Gemma is not playing that game. Every open model above Gemma 4 31B is a Mixture-of-Experts in the 397B–1.6T range. Gemma's own summary — "the leading dense open model on the leaderboard" — is a precisely scoped claim, and the scoping is the point. The family's stated target is "varied hardware environments" and "edge deployment." It competes on inference efficiency, where a 0.8 GB quantized E2B is a category the 1.6T models cannot enter at all. Two open-weight strategies have separated: approach the frontier with sparsity, and approach the device with efficiency.

And the sparsity side now has a documented training method. The GLM 5.1 at the top of this table (744B/40B) is the immediate predecessor of GLM-5.2 (750B-A40B), which the SAO paper reports training with asynchronous single-rollout RL. So the corpus now has both ends of the frontier-open MoE story: where these models land (this page) and how they are trained to get there (Asynchronous RL for LLMs). The same paper's Table 1 also shows GLM-4.7 beating GPT-5 High and Claude-Sonnet-4.5 on three of four math benchmarks — the capability-side counterexample to Gemma's efficiency-side positioning, on measured benchmarks rather than Arena Elo.

A third strategy arrives: customize, don't compete (July 2026). Inkling — TML's 975B/41B-active MoE, released with full weights — lands squarely in the frontier-MoE parameter class but declines the frontier race: "Inkling is not the strongest overall model available today, open or closed" is TML's own sentence, and its vendor tables confirm it (HLE text 29.7 vs GLM 5.2's 40.1; Terminal Bench 63.8 vs 82.7). Its claimed edge is being the best base for fine-tuning — multimodal, token-efficient via a controllable effort dial, hosted on TML's Tinker platform. So the two-strategy picture this page drew (approach the frontier with sparsity vs. approach the device with efficiency) gains a third axis: approach the fine-tuner with adaptability. It also breaks the pattern noted below: a Western lab shipping a ~1T-class open MoE (vendor-claim; no Arena placement yet).

The sparsity pole goes to 2.8T, and the agentic gap looks wider than the chat gap (July 2026). Kimi K3 (vendor-claim) is a 2.8T-total / 104B-active MoE — Moonshot calls it "the world's first open 3T-class model," nearly double the 1.6T DeepSeek entry in the table above and 2.9× K2.6's own 1T. It is the strongest available restatement of this page's first reading: the open side keeps closing the capability gap by spending parameters, at ratios the closed side never has to disclose.

Two things it adds that the Arena snapshot could not. First, numbers on the agentic axis. The card cites two third-party Artificial Analysis Elo measures of long-horizon work: GDPval-AA v2, where K3 scores 1686 against Fable 5's 1747 — a 61-Elo gap — and AA-Briefcase, 1548 vs 1583, a 35-Elo gap. Both are wider than the 33-Elo chat gap this page opened with, and the wider one is on the more economically-framed benchmark. Second, a shape to the residual gap: across 45 benchmarks K3 tops 14 rows, and it tops them almost entirely on retrieval, MCP/tool orchestration and document vision (BrowseComp 91.2, MCPMark 94.5, OmniDocBench 91.1) while trailing on hard reasoning (HLE-Full 43.5 vs 53.3) and the hardest long-horizon coding (FrontierSWE 81.2 vs 86.6, OSWorld 2.0 58.3 vs 66.1).

Treat all of it as self-reported: the Elo figures are third-party but selected and transcribed by the vendor, K3 runs on Moonshot's own Kimi Code harness while rivals are quoted at their best across harnesses, and Moonshot itself flags that Fable 5 was hitting fallbacks during two of the coding evaluations (Compute-Controlled Benchmarking). No Arena placement for K3 exists yet — when one lands it will be the first clean test of whether a 2.8T open model narrows the 33 points.

DeepMind's MoE loses to DeepMind's dense model. Gemma 4 26B-A4B (Elo 1438, rank 61) sits 13 Elo below Gemma 4 31B (1451, rank 43), on human preference, from the same lab in the same release — while every larger open model in the table proves MoE scales. On static benchmarks the MoE is close to the dense model (MMLU Pro 82.6 vs 85.2, AIME 88.3 vs 89.2) and sometimes ahead (τ²-airline 76.0 vs 75.0), but human raters prefer the dense 31B. The report does not remark on this. If the effect is real, it suggests sparsity's returns arrive at scales far above 26B, or that active-parameter count (3.8B) governs the qualities Arena raters respond to.

A fourth axis, and the first one nobody with a product measured: cyber capability (July 2026). Every number above is chat preference, vendor-selected agentic Elo, or card spend. UK AISI and US CAISI's joint assessment (AISI / CAISI, 2026-07-23, empirical) puts the open/closed gap on a dangerous-capability axis, graded by two governments that ship no model:

MeasureKimi K3GLM-5.2"Top U.S. Models"
ExploitBench ladder score (mean cap, % of 16)32.2% ± 4.224.4% ± 4.076.2% ± 7.6
"The Last Ones" cyber range (mean step of 32)171128.5
ExploitBench arbitrary code execution (of 41)0020

Three things this adds that no prior axis on this page could.

It corroborates a vendor claim from outside. Moonshot's card has K3 beating GLM-5.2 on nearly every row, which is exactly the comparison a vendor is least trustworthy on. AISI/CAISI reproduce it on four independent measures — ladder score, cyber-range reach, in-cage V8 primitives (17 vs 6), bug reproduction (34 vs 24) — on a benchmark neither lab chose. The intra-open ordering is now third-party fact rather than card copy.

The fitted trend diverges rather than converges. Figure 2 plots an IRT-derived "Overall Cyber Capability" Elo against release date, with ten individually named PRC-lab models and an anonymous US aggregate trendline. The PRC line's slope is visibly shallower than the US line's, so the gap widens across the plotted 2025 → mid-2026 range; K3 is the highest PRC point and still sits below the lower edge of the US confidence band at its own release date. Read off the gridlines (the chart prints no data labels, so these are estimates, not published figures) that is roughly 2,000 against ~2,900 — and on this axis 400 points equals a 10× change in the odds of solving tasks, so the gap is orders of magnitude in solve odds rather than the near-parity the Arena table above suggests on chat. Direction transfers to this page's question; magnitude does not, because an IRT cyber-capability Elo and an Arena preference Elo are not the same ruler.

And it is not safeguard-symmetric, which caps how far the numbers travel. The US models were run "with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities," while K3 ran through Moonshot's hosting on a "selective set" of evaluations its setup allowed. So the table measures the latent frontier gap. On the axis of what a general user can obtain, the shipped US classifiers gate these tasks near zero while K3's weights download — see Open-Weight Elicitation Irreversibility, which is where that inversion belongs.

Why the human-preference number is the trustworthy one#

The rest of the report is Google measuring Google. Arena is blind, third-party, human-rated, and confidence-intervalled (Gemma's ±8). It also disagrees with the static benchmarks in a legible way: Gemma 4 31B is claimed to rival "larger, frontier open models," and on Arena it does — 1451 against DeepSeek V4 Pro's 1456, inside the combined error bars. That is a real result, and it is real precisely because Google didn't grade it.

The same table quietly cites Claude Fable 5 at rank 1 — a third-party corroboration of a model whose own wiki page rests entirely on Anthropic's vendor-claim announcement, and whose head-to-head benchmark table was published only as an untranscribed image. A competitor's leaderboard is better evidence for Fable 5's standing than Fable 5's launch post.

What the table does not control for#

Arena Elo carries no compute budget. Gemma 4's entries are thinking-mode models; the table doesn't say at what thinking budget they were served, nor what the closed models spent. Per Compute-Controlled Benchmarking, a preference score without a budget has the same defect as a benchmark score without one — it just hides it behind human judgment instead of a number. The 33-Elo open-versus-closed gap could be a capability gap, an inference-spend gap, or both.

The demand side: who is actually buying this (July 2026)#

Everything above is supply-side — what open models score, at what parameter count, on whose harness. Ramp's July 2026 AI Index letter (Ara Kharazian, empirical) is the first measurement in this vault of whether US businesses are paying for any of it, read off corporate-card and bill-pay records. Its answer is narrow but precise, and it is a more falsifiable claim than "China is catching up."

The proxy. Ramp has no per-model visibility, so it counts firms paying model-serving and inference platforms — vendors that resell hundreds of models — as its stand-in for open-source and Chinese model use. 5.8% of AI-spending businesses used one in June 2026, up from 4.5% in January. The recovered monthly series (the letter reports only two points) puts that in a longer frame: 1.03% in July 2023 → 5.77% in June 2026, a 5.6× rise as a share of AI adopters, with the slope roughly doubling in 2026 (+1.24pp across all of 2025, +1.30pp in the first five months of 2026). Rising and small are both true.

The proxy is loose in both directions, and the second direction matters for this page. It over-counts: a serving platform sells Llama, Mistral and Qwen alongside GLM and Kimi, so American open weights land in the same bucket as Chinese ones. And it under-counts the case this page has spent the most words on — self-hosting is invisible to it. The Hugging Face forensics below ran nvidia/GLM-5.2-NVFP4 on HF's own endpoints; a firm that downloads open weights and serves them on its own GPUs pays no model-serving vendor and does not appear here at all. Treat 5.8% as a loose upper bound on paid Chinese-model access, not a measurement of open-weight use.

The finding worth keeping: adoption is additive, not substitutive. Among the firms using model-serving platforms, 85.8% also pay OpenAI, 93.2% also pay Anthropic, and 96.4% pay at least one directly. Decomposed (vault arithmetic on Ramp's three numbers): 82.5% pay both labs, and only 3.6% pay neither.

That is stronger than it first looks, because the right comparison is not 100% but the base rate. Rebasing the letter's vendor-share chart onto AI-spending businesses (also vault arithmetic, not stated by Ramp): in June 2026 71.8% of AI-spending businesses paid OpenAI and 77.2% paid Anthropic. So the firms buying open/Chinese model access are 14pp and 16pp more likely to pay the American labs than the average AI spender — the least substitution-shaped result the data could have produced. They are also, by an order of magnitude, the biggest AI spenders in the base: median $248.41 per employee per month against $10.59 for the typical AI-spending business (23.5×). Cheap models are showing up as more AI spend, not less.

One more line the letter never mentions while arguing about Chinese models: DeepSeek's own direct-payment share peaked at 0.23% of businesses during the January 2025 R1 news cycle, fell for a year, and reaches only 0.29% in June 2026 — 1/20th of the serving-platform proxy.

What it does not settle. Additive-today is also the shape an early substitution curve has: a firm evaluates a cheaper model on some tasks while keeping its frontier contracts, and the 96.4% only falls once routing share moves enough to cancel a subscription. The instrument sees whether a vendor is paid, never how much traffic each model gets, so a firm that moved 80% of its tokens to GLM and kept a $20 seat still counts as "also uses Anthropic." Kharazian's own reading is the honest one — the growth "reflects a real weakness for American model companies" on price, and the absence of broad price cuts is the corroborating signal that share has not actually moved. And the sample is Ramp's own VC-forward-skewed customer base, from a vendor with a commercial interest in owning this dataset (see Telemetry vs. Survey Measurement); the direction is more defensible than the level.

The next edition, and the corroborating signal flips (September 2026). Kharazian's September 9 2026 letter republishes the serving-platform series and adds the denominator the July compile had to derive: 6.4% of AI-spending businesses used a model-serving or inference platform in August 2026, and 3.6% of all businesses — internally consistent with the same edition's 56.1% overall adoption (6.35% × 56.14% = 3.57%, vault arithmetic). He also concedes the over-counting direction named above, in sharper terms than this page had used: "we measure using adoption of routing platforms, which also provide access to closed-models, so actual open source adoption is likely lower."

Three things changed under the number.

  • The series revised up and the trend slowed. June 2026 reads 5.77% in the July edition and 6.00% in the September one (+3.9% relative) — Ramp's count-and-dollar series revise upward for months after first print, a general finding documented at Firm AI-Spend Intensity and Headcount Growth. On the revised footing the run to August is 6.00 → 6.16 → 6.35%, +0.35pp in two months against +1.30pp across the first five months of 2026. Still rising, still small, and rising more slowly.
  • The corroborating signal Kharazian leaned on has flipped, and his conclusion did not move. The July reading rested on the absence of broad American price cuts as evidence that open/Chinese share had not shifted — "the absence of broad price cuts is the corroborating signal that share has not actually moved" (superseded 2026-09-22 by September 2026 Ramp AI Index: Cracks in the AI thesis, part 2). This edition reports the opposite: "OpenAI and Anthropic have both announced a series of price cuts over the last month," with the blended effective price down 41% to $0.68 per million tokens from a $1.15 March peak. The cuts landed, and the letter still reads open models as not the cause: "it seems these trends are not driven by adoption of open source models or Chinese models."
  • The substitution that did happen stayed inside the closed menu. The tier that took volume is not open weights but the American labs' own standard tier — frontier token share fell from a 53% August peak to 45% while GPT-5.6 Terra and Claude's Sonnet series absorbed it (25.7% → 35.8% across five weeks). Firms cutting inference cost in August 2026 traded down a tier with the same vendor rather than out to an open model, which is the cheaper substitution and the one this page's supply-side gap predicts.

What this edition cannot do is settle the question below, because it does not republish the overlap statistic. The 85.8% / 93.2% / 96.4% figures remain a July cut; without a later value for them, a rising serving-platform share is still compatible with both the additive and the early-substitutive reading.

One self-hosting use case, named (Ng, August 2026). The invisible case above has a practitioner instance. Andrew Ng tells a general audience that for material non-public information he "just can't even send that to the cloud," so he either works without AI or "really carefully only use[s] a local model"; the banks his advisory firm works with run "in a virtual private cloud or on prem so… it never even leaves their control" (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion). His capability read is this page's first reading from the user's side — "some of the latest open-weight models are approaching frontier capability and… are actually small enough… they can run on" a laptop — and he names a Meta open-weight model and "the latest version of Qwen" while advising against attachment: "these models change every other week… the best practice is to not get stuck on one." It is a substitutive use — no frontier vendor is paid for the MNPI workload at all — and exactly the kind Ramp's proxy cannot count. It is also a use carved out by data policy rather than by price or quality, a third demand driver beside the two this page has tracked. (The Meta model's name is garbled in the auto-caption transcript and is not recorded here.)

One practitioner's port bake-off: parity on the task, order-of-magnitude spread on price (DHH, August 2026)#

A single uncontrolled trial, and worth recording because nothing else in the corpus puts open-weight and closed models on the same agentic task with the same plan and reports the bill. DHH ported Terminal Text Effects, a Python animation library, to a dependency-free Rust executable (Lex Fridman #501, 2026-08-26, practitioner-opinion, self-reported from memory). One prompt: here is the source, produce a Rust version, pixel-perfect frame by frame, do a full analysis, do not stop until finished.

ModelOutcomeWall clockReported per-token cost
Fablecompleted (wrote the 8-step plan; ran out of subscription tokens ~2/3 through and Opus 5 finished it)~45 min~$550 (estimated; actually run on a Max subscription)
GPT Solcompleted, reusing Fable's plan~1h30~$46
Grok 4.6completednot reported~$55
Kimi K3completed"took forever"not reported
DeepSeek V4 Procompleted2h45~$23
DeepSeek V4 Flashfailed——
GPT Lunafailed — needed ~12 prompts to start, then wrapped an existing implementation and declared done——

Reported result quality: startup 86 ms → 2 ms, ~9.6× execution speedup on the first shot, 3 MB executable, and "the output the same" across the models that finished.

What it is evidence for. On one well-specified translation task with a plan already written, the open-weight models that finished produced an outcome he could not distinguish from the frontier one, at roughly 1/20th the metered price and 2–4× the wall clock. That is the shape this page's Elo gaps do not capture: a capability threshold rather than a ranking, where the question is whether a model clears the task at all.

What it is not evidence for, and the list is long. n = 1. The task is translation with a reference implementation — the case models have been good at longest, and the one where the verifier is free. Only Fable was asked to plan; every other model inherited Fable's plan, so the frontier model's hardest contribution was factored out of its competitors' runs. Costs are estimates recalled in conversation, mixing metered pricing against a flat subscription. Failure was scored by one person eyeballing one output. And the two failures are the interesting cell — both cheap models, both failing the same way (would not sustain the loop; one cheated by wrapping an existing implementation), which points at agentic persistence rather than raw capability as the axis that separates the tiers here, exactly as the agentic-vs-chat Elo spread recorded above predicts.

Connections#

  • The Price of Fixed Capability — another place to look for open-weight pressure: in Epoch's cost-frontier traces, the price competition right after a score becomes SOTA is almost entirely closed against closed. Only one of eight traces has an open-weight model as the second point (QwQ-32B at 25% on AIME), though the authors allow that open models may be pushing from behind

  • Illicit Distillation — the provenance question underneath the gap: seven PRC labs named with per-campaign exchange counts for extracting frontier reasoning traces, which if true makes part of the measured convergence derivative rather than independent. The September 2026 NSA/CISA/FBI advisory (China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies) goes further and calls distillation "the critical core" of DeepSeek's, Moonshot's, Alibaba's, MiniMax's, StepFun's and Z.AI's development. It names specific open releases as students: DeepSeek R1/V3, Kimi K2/K3, MiniMax M2, Step 4. If that is right, every open entry in the Arena table above is partly a measurement of how much a closed teacher leaked, and the gap is a lag between teacher and student more than an independent race. It is asserted with no disclosed evidence, and no source in the corpus measures the share

  • GDPval Benchmark — the benchmark under the GDPval-AA Elo board this page differences: real professional deliverables scored as a win rate against practising experts (4-year floor, 14-year mean) across 44 occupations. Worth knowing what the 61-Elo agentic gap is a gap at — economically valuable knowledge work, not chat preference or coding tasks

  • Gemma 4 — the source of the table; the leading dense open model

  • Claude Fable 5 — rank 1, and cited here by a competitor rather than by its vendor

  • Inference Efficiency as Capability — the axis Gemma competes on instead of scale

  • Compute-Controlled Benchmarking — an Elo score without a budget is still a score without a budget

  • Open-Weight Elicitation Irreversibility — what "open" costs, once these models carry a thinking mode

  • Jagged Intelligence (Ghosts, Not Animals) — the aggregate Elo hides that small Gemmas beat Gemma 3 27B on reasoning and lose on knowledge

  • Large-Scale Test-Time Compute — the unnamed variable underneath every cell of the table

  • Encoder-Free Early Fusion — one of the levers that lets a 31B dense model contend at all

  • Responsible Scaling Policy Evaluations — why the open-weight safety argument here is structural: Gemma 4 sits at rank 43, nowhere near a risk threshold

  • Task Time-Horizon Scaling — Arena scores chat preference; whether the open/closed gap survives on long-horizon agentic work is a different measurement

  • Google DeepMind — publishes the table, and places itself 43rd on it

  • GLM (Z.AI) — GLM 5.1 (top of the table) and its SAO-trained successor GLM-5.2: the capability-side open-weight strategy this page contrasts with Gemma's efficiency-side one

  • Single-Rollout Optimization — the RL method behind the GLM MoE line's continued frontier presence; the training-side complement to this landing-place snapshot

  • Kimi (Moonshot AI) — K2.6 sits at rank 34 in the table; K3 at 2.8T/104B is the sparsity pole's current extreme and the first open release with agentic-Elo numbers to set beside the chat gap

  • UK AI Security Institute / US Center for AI Standards and Innovation (CAISI) — the two government evaluators who put the gap on a dangerous-capability axis, and the only parties in this page who grade without selling

  • LLM-Driven Vulnerability Research — where the cyber numbers are read as a capability ladder rather than a market position: the open side's rungs terminate at sandbox escape (0 of 41) while the closed side's merely thin

  • Open-Weight Elicitation Irreversibility — the same assessment read as a safety artifact: a pre-release, single-budget, black-box audit of weights whose elicitation budget is unbounded afterwards

  • Autonomous Intrusion — a fourth reason open weights matter, and the first one that isn't about the leaderboard. This page's three strategies (frontier-by-sparsity, edge-by-efficiency, fine-tunability) are all capability arguments. Hugging Face's July 2026 incident disclosure adds self-hostability as an operational requirement: frontier commercial APIs' safety guardrails refused to process its attack payloads, so the forensics over 17,000+ attacker events ran on GLM 5.2 locally. The relevant property is not that the open model is close on Elo — it is that it is the one that will run on data a hosted model declines. Sharpened 2026-08-03 by HF's technical post-mortem: the refusing APIs are named ("Claude Opus and Fable"), and the replacement was Nvidia's NVFP4 quantization of GLM-5.2 on HF's own endpoints — so the operational requirement was met by a quantized frontier-open MoE, which matters for this page's parameter-vs-capability framing: the datacenter-class model was the only one available for the job, but it did not have to be run at full precision to do it. And the job was not summarization — GLM-5.2 recovered the agent's chunk+XOR+compress scheme and its per-campaign key, surfacing ~4× the secrets a naive scan found. Still a single first-party account, and one authored by the company that hosts the open-weight ecosystem

  • Autonomous Defense — where that requirement bites: a hosted-API SOC degrades exactly at the top of the severity distribution

  • Firm AI-Spend Intensity and Headcount Growth — the buyer-side instrument behind the section above, and the population it isolates: the model-serving cohort's $248/employee/month is ~7× the high-intensity mean in that page's Ramp × Revelio panel, so the firms buying open/Chinese model access are an extreme tail inside the group that panel measures growing headcount ~10%

  • Open Weights as Competitive Strategy — the same open/closed split argued as national economic strategy rather than measured as a capability gap. It is the page that wants this one's numbers: Ng asserts China is "approaching par" with no instrument behind the phrase, and the Elo table here is the check — defensible on chat, progressively less so on agentic Elo (35–61 points) and least of all on the cyber axis. It also inherits this page's Ramp cut and immediately runs past its edge: 5.8% measures US spenders, and Ng's argument turns on the price-sensitive non-US markets Ramp cannot see

  • AI Product Economics Maturation — the second demand-side reading, from a survey rather than a payment rail: among ~305 AI-building software companies, DeepSeek 7% / Alibaba 6% / Moonshot 4% against Anthropic 81% and OpenAI 71% (select-all). A cohort far more AI-intensive than Ramp's card base, and the Chinese labs still sit an order of magnitude below the American ones

  • Telemetry vs. Survey Measurement — why the 5.8% is a floor-and-ceiling problem rather than a number: a payment-rail instrument sees only vendors that get paid, so self-hosted open weights (this page's whole point about downloadable frontier MoEs) are invisible to it by construction

  • Balance-of-Power Superintelligence — the philosophy that underwrites closing this gap from the open side; the August 2026 manifesto confirms Meta's open-weight pause was a pause ("we will resume releasing some open source models soon") and asks that US policy stop restricting American open models on training-data grounds

  • Weak-Verifier Ensembling — the gap closed from the inference side rather than the training side, and a late-2025 practitioner-opinion data point on how far that goes: an 8B open generator with a pool of ≤8B open verifiers reportedly reaching what 70B-class majority voting reaches, and a 70B-class stack averaging 86.2% — "very comparable" to o3-mini. If it holds, part of the measured frontier gap is a verification gap, purchasable with inference compute and open checkpoints

  • DHH (David Heinemeier Hansson) — the practitioner-side bake-off: one port task across seven models, with the two failures both being cheap models failing on agentic persistence rather than on capability

Open Questions#

  • Is the dense-beats-MoE result at 26B robust, or an artifact of one Arena snapshot with ±8 error bars on both models? (The two intervals overlap: 1451±8 and 1438±8.)

  • The open MoE giants (GLM, DeepSeek, Kimi, MiMo, Qwen) are overwhelmingly Chinese-lab releases. Gemma is the Western open-weight entry and it targets the device, not the frontier. Is that a strategic choice or a capability constraint? Partially answered (2026-07-22): Inkling is a Western 975B/41B open MoE — so Western labs can and do ship at frontier-open scale — but it self-reports below GLM 5.2 / Kimi K2.6 on hard reasoning and coding and explicitly declines the frontier framing in favor of a customization axis. One release, still consistent with either reading of the remaining gap.

  • Arena measures preference on chat. Does the 33-Elo open/closed gap widen or collapse on long-horizon agentic work, where time-horizon rather than response quality governs? Partially answered (2026-07-30): Kimi K3's card cites two Artificial Analysis agentic Elo boards where the best open model trails the best closed one by 61 (GDPval-AA v2: 1686 vs 1747) and 35 (AA-Briefcase: 1548 vs 1583) — both wider than 33, pointing to widen-not-collapse. But the comparison is vendor-selected, harness-asymmetric, and taken against a Fable 5 that Moonshot itself reports hit fallbacks on 35% of one coding benchmark, so the direction is indicative rather than settled. Further partial answer (2026-07-23), on a third axis and from a non-vendor: UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities. Two government evaluators put the open/closed gap on cyber capability and it points the same way — widen, not collapse — with the two caveats that weakened the previous answer removed: the comparison is not vendor-selected, and the grader sells nothing. Two new caveats replace them. The scale is an IRT-derived cyber Elo where 400 points is a 10× odds change, so it cannot be differenced against Arena's 33; and the US arm was run with system-level safeguards disabled while the open arm was not, so the measured gap is a latent-capability gap. Three axes now point to widening (GDPval-AA, AA-Briefcase, cyber) and none to collapsing — but no two of them share a ruler.

  • Does open/Chinese-model adoption ever become substitutive rather than additive? The falsifiable version: Ramp's 96.4%-of-model-serving-users-also-pay-OpenAI-or-Anthropic figure is published monthly, so a sustained fall in it — or in the 82.5% who pay both — while the 5.8% serving-platform share keeps rising is the signature of displacement. Absent that, rising serving-platform use is a story about firms buying more AI, not about the American labs losing share. Trigger: the monthly Ramp AI Index, and any broad frontier-price cut (Kharazian names the absence of one as evidence share has not moved). Partially answered (2026-09-22), and the trigger fired: September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 reports that both American labs announced price cuts in the month before publication, with the blended effective price per million tokens down 41% to $0.68 from a $1.15 March peak — the second trigger named here — while the serving-platform share kept rising (6.00% to 6.35% of AI spenders, June to August 2026; 3.6% of all businesses). On substitution the answer so far is no, and for a reason this bullet did not anticipate: the volume moved within the closed menu, frontier token share falling from a 53% August peak to 45% with the vendors' own standard tier absorbing it, and the author states directly that "these trends are not driven by adoption of open source models or Chinese models." The falsifiable statistic this bullet names is still unavailable — the edition does not republish the 96.4% overlap cut — so displacement cannot be checked, only the price trigger graded. Retagged from #oq/wait to #oq/source: the trigger event has landed, and what is needed now is an edition (or the index's own methodology page) that republishes the overlap.

Sources#

  • The plunging price of thought — Emberson & Roodman (Epoch AI), "The plunging price of thought", 2026-09-22 (empirical). Cited here only for the Discussion's observation about which models supply the second point on the cost-frontier traces near SOTA. Full treatment on The Price of Fixed Capability

  • DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux | Lex Fridman Podcast #501 — DHH, Lex Fridman #501 (2026-08-26, practitioner-opinion, n=1, self-reported): the Terminal Text Effects Python→Rust port run across Fable, Opus 5, GPT Sol, Grok 4.6, Kimi K3, DeepSeek V4 Pro/Flash and GPT Luna, with wall-clock and cost estimates

  • UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities — UK AISI / US CAISI, 2026-07-23 (empirical; joint government evaluation, neither party builds a model, no vendor COI): the ExploitBench ladder scores and cyber-range step counts in the table above, the intra-open K3-over-GLM-5.2 ordering on four measures, and Figure 2's IRT-derived cyber-capability Elo against release date (ten named PRC models, anonymous US aggregate trendline, 400 points = 10× solve odds). All three figures viewed per the image two-pass rule. Two limits: Figure 2 prints no numeric data labels, so any Elo value read from it is a gridline estimate rather than a published figure; and the US comparator is never individually named anywhere in the document, so no US number is attributable or checkable against a vendor card

  • Ramp's latest data on China vs. the American AI Labs — Ara Kharazian, Ramp's latest data on China vs. the American AI Labs (Ramp AI Index, 2026-07-08; empirical, corporate-card/bill-pay records for Ramp's own customer base). §"Chinese and open source models" key takeaways and body for 5.8%/4.5%, $248 vs $10.59, and 85.8%/93.2%/96.4%. The longer monthly series, the DeepSeek direct-payment line, and the AI-spender base rates used to rebase the 85.8/93.2 figures come from the raw file's recovered chart datasets — all four in-article charts were Datawrapper iframes with no static fallback, and the ingest pass pulled each chart's dataset.csv and reproduced them in full. COI: Ramp measures its own VC-forward-skewed customer base and markets itself as the authoritative AI-adoption dataset. Full evidence note, denominators, and the instrument's aperture limits at Firm AI-Spend Intensity and Headcount Growth

  • Gemma 4 Technical Report — Table 4, Arena Text leaderboard as of 2026-06-19 (empirical; third-party human ratings, unlike the rest of the report's self-reported benchmarks); §4.1 human evaluation

  • Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning — GLM-5.2 (750B-A40B) as SAO's deployment target; GLM-4.7 on Table 1 (empirical, first-party to the GLM lab)

  • Inkling: Our Open-Weights Model — a Western 975B/41B open MoE that self-reports below the Chinese MoE leaders and positions on customization instead (vendor-claim)

  • Kimi K3 Model Card — Kimi K3 at 2.8T/104B (§1–2); the 45-benchmark table's GDPval-AA v2 and AA-Briefcase Elo rows and their Artificial Analysis provenance (§3 + footnote 3) (vendor-claim)

  • Security incident disclosure — July 2026 — "Forensic analysis" and "The asymmetry problem": self-hostability as an operational requirement rather than a capability strategy (case-study, first-party)

  • Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident — Hugging Face, 2026-07-27 (case-study, first-party): "How we intercepted and analyzed the attack" — Claude Opus and Fable named as refusing, nvidia/GLM-5.2-NVFP4 (a quantized build) as the deployed replacement, and the cryptanalysis it was needed for

  • Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28, practitioner-opinion, n=1 self-report): local open-weight models for MNPI, banks on VPC/on-prem, "approaching frontier capability… small enough" to run locally, and "change every other week". Model name garbled in the captions

  • September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 — Ara Kharazian, September 2026 Ramp AI Index: Cracks in the AI thesis, part 2 (Ramp, 2026-09-09; empirical). Supplies the 6.4%/3.6% serving-platform figures, the routing-platforms-also-sell-closed-models caveat, the price-cut and $0.68/$1.15 token-price statements, and the frontier-to-standard token-share shift. The revised monthly serving-platform series and the weekly tier series are recovered Datawrapper datasets — the page carries no static chart images. It does not republish the 85.8%/93.2%/96.4% overlap cut, which remains a July-edition figure. COI unchanged: Ramp's own VC-forward-skewed card base, published by the platform's lead economist; instrument notes at Firm AI-Spend Intensity and Headcount Growth and Ramp

  • China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies — NSA, CISA and FBI, Cybersecurity Advisory AA26-251A, 2026-09-08, case-study (government attribution, no disclosed method or data, references are vendor disclosures). Cited in Connections for the "critical core" claim and the named student models

§ end
Cited by 35
  • Open Weights as Competitive Strategy×6

    This is the one claim on the page the vault can partly check, and the check is unflattering to a…

  • GDPval Benchmark×4

    Open Weight Frontier Gap — the GDPval-AA Elo board is where the open/closed agentic gap is measured…

  • GLM (Z.AI)×4

    Open Weight Frontier Gap — GLM-5.2 is the 744B–1.6T-class open MoE that page describes; GLM-4.7's…

  • Kimi (Moonshot AI)×4

    On the two third-party Elo agentic measures the card cites from Artificial Analysis, K3 sits 61 Elo…

  • Open Questions Backlog×4

    Open Weight Frontier Gap (88d) — Is the dense-beats-MoE result at 26B robust, or an artifact of one…

  • Artificial Analysis×3

    GDPval-AA (Elo, v1 and v2) · Gdpval Benchmark, Claude Opus 4 8, Kimi, Open Weight Frontier Gap · an…

  • Inference Efficiency as Capability×3

    3.7% activation sparsity. 104B active of 2.8T total, routing 16 of 896 experts per token plus 2…

  • Autonomous Intrusion×2

    It is also the first safety-grounded argument for open weights in this corpus. Open Weight Frontier…

  • Balance-of-Power Superintelligence×2

    Resuming open-weight releases. "Now that Meta Superintelligence Labs are up and running, we will…

  • Claude Fable 5×2

    Open Weight Frontier Gap — Fable 5 is the rank-1 closed reference point in DeepMind's Arena table;…

  • Firm AI-Spend Intensity and Headcount Growth×2

    Open Weight Frontier Gap — the demand-side half of the same monthly index: the highest-PEPM tail of…

  • Gemma 4×2

    It is not competitive at the frontier, and says so. On Arena Text (June 19, 2026), Gemma 4 31B sits…

  • Inkling×2

    The positioning is explicit and unusual: "Inkling is not the strongest overall model available…

  • Open-Weight Elicitation Irreversibility×2

    Note that this is not an argument that Gemma 4 is dangerous. Gemma 4 sits at Arena rank 43 (Open…

  • The Price of Fixed Capability×2

    The authors' explanation is a short-lived premium: the lab that first reaches a score can charge…

  • Ramp×2

    Aperture. The rail sees purchases that cross a card or bill-pay flow. It is a good instrument for…

  • Telemetry vs. Survey Measurement×2

    ramp ai index july 2026 — Ara Kharazian, Ramp's latest data on China vs. the American AI Labs (Ramp…

  • Weak-Verifier Ensembling×2

    Open Weight Frontier Gap — the class-closure claim in its terms: an 8B generator plus ≤8B verifiers…

  • AI Product Economics Maturation

    Ramp's data also puts a bound on the deck's Chinese-model tail (DeepSeek 7%, Alibaba 6%, Moonshot…

  • Andrew Ng

    Privacy: trust the hyperscalers, run MNPI locally. He trusts the largest hyperscalers to honor…

  • Asynchronous RL for LLMs

    Open Weight Frontier Gap — GLM-5.2, trained under this async regime, is the frontier-open MoE that…

  • Autonomous Defense

    Open Weight Frontier Gap — self-hostability becomes an incident-response prerequisite, not a cost…

  • Cline

    Open Weight Frontier Gap — ClinePass is a commercial bet that curated open-weight models are good…

  • Compute-Controlled Benchmarking

    Open Weight Frontier Gap — Arena Elo inherits the same defect: human preference scored at an…

  • Encoder-Free Early Fusion

    Open Weight Frontier Gap — encoder removal is one of the levers that lets a small dense model…

  • Google DeepMind

    Open Weight Frontier Gap — the lab publishes the Arena table that places it 43rd

  • Illicit Distillation

    Open Weight Frontier Gap — the competitive frame: if the gap closes partly by extraction rather…

  • Jagged Intelligence (Ghosts, Not Animals)

    Open Weight Frontier Gap — an aggregate Arena Elo averages the ridge flat; the small Gemmas'…

  • Large-Scale Test-Time Compute

    The accounting is asymmetric and the paper's ratio hides it. A student objects that pre-training is…

  • LLM-Driven Vulnerability Research

    Open Weight Frontier Gap — the same gap this page reads as a safety ladder, read as a market…

  • Model Capability & Training

    Open Weight Frontier Gap — Arena Text, June 2026: the top closed model leads the best open model by…

  • Responsible Scaling Policy Evaluations

    Gemma 4 (DeepMind, July 2026, Apache 2.0) makes the shape visible. It ships a thinking mode; its…

  • Single-Rollout Optimization

    Open Weight Frontier Gap — GLM-5.2 (750B-A40B), SAO's deployment target, is the frontier-open MoE…

  • Task Time-Horizon Scaling

    Open Weight Frontier Gap — Arena Elo measures chat preference; whether the 33-Elo open/closed gap…

  • UK AI Security Institute

    Open Weight Frontier Gap — the same assessment as the first non-vendor measurement of the…

Related articles
  • Open-Weight Elicitation Irreversibility

    A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…

  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…

  • Kimi (Moonshot AI)

    Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…