H
Howardism
Plate IIEntities中文HOWARDISM

Cline

Open-source coding agent and harness (VS Code extension, bring-your-own-key or ClinePass subsidized inference) that publishes its benchmark hill-climbing as a practice — a Feb 2026 playbook, a Jan 2026 Opus 4.5 campaign (47%→57% on Terminal-Bench, four engineers, two weeks), a July 2026 one-prompt autonomous campaign that took Kimi K3 from 77.5% to 88.8% on Terminal-Bench 2.1, and a September 2026 account of migrating its 11M-install VS Code extension off a 76k-line legacy core onto the shared Cline SDK via a self-built dual-bundle rollout, with a controlled A/B showing task.mistake_limit_reached falling 6.34%→0.62% (10x)

Article metadata
Publication details
Published:August 3, 2026
Filed:Entity
Domain:Entities
Reading:12 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Cline

Sources#

What it is#

An open-source coding agent distributed primarily as a VS Code extension, model-agnostic by design: users bring their own provider keys (OpenRouter, Anthropic, OpenAI, local) rather than being tied to one lab's model. That model-agnosticism is the product position — Cline competes on harness quality rather than on owning the weights, and it markets specific model/harness pairings ("Cline is the best harness to run Kimi K3"). It ships ClinePass, a $9.99/month subscription giving subsidized inference on curated open-weight models (Kimi K3, DeepSeek, GLM, MiniMax, Qwen) at 2–5× standard rate limits with no separate provider keys — an open-weight-first commercial bet, distinct from the frontier-lab subscriptions it sits alongside.

Cline is also an MCP client: it is one of the two clients (v3.35.0) used in the MCP tool-poisoning benchmark documented in MCP Tool Poisoning, so it appears in this corpus both as an attack surface and as a benchmarking party.

Hill climbing as a published practice#

What makes Cline unusual as a source is that it treats benchmark hill-climbing as a documented methodology rather than a marketing output, and publishes the losing experiments alongside the winning ones.

  • Jan 2026 — took Opus 4.5 from 47% → 57% on Terminal-Bench. Four engineers, roughly two weeks, by hand: reading model traces, forming hypotheses, testing fixes.
  • Feb 2026 — published the resulting playbook ("A Practical Guide to Hill Climbing").
  • Jul 2026 — ran the same climb autonomously: one prompt, GPT-5.6-Sol as leader model, 17 unattended hours, ~1B tokens and ~$680, moving Cline + Kimi K3 from 77.5% ($79) to 88.8% ($49.8) on Terminal-Bench 2.1 and matching Moonshot's own vendor-reported SOTA of 88.3%. The full account is in Agent-Authored Harness Optimization.

Cline says it is making this a standard release ritual: a baseline run on every new model, then an autonomous optimization campaign against it.

The 77.5% → 88.8% claim is contested as of 2026-08-04, and the numbers are not in dispute — the attribution is. Wang et al. (arXiv 2607.12227, Ai2 / UW, empirical) run budget-matched baselines on the same benchmark, Terminal-Bench 2.1, with Claude Opus 4.6 and GPT-5.4. Give plain parallel sampling the same inference budget an evolution loop would spend and it beats automatic harness evolution on every model tested (72.3 vs 67.4 average pass@1 without unit-test feedback; 86.0 vs 75.8 with), and an evolved harness transfers +0.6pp to held-out tasks from the same suite. Cline's campaign reports no budget-matched baseline and no held-out split, so it cannot separate "the harness got better" from "we spent 17 hours and ~1B tokens searching." Two qualifications in Cline's favour: Wang et al. never ran Cline's harness or Cline's model, and Cline's five fixes were repairs to genuine defects in a mature production harness — a different task from improving an already-adequate minimal one, and the one place this critique has least purchase. The full three-way weighing is on Agent-Authored Harness Optimization.

How to weight Cline as a source#

case-study. Cline benchmarks its own harness, publishes its own scores, and does so on a suite it has been optimizing against since at least January 2026. Traces, cost breakdowns and the merged PR are posted publicly, which is more disclosure than most vendors offer and still not third-party replication. The July 2026 post also reports its own negative results — an experiment that got no causal credit, a fix that only partially worked, two runs invalidated and thrown out — which is the main reason to extend it credit. Comparisons it draws against other models' benchmark costs are not harness-controlled (Compute-Controlled Benchmarking).

The missing control now has a name and a measurement behind it. What Cline never ran is a budget-matched baseline — the same compute spent sampling repeatedly instead of rewriting the scaffold — and an empirical third party has since shown that on this exact benchmark the baseline wins. That does not make the 88.8% wrong; it makes the causal story ("the harness improved") unsupported by Cline's own design. Read the campaign as a build log with a score attached, not as evidence about what harness evolution buys.

An outside observation of Anthropic's fallback#

Cline attempted the same autonomous campaign with Claude Fable 5 as leader model and abandoned it: the safety classifier "kept downgrading the model to Opus-4.8." One vendor's passing remark rather than a measurement, but it is a third-party instance of the routing mechanism in Capability-Gated Model Fallback biting a real workload — AI evals research on a coding harness, with no cyber or bio framing.

Migrating the VS Code extension: rollout without a platform lever#

A different kind of source from the two above — not an agent optimizing a benchmark, but Cline's own engineering team rewriting the flagship 11M-install VS Code extension and proving the rewrite's value with a dashboard A/B (How We Migrated 11 Million Users to Cline's Biggest Harness Upgrade, Saoud Rizwan, Cline blog, 2026-09-02, empirical, vendor voice).

The problem being solved. The VS Code extension ran on a private ~76,000-line core dating to 2024, while every newer Cline surface (CLI, Desktop, JetBrains) already ran on the shared Cline SDK. A first migration attempt shipped, broke badly, and was rolled back. The harder problem than rewriting the code was that the VS Code Marketplace has no gradual-rollout lever: publish is 100%-or-nothing, versions are monotonic (no revert), and a bad release's only fix is a new higher-numbered release — a remediation loop measured in hours to days.

The workaround: bundle both versions, make the extension its own rollout mechanism. Every install ships a ~46 KB loader plus both legacy/ and next/ bundles; a PostHog feature-flag percentage decides which one activates per window, checked on launch and refreshed only between sessions (never mid-task). If next crashes on activation the loader falls back to legacy in the same window and pins that machine — a functioning product plus a captured crash report. The flag is the kill switch (dial to 0% to abort), and both bundles read/write the same settings, credentials and task storage so cohort membership is stateless. Because VS Code extensions declare their IDE surface statically in one package.json, the build generates a union manifest from both bundles and hard-fails if their declared views or settings schemas ever diverge — a mechanical invariant (see Agent Harness Engineering's "enforce invariants, not implementations") that is what kept two divergent codebases swappable on reload behind one manifest.

The result, measured. Every telemetry event carries extension_variant: next | legacy, letting the pre-existing legacy cohort serve as a live randomized control instrumented after the fact (the team retrofitted task.mistake_limit_reached — three-consecutive-mistakes-and-ask-the-human — onto the old harness with identical semantics specifically to make this comparison possible). Flag steps: 1% (Jul 30) → 5% (Aug 2) → 15% (Aug 5) → 30% (Aug 6) → 50% (Aug 11) → 100% (Aug 23), read off the rollout-dial chart, the article's only dated timeline. With both cohorts near 50/50 and each seeing a full day of production traffic: 6.34% of tasks on the old harness hit the mistake limit versus 0.62% on the new one — 10x fewer, with a per-model breakdown (claude-sonnet-5 11x, deepseek-v4-flash 10x, deepseek-v4-pro 10x, claude-sonnet-4-6 6x, gpt-5.6-sol 6x); the article itself flags event attribution as biased against the new engine, so these are conservative. Cline's own reading, and the part it argues generalizes beyond this one migration: the legacy harness parsed tool calls out of XML tags in the model's text stream, the reliable mechanism for 2024-era models — 2026 models are RL-trained to call tools natively, and wrapping one in a 2024-shaped harness pays a tax on every turn in format issues, parse failures, retries, and eventually the mistake limit. See Harness Shrinkage as Models Improve for the broader thesis this instance sits inside, with a twist: this isn't a prompt getting pruned, it's a tool-invocation mechanism being replaced outright.

Evidence weighting. Frontmatter tags empirical, and the full read doesn't contradict that — this is a genuine controlled comparison (a feature-flag-defined cohort split, dated rollout telemetry, a per-model breakdown, an event-attribution caveat volunteered against its own result) rather than an announcement. It is still a vendor measuring its own product with no third-party replication; weight accordingly, same caveat class as Writer's harness-swap paper and DarwinX elsewhere in the corpus. One source-internal correction: the article's prose says the 50/50 split "held for over a week" — the rollout-dial chart's own dates put it at Aug 11→23, 12 days. Two figures for the legacy core's size don't reconcile: prose and the diagram both say "~76k lines," a separate chart says "75k lines" (same rounded figure, not a real conflict) — but that same chart's "SDK adapter · 42k lines" for the new VS Code-side code does not reconcile with the package-breakdown diagram's "21k-line adapter" (the 21k figure explicitly excludes "the UI, kept intact"; the 42k figure's scope is unstated). Flagged in the raw as unresolved; treat both adapter figures as uncertain rather than picking one.

This is a different Cline source from the July 2026 recursive-self-improvement campaign above and shouldn't be conflated with it: that one is an agent editing Cline's harness against a benchmark score (treated on Agent-Authored Harness Optimization); this one is Cline's human engineering team rewriting the harness and proving it with production telemetry, plus a reusable trick for shipping a risky change on a platform with no native gradual-rollout support.

Connections#

  • Agent-Authored Harness Optimization — Cline's July 2026 campaign is the corpus's first end-to-end instance, and the page that weighs the result
  • Agent Harness Engineering — Cline's competitive position is harness quality on top of other labs' models; the five bugs its agent fixed are textbook harness defects; the September 2026 migration's union-manifest build contract is a second, mechanical instance of "enforce invariants, not implementations"
  • Harness Shrinkage as Models Improve — the migration's core argument (a 2024-shaped, XML-tool-call harness taxes 2026 native-tool-calling models) is this page's thesis with the twist that the fix is a mechanism swap, not a prompt prune
  • Shared Harness, Differentiated Surfaces — the migration's package breakdown (@cline/core, @cline/agents, @cline/llms, @cline/shared feeding VS Code/CLI/JetBrains/Desktop/Hub) is a fourth vendor instance of one shared runtime under many product surfaces
  • Harness Build-vs-Buy — the same package-level line counts, read as the price of the rewrite Cline chose to build rather than buy
  • Agent Quality Flywheel — a different discipline than the flywheel's eval-fix loop, but the same instinct: trust a measured delta (task.mistake_limit_reached pre/post) over an absolute score
  • Kimi (Moonshot AI) — the open-weight model Cline pairs itself with, both in the benchmark campaign and in ClinePass
  • Capability-Gated Model Fallback — Cline is the outside party that reports abandoning Fable 5 for evals research because of classifier downgrades
  • MCP Tool Poisoning — Cline v3.35.0 as one of the two MCP clients in the tool-poisoning benchmark
  • The Open-Weight Frontier Gap — ClinePass is a commercial bet that curated open-weight models are good enough for daily coding work at subscription prices

Sources#

§ end
Cited by 12
Related articles
  • Agent-Authored Harness Optimization

    An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Nine instances (Cli…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…