資料來源#
- Gemma 4 Technical Report
- Interaction Models: A Scalable Approach to Human-AI Collaboration
- Introducing System One Models & Jev
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
摘要#
Rich Sutton 2019 年的文章指出:運用運算力的通用方法(搜尋、學習),最終會勝過將人類知識和手工設計結構直接納入的方法;而且隨著運算力增加,優勢會大幅擴大。所謂「苦澀」之處在於,這個結果不斷讓投入巧妙領域結構的研究人員感到意外,因為那種結構會成為上限,而非根基。
這個頁面之所以存在,是因為這項原則在 wiki 中一再成為關鍵論據——它被明確援引,作為將 harness 融入模型的理由。
本頁引用之處#
- Interaction Models — TML 直接引用「苦澀的教訓」:手工打造的互動系統(VAD、回合偵測、對話管理 harness)「終將被通用能力的進步超越」,因此「若要讓互動性隨智慧擴展,就必須把它納入模型本身」。參見 Turn-Based Interface Bottleneck。
- Encoder-Free Early Fusion — 在單一 transformer 中從頭共同訓練所有模態元件,而不是拼接預訓練編碼器/解碼器:手工設計的模組邊界更少。
- Time-Aligned Micro-Turns — 移除人為的回合邊界,讓互動模式成為可擴展的模型行為,而非各模式專屬的 harness 程式碼。
- Harness Shrinkage as Models Improve — 將相同邏輯套用到程式碼代理 harness:提示詞鷹架用來彌補模型目前做不到的事,並應隨模型進步而縮減。(但須注意:機械式驗證——測試、型別、linter——不會遷移到模型內部。)
- Agent Harness Engineering —「強制不變條件,而非實作方式」:讓模型自行找出路徑;harness 只編碼必須成立的條件。
標準但書#
苦澀的教訓談的是能力與結構遷移到模型內部,而不是「harness 毫無用處」。合理地留在模型外的項目包括:機械式驗證(Harness Shrinkage as Models Improve 的綜整)、組織專屬政策/風格、安全邊界,以及依據 Claude Character as Product 所述的刻意性格/人格設計。每個 harness 元件都要面對的開放問題是:它屬於界線的哪一側?
部署豁免#
Gemma 4 同時朝兩個方向違背這條界線,讓它更加清晰。同一份報告一方面移除手工設計的結構——550M 視覺編碼器變成 35M matmul,305M 音訊 conformer 則完全刪除(Encoder-Free Early Fusion)——另一方面也增加大量結構:5:1 的區域至全域注意力比例、p = 0.25 的 p-RoPE、全域層中的 values = keys、逐層叢集量化位元寬度,以及 drafter head 中依 token 叢集進行的 top-k 投影。
這並不矛盾,說清楚原因很有幫助。苦澀的教訓針對的是編碼了人類對任務先驗的結構——模態編碼器主張模型看見音訊之前,應先以特定方式處理它;這項主張便會成為上限。KV-cache 和量化技巧對任何事都不編碼先驗;它們只是執行模型已學會之網路的算術技巧。部署工程不受此限;Inference Efficiency as Capability 因此主張,這類工程會持續累積,而非縮減:它沒有可供遷移的「內部」。
對 Harness Shrinkage as Models Improve 的推論是,模型外項目清單中還要再加一項:除了機械式驗證,還有推論路徑本身。
同樣的豁免也適用於訓練迴圈。SAO 在 RL 端進行相同的雙向變動——它移除機制(舊策略模型 π_θ_old、GRPO 的群組基線),同時增加不少東西(凍結注意力評論器、skip-observation GAE、長度自適應 λ、依任務設定的裁切不對稱)。它手工打造的評論器在推論時會被丟棄;就像 KV-cache 技巧一樣,這種結構沒有「內部」可供遷移。界線依然成立:模型學到的任務知識會遷移到內部;負責產生或服務模型的鷹架則不會。
業者的「最苦澀教訓」(TypeSafe,2026 年 9 月)#
TypeSafe AI(2026-09-15,vendor-claim)提出一種反轉說法:**「最佳化正確的任務,比資料、運算力或演算法更重要。」**其論點是,每家實驗室的 RLHF 都在最佳化產出人類評分者偏好的文字——這對聊天而言正確,對自動化而言卻錯誤——因此無論規模多大,都無法讓聊天模型成為可靠的決策元件。這與其說是否定 Sutton,不如說是在主張目標位於界線的哪一側:規模仍是引擎,但獎勵承載著模型無法靠學習擺脫的任務結構。這與上文的部署豁免有相似之處——不屬於任務解法先驗的結構得以保留——但除了業者自己的評測(Typed Decision Verifiers)之外,沒有提出其他證據;而這家公司銷售的產品正是另一種目標。
延伸閱讀#
- Rationale Bootstrapping (STaR) — 這項教訓出現在一種旨在超越它的方法之中:CS329A 的結課講座指出,大型模型更能吸收自我改進飛輪的效益;這項觀察分別適用於 Absolute Zero 自行提出的課程,以及 SWiRL 的逐步 RL。旨在越過資料牆、製造訓練資料的迴圈,其效益會隨基礎模型規模增加而上升;因此,自我改進是擴展的互補手段,而非替代方案。
- Why AI Lags at Design —「這些模型會變得擅長設計」是將苦澀教訓套用於設計落差的押注。
- Evolutionary Proof Search — 苦澀的教訓正預測,這種特製的演化裝置會被吸收進模型。
- Interaction Models — 最近最明確的引用。
- Turn-Based Interface Bottleneck —「較不聰明的 harness 會輸給擴展。」
- Harness Shrinkage as Models Improve — 程式碼代理版本,並附有機械式驗證的但書。
- Agent Harness Engineering — 以不變條件而非實作方式為核心,是理解苦澀教訓後的設計規則。
- Encoder-Free Early Fusion / Time-Aligned Micro-Turns — 以此原則作為架構選擇的依據。
- Claude Character as Product — 一個可能的反例:性格或許不會遷移到模型內部。
- Model Spec Midtraining (MSM) — 對齊從 harness 提示詞注入移至模型內化價值觀,是沿著對齊軸線實踐苦澀教訓。
- Compute Allocator — 指出留在人類界線一側的事物:資源分配決策,以及支援這項決策、面向人類的鷹架,都不會遷移到模型內部,即使面向模型的結構會遷移。
- HTML as the New Markdown —「為模型保留驚喜空間」是這項教訓在提示詞層級的呈現;但書是面向人類的可讀性(HTML 產物)屬於不會融入模型的一側。
- Prototype Over PRD —「說明為何,而非做什麼」是提示詞粒度上的這項教訓:少規定要做什麼,讓模型能力補足,而非將先驗假設照抄到 UI 中。
- Deep Modules for Agents — 模組邊界是能保留下來的部分;教訓要消解的是模組內部手工設計的任務結構。
- MCP and Computer Use — Boris Cherny 所說的「對模型而言,那只是 token」,表示基底選擇(MCP/API/電腦操作)是模型的決策,而非 harness 的決策;這是工具派送上的苦澀教訓終點。
- Agentic Loops Overtake Bespoke Systems — 語料中最明確的實證確認:隨著 LLM 進步,DeepMind 的簡單代理迴圈在開放式數學問題上,追上了其特製訓練系統(AlphaProof + 演化搜尋)。
- AI R&D Autonomy Evaluation (AECI) — 若苦澀教訓一路成立,擴展通用方法最終也會改進自身;AECI 是 Anthropic 用來衡量是否接近該門檻的方法。
- Recursive Self-Improvement — 這項原則最遠的推演:「研究進展大多取決於工具與資源」,因此努力(99%)將能自動化。
- AI Accelerating AI Development — 實證案例:核心最佳化迴圈從 3× 提升至 52×,呈現擴展通用方法勝過手動調校,並有實測佐證。
- The Data Wall and the Validation Commons Are One Supply Constraint — 擴展互補性的發現成為預測論證的關鍵:由於大型模型更能吸收自我改進飛輪的效益,合成資料既無法挽救停滯的擴展曲線,也不會限制正在運作的曲線——因此資料牆被重新估算為驗證供應限制,而非運算力限制。
- Research Taste as the Human Bottleneck — 對最後一道防線的未解押注:研究品味是真正的上限,還是苦澀教訓即將消解的下一種結構?
- Build for the Next Model — 產品策略上的推論:既然能力會在各次發布間遷移到模型內部,就先製作「幾乎能用的東西」,讓下一個模型消解落差,而不是靠工程手段繞過它。
- Task Time-Horizon Scaling — 通用基準能力持續提升,正是手工搭建的鷹架優勢不斷縮小的曲線。
- The Verifiability Thesis — Karpathy 解釋為何擴展 RL 勝過手工工程:實驗室把運算力投入可驗證獎勵的環境。
- Software 3.0 — 將神經網路視為宿主程序的推演,是將苦澀教訓一路推至硬體層。
- Universal AI (AIXI) — 形式化版本:「智慧是在假設/策略空間中搜尋」;透過交錯執行 AIXI 近似算法,可保證隨運算力增加而改進(但暴力搜尋的代價高得無法負擔)。
- Effective Compute Scaling —「擴展就夠了嗎?」是以預測問題的形式提出苦澀教訓;DeepMind 提醒天真暴力搜尋在玩具領域之外會失敗,呼應 Sutton 所說的「搜尋需要良好的先驗」。
- Andrej Karpathy — 經常引用這項原則(可驗證性、幽靈、Software 3.0 都以此為基礎)。
- Repository Exploration Subagent — 一個實際測試案例:FastContext 訓練專門的探索器(手工設計結構),但其「同模型探索」基準顯示,持久的收益來自架構分離,而非訓練出的模型——苦澀教訓預測,隨著基礎模型成本下降,這項收益將會縮小。
- Inference Efficiency as Capability — 例外:不編碼任務先驗的結構(KV cache、量化、drafter)永遠不會遷移到模型內部,而會持續累積。
- Gemma 4 — 同一版本中移除編碼器並增加推論路徑結構,清楚劃出了界線。
- Single-Rollout Optimization — 訓練迴圈中同樣的雙向變動;手工打造的評論器是鷹架,推論時會被丟棄,因此永遠不會遷移到模型內部。
- Asynchronous RL for LLMs — DIS 藉由進一步移除(
π_θ_old、檢查點歷史)而「變得更簡單」;RL 管線和推論路徑一樣,不受這項原則限制。 - Group Relative Policy Optimization (GRPO) — 移除評論器本身就是一次苦澀教訓式的調整;SAO 以實測說明這麼做何時會矯枉過正。
- Typed Decision Verifiers — TypeSafe 的「最苦澀教訓」:決定模型能否成為自動化元件的是訓練目標,而非規模——這是一種業者主張的反轉,已記錄於上文。
資料來源#
- Interaction Models: A Scalable Approach to Human-AI Collaboration(明確引用「苦澀的教訓」)
- Gemma 4 Technical Report — §2(手工設計的推論路徑)及 §2.3(移除編碼器)(
empirical) - Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning — §3(移除
π_θ_old和群組基線;加入凍結注意力評論器與 skip-observation GAE):訓練迴圈的雙向變動案例(empirical) - Introducing System One Models & Jev — Diogo Almeida、TypeSafe AI、2026-09-15、
vendor-claim:FAQ 中「最苦澀教訓」的框架說法
Cited by 47
- Agentic Loops Overtake Bespoke Systems×6
The headline empirical finding of DeepMind's Ai Driven Formal Proof Search paper, and its clearest…
- Opinions on Using AI Tools & the Future of the Software Engineering Role×3
The bitter lesson recurs. The Bitter Lesson: scaled general methods beat hand-engineered structure.…
- Compute Allocator×3
What doesn't migrate inward — The Bitter Lesson dissolves model-facing structure; the allocation…
- HTML as the New Markdown×3
The Bitter Lesson dissolves model-facing structure; it does not dissolve the human-facing structure…
- Single General Agent vs. Multi-Agent Coding Architecture×3
The Bitter Lesson: scaled general methods beat hand-engineered structure over time; the structure…
- Agent Harness Engineering×2
Does a single general-purpose coding agent outperform a multi-agent architecture with specialized…
- Build for the Next Model×2
This is the product-side expression of The Bitter Lesson and Harness Shrinkage As Models Improve:…
- The Data Wall and the Validation Commons Are One Supply Constraint×2
The escape route is not a substitute for scaling. Both instructors state independently that larger…
- Effective Compute Scaling×2
The Bitter Lesson — "is scaling enough?" is the bitter lesson as a forecasting question; search…
- Encoder-Free Early Fusion×2
The Bitter Lesson — "co-train from scratch, drop the modular encoders" is a bitter-lesson move
- Does the Human-Facing Harness (HTML Artifacts) Hit Its Own Bloat Ceiling?×2
The model-facing harness can shrink toward zero as capability migrates inward (Harness Shrinkage As…
- Inference Efficiency as Capability×2
The report simultaneously removes hand-engineered structure (encoders, per Sutton's logic) and adds…
- Interaction Models×2
The central bet: interactivity should scale alongside intelligence. If interaction is part of the…
- MCP and Computer Use×2
This connects to The Bitter Lesson: as models improve, the boundary between "use an MCP" and "use…
- Rationale Bootstrapping (STaR)×2
And the scale term cuts against the flywheel's most interesting reading. Both instructors state…
- Recursive Self-Improvement×2
Perspiration is becoming automated. AI advances rarely come from "eureka" moments; paradigm shifts…
- Research Taste as the Human Bottleneck×2
The Bitter Lesson — "research progress is mostly tools and resources" is the bitter lesson aimed at…
- Software 3.0×2
Pushed to the limit: a "completely neural computer" — raw video/audio in, diffusion rendering a UI…
- Thinking Machines Lab×2
Position: interactivity should scale with intelligence → it must be in the model, citing The Bitter…
- Turn-Based Interface Bottleneck×2
The Bitter Lesson says these hand-crafted systems get outpaced by general capability growth → the…
- Universal AI (AIXI)×2
The Bitter Lesson — "intelligence as search through hypothesis/policy space" is the shared premise;…
- What Scaffolding Survives Model Improvement — and How Do You Know When a Line Turns Harmful?×2
The question's examples (org style, security rules, brand voice) all survive, and the sorting rule…
- AI Accelerating AI Development
The Bitter Lesson — "research progress is mostly a function of tools and resources" is the bitter…
- AI R&D Autonomy Evaluation (AECI)
The Bitter Lesson — the acceleration AECI tracks is what makes "scaled general methods improve…
- Andrej Karpathy
The Bitter Lesson — the Sutton principle his neural-net-as-host-process extrapolation rests on
- Asynchronous RL for LLMs
The Bitter Lesson — DIS removes hand-built machinery (π_θ_old, checkpoint history) — "simpler by…
- Authority and Audit Survive Abundance
The shared shape: both ask whether abundance retires a layer — model capability retiring authority…
- Claude Character as Product
The Bitter Lesson — character is a candidate counterexample: a deliberately hand-crafted asset that…
- CS329A: Self-Improving AI Agents (Stanford)
The scale term, stated twice and followed up by neither instructor. Absolute Zero reports larger…
- Deep Modules for Agents
Single Vs Multi Agent Coding Architecture — the Sandcastle Planner/Implementers/Reviewer/Merger…
- Evolutionary Proof Search
The Bitter Lesson — Elo/P-UCB/evolution is exactly the hand-engineered structure the bitter lesson…
- The Future of Agent Interfaces
Interaction Models is the strongest claim in the covered pages. It says interactivity should be…
- Gemma 4
The Bitter Lesson — the report both obeys it (drop the encoders) and defies it (hand-engineer the…
- Group Relative Policy Optimization (GRPO)
The Bitter Lesson — GRPO's critic-free design was itself a "remove hand-built structure" move…
- Harness Shrinkage as Models Improve
The Bitter Lesson — the underlying principle: hand-crafted scaffolding gets outpaced by scaled…
- Jeff Dean
The Bitter Lesson — his TPU design rule (specialize to the arithmetic, not to the architecture) is…
- Jev
The Bitter Lesson — the vendor's counter-slogan, "the bitterest lesson": optimizing for the right…
- Model Capability & Training
The Bitter Lesson — Sutton 2019: scaled general methods beat hand-engineered structure; recurring…
- Model Spec Midtraining (MSM)
Underlying principle: The Bitter Lesson — moving alignment from harness-prompt-injection of values…
- Prototype Over PRD
The "why not what" rule is load-bearing: it leaves the what to the model, so the prototype can…
- Repository Exploration Subagent
The Bitter Lesson — a live tension: training a specialized explorer adds hand-built structure,…
- Single-Rollout Optimization
The Bitter Lesson — SAO both removes structure (drops the group baseline, drops π_θ_old) and adds a…
- Task Time-Horizon Scaling
The Bitter Lesson — rising capability on general benchmarks is what makes hand-built scaffolding a…
- Time-Aligned Micro-Turns
The Bitter Lesson — "no turn boundaries → interaction modes become scalable model behavior" is a…
- Typed Decision Verifiers
The Bitter Lesson — TypeSafe's "bitterest lesson" (the right training objective beats data, compute…
- The Verifiability Thesis
The Bitter Lesson — RL-at-scale in verifiable environments is the general method outrunning…
- Why AI Lags at Design
The Bitter Lesson / Build For The Next Model — "these models will get good at design" is the…
Related articles
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
