資料來源#
- An open-source spec for Codex orchestration: Symphony.
- Anthropic's Boris Cherny: Why Coding Is Solved, and What Comes Next
- Auto mode for Claude Code
- Best Practices for Claude Code
- Fable's judgement
- Full Walkthrough: Workflow for AI Coding — Matt Pocock
- How Anthropic's product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
- Introducing Claude Opus 4.7
- Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
- Tips & Best Practices
- Tutorial: Team Telegram Assistant
摘要#
Anthropic 官方的 Claude Code 有效使用指南,圍繞一項核心限制展開:脈絡視窗很快就會填滿,效能也會隨之下降。所有最佳實務都源自管理這項稀缺資源,包括以驗證為導向的開發、結構化脈絡(CLAUDE.md)、積極管理工作階段,以及透過平行工作階段水平擴展。
詳細說明#
脈絡視窗是首要限制#
脈絡視窗容納整段對話:訊息、檔案讀取內容、命令輸出。一次除錯工作階段就可能耗掉數萬個 token。脈絡逐漸填滿時,Claude 會「忘記」較早的指示,也更容易出錯。歸根究柢,每項最佳實務都是在管理這項資源。底層機制請參閱脈絡視窗智慧區(注意力的二次方縮放、約 100K token 的智慧區標記)。
模型層級的放大因素(隨 Claude Opus 4.7 推出,在目前的 Claude Opus 4.8 中仍適用):更新後的 tokenizer 會將相同輸入映射為多出 1.0–1.35 倍的 token,而 Opus 4.7「在較高努力等級下會思考得更多」——尤其是代理程式情境中的後續輪次。Claude Code 的預設努力等級已提高至 xhigh。這些因素會疊加:原本在 4.6 的 high 下放得進去的工作階段,到了 4.7 的 xhigh 可能就會明顯更吃緊。沿用對 4.6 的直覺之前,先用真實流量測量。可採取的反向調整:降低努力等級、設定任務預算(API)、明確提示要簡潔,或用簡潔風格的輸出上限(參閱依規模而變的提示敏感度)。
以驗證為導向的開發#
最能發揮效益的單一做法:讓 Claude 有方法驗證自己的工作成果。提供測試、螢幕截圖、預期輸出或 linter 命令。缺少驗證時,Claude 會產生看似合理、實際卻有問題的程式碼,而人就成了唯一的回饋迴路。
重要做法:
- 提供具體測試案例,列出輸入與預期輸出
- 修改 UI 時,貼上螢幕截圖,請 Claude 比較成果
- 提供錯誤訊息以處理根本原因,不要只說「建置失敗」
- 使用 Claude in Chrome 擴充功能自動測試 UI
探索 → 規劃 → 編碼工作流程#
將研究與實作分開。涉及多個檔案的修改或不熟悉的程式碼時,使用 Plan Mode。範圍明確且能用一句話描述差異時,就跳過規劃。
更積極的變體是設計概念盤問(Matt Pocock 的 grill-me skill):不再是「請代理程式提出計畫」,而是「讓代理程式持續訪談你,直到雙方在計畫形成之前取得共識」。另請參閱供代理程式使用的垂直切片追蹤彈,了解如何把產出的 PRD 切成代理程式可直接接手的 Kanban 工單;以及代理程式的深層模組,了解如何讓程式碼庫結構對代理程式更友善。
環境設定#
- CLAUDE.md:每個工作階段都會載入的持續性指示。只納入 Claude 無法從程式碼推知的內容——bash 命令、非預設的程式碼風格、工作流程規則、架構決策、注意事項。大刀闊斧地精簡:若沒有指示 Claude 也能把事情做好,就刪掉。逐行精簡是必要條件,卻還不夠;如今這項缺失的限制已有量化結果:Instruction Compounding 根據 Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models(
empirical,五個模型)記錄了容量下限:所有規則都遵守的比例在同時可驗證指示達 40 條時急遽下降,到 80 條時降為零;markdown、純文字、散文、表格呈現方式皆相同。實驗中的每條規則都彼此不同且可單獨滿足,因此逐行消融也無法察覺這個問題。CLAUDE.md 有兩項推論:限制單位是指示的數量,不是 token;格式也不是著力點——放置位置的影響更大(將相同區塊在 system prompt 與 user turn 之間移動,遵循率最多改變 8.7 個百分點),且不同模型的變化方向不一,必須實測而不能想當然。把它當成程式碼:出了問題就檢視,並透過觀察行為變化來測試。使用@path匯入以模組化。對創辦人/個人建置者而言,每次工作階段一開始就把 CLAUDE.md 當作架構脈絡,結束時再更新它;這種更嚴謹的做法,是防範代理程式技術債的主要方式——這種債務會複利累積(不只是增加),因為脈絡若未保存,每次工作階段都得重新推導基礎決策。 - Skills(
.claude/skills/):依領域而定的知識與可重複使用的工作流程,需要時載入,而非每個工作階段都載入。以/skill-name呼叫。 - Subagents(
.claude/agents/):在隔離脈絡中執行、工具範圍受限的專門助理。適合用來處理需要讀取大量檔案、又不想讓主要脈絡變得雜亂的任務。 - Hooks:在 Claude 工作流程特定時點執行的確定性腳本。與僅供參考的 CLAUDE.md 不同,hooks 能保證執行。
- MCP servers:透過
claude mcp add連接外部工具(Notion、Figma、資料庫)。 - Plugins:從市集安裝、打包在一起的 skills、hooks、subagents 與 MCP。
- Permissions:auto mode(以分類器核准,介於預設提示與
--dangerously-skip-permissions之間的折衷方案)、允許清單,或作業系統層級的沙箱。
工作階段管理#
- 不相關任務之間使用
/clear,避免脈絡遭到污染 - 使用
/compact <instructions>,依指定內容進行摘要 - 使用
/rewind或Esc+Esc,還原至任一檢查點的對話、程式碼或兩者 - 透過 subagents 進行調查,在獨立脈絡中探索,並回報摘要
- 使用
/btw提出不會進入對話歷史的附帶問題 - 同一問題連續修正兩次都失敗後,使用
/clear,並將學到的內容納入提示後重新撰寫
擴展模式#
- 非互動模式:在 CI、腳本、pre-commit hooks 中使用
claude -p "prompt"。支援 JSON 與串流輸出。 - 平行工作階段:桌面應用程式(隔離的 worktrees)、網頁(隔離的 VMs),或代理程式團隊(協調並共用任務的工作階段)。
- 撰寫者/審查者模式:一個工作階段負責實作,另一個以全新脈絡進行審查(不會偏袒自己的程式碼)。
- 分散式擴展:針對大型遷移逐檔執行
claude -p迴圈。使用--allowedTools限定權限。 - 無人值守執行時使用 Auto mode:分類器會擋下高風險操作,允許例行工作。非互動模式若反覆遭到阻擋,就會中止。
- 迴圈與例行工作:
/loop(在 CLI 中依 cron 排程重複工作)與 routines(伺服器端版本)。以 AFK 方式清空 Kanban 待辦清單;這是將規劃成本分攤至多次執行的主要機制。另請參閱代理程式迴圈模式。
平行生態系與跨工具概念對照#
Claude Code 是多個逐漸匯流的程式設計代理程式生態系之一。與 Hermes Agent(Nous Research)和 Codex(OpenAI)的功能對應如下:
| 功能 | Claude Code | Hermes | Codex |
|---|---|---|---|
| 專案脈絡檔案 | CLAUDE.md | AGENTS.md(專案)+ SOUL.md(個性,分開存放) | AGENTS.md |
| 工作階段壓縮 | /compact <instructions> | /compress | (透過 Codex App Server thread compaction) |
| 工作階段中途切換模型 | /model | /model | 工作階段層級設定 |
| 平行 subagents | .claude/agents/ 中的 subagents | delegate_task | 透過 Symphony orchestrator 啟動 |
| 非互動/程式化使用 | claude -p、Claude Agent SDK | 腳本中的 hermes CLI | Codex App Server(JSON-RPC stdio) |
| 多使用者團隊部署 | 每個工作階段各自執行 claude -p | Hermes Gateway(Telegram/Discord/Slack/WhatsApp),搭配允許清單或私訊配對 | Symphony(由問題追蹤器驅動的常駐服務) |
| 權限控管 | auto mode 分類器 | 依模式核准(once/session/always/deny);使用容器後端時略過 | 由各 Symphony 規格實作決定 |
| 記憶體模型 | 對話 + CLAUDE.md | 有界的 MEMORY.md(約 2,200 字元)+ USER.md(約 1,375 字元) | 由檔案系統驅動 |
這三者共通的結構性洞見是:代理程式行為透過以儲存庫版本控管的 markdown 檔案設定(CLAUDE.md / AGENTS.md / SOUL.md / WORKFLOW.md)。各家廠商都採用這種模式,足以顯示它正在成為新興標準。(計畫另寫一篇專門的Agent Context Files概念頁,將此模式正式整理出來。)
架構上最大的差異是:Claude Code 以工作階段為先,另提供選用的非互動模式;Hermes Gateway 和 Symphony 在團隊規模部署時則以常駐服務為先。工作階段與常駐服務的分野,是 2026 年部署架構的主導選擇。
常見失敗模式#
| 模式 | 修正方式 |
|---|---|
| 百寶袋式工作階段(混雜不相關任務) | 任務之間使用 /clear |
| 反覆修正(超過兩次都失敗) | 使用 /clear,把學到的教訓寫進提示後重寫 |
| CLAUDE.md 過度規定 | 精簡;若要確保確定性,改用 hooks |
| 信任與驗證之間的落差 | 一律提供驗證標準 |
| 無止境探索 | 縮小範圍或使用 subagents |
相關連結#
- Agent Harness Engineering — Claude Code 的 CLAUDE.md、skills 與 hooks,是 OpenAI 和 Anthropic 研究團隊所述 harness 工程模式的實際應用
- LLM-as-Compiler Knowledge Base — CLAUDE.md 檔案在此知識庫的 LLM-as-compiler 架構中擔任綱要層
- LLM-Driven Vulnerability Research — Claude Code 是 Anthropic 漏洞研究支架的執行環境;所有 Mythos Preview 研究發現都使用了 Claude Code 的代理程式能力
- Client-Side Agent Optimization — 直接挑戰「使用最強模型」這項預設:在 HotpotQA 上,Claude Opus 4.6 搭配較便宜規劃器的組合,比全程使用 Opus 高出超過 40 個百分點。AgentOpt 的 httpx 攔截與
claude -p非互動模式相容 - 依規模而變的提示敏感度 — 補充脈絡視窗管理:簡潔限制既能提高容易過度思考問題的準確率,也能保留脈絡預算。大型模型的冗長輸出可能掩蓋推理錯誤,因此以驗證為導向的開發尤其重要
- Claude Code Auto Mode — 環境設定與擴展模式中提及的「auto mode」權限選項完整說明
- Claude Opus 4.7 — 引入字面遵循指示能力與 tokenizer 膨脹,改變了 CLAUDE.md 和工作階段管理的撰寫方式
- Claude Opus 4.8 — 目前多數 Claude Code 工作採用的模型(自 2026-05-28 起普遍開放);它是 4.7 的直接升級版,因此上方的脈絡預算指南仍完全適用。環境設定一節有一項注意事項:4.8 對提示注入的抵抗力比 4.7 弱,因此縮限工具權限更有價值
- Hermes Agent — 來自 Nous Research 的平行生態系;許多 Claude Code 模式都可直接對應(
/compress↔/compact、delegate_task↔ subagents、AGENTS.md↔CLAUDE.md);差異(Gateway 常駐服務、有界記憶體檔案、拆分的SOUL.md)則突顯各自的設計選擇 - Codex App Server Protocol — OpenAI 端對應於
claude -p+ Claude Agent SDK 的方案;兩者都能讓外部協調器驅動工作階段,但 App Server 對穩定的 JSON-RPC stdio 通訊協定有更明確的規範 - Symphony — 以常駐服務為先的部署典型;Claude Code 的對應做法是把
claude -p和 subagents 接上問題追蹤器,就像 Symphony 把 Codex 接上 Linear 一樣 - Ticket-Driven Agent Orchestration — 非互動模式穩定後,自然形成的協調模式;將單一工作階段的最佳實務延伸至團隊規模部署
- 脈絡視窗智慧區 — 驅動本文所有脈絡管理做法的底層限制
- 設計概念盤問 — 以探索→規劃→編碼為基礎、更積極且優先對齊的變體
- 垂直切片追蹤彈 — 將 Kanban 待辦清單拆分為迴圈基本操作可清空的任務
- 代理程式的深層模組 — 讓 Claude Code 的審查與驗證模式可靠運作的程式碼庫結構;指示傳遞的推送與提取方式
- 代理程式迴圈模式 —
/loop與 routines 是取代逐步提示的新一代基本操作 - 模型進步時的 harness 縮減 — 解釋最佳實務提示與 CLAUDE.md 區段為何隨每次模型發布而縮減;Cat Wu 每次推出新版本都會精簡內容
- Claude Code — 實體層級頁面
- AI Native Product Cadence — 這些最佳實務產物,是團隊依照該內部節奏運作的公開成果
- Engineer PM Convergence — 本指南隱含面向、兼具產品品味的工程師角色
- 代理程式技術債 — CLAUDE.md 主要防範的失敗模式;創辦人手冊特別提及此問題
- AI 原生新創生命週期 — 創辦人階段的觀點,將 CLAUDE.md 從「最佳實務」提升為「MVP 生存準則」
- MCP 與電腦操作 — 支援「以自訂工具擴充 Claude Code」擴展模式的連接器基礎;MCP 和電腦操作能讓外部系統成為代理程式可採取行動的介面之一
- Evals 作為產品規格 — 「以驗證為導向的開發」的嚴格形式:十個優質 evals 能在功能層級定義完成條件,補足本文提出的工作流程層級驗證
衍生內容#
- 何時在工作中使用 Claude Opus 4.6 — 將脈絡視窗視為首要限制的觀點,啟發 Claude Code 的推論:Opus 的冗長輸出會更快耗盡預算
- Opus 4.6 → 4.7 變化與多代理程式編碼考量 — 將 subagents、撰寫者/審查者模式和擴展模式指南應用於 Opus 4.7 多代理程式團隊
- 學習與 AI 協同工作:軟體工程師實務指南 — 將最佳實務提煉成供個別工程師培養技能的實務指南(六個技能群組、日常做法、反模式、90 天計畫)
- App Server 與 MCP,以及 Claude 端的對應方案:驅動代理程式的三種邊界 — 說明
claude -p+ Agent SDK 如何成為 Codex App Server 在 Claude 端的對應組合,並提供「驅動 CLI 或建置於 SDK 之上」的決策規則 - 撰寫者/審查者模式與代理程式對代理程式審查 — 將撰寫者/審查者擴展模式拆解為四個設計變數,並修正其原本的理由:偏袒自己程式碼是模型家族的特性,而非工作階段的特性,因此換用新脈絡無法消除偏誤(8.9 個百分點的高嚴重度召回率來自跨家族路由)
尚待解答的問題#
- 指示數量上限是否適用於條件式政策,也就是任何一輪中只有少數規則相關的情況?Eliav 實驗中的每條規則都會同時套用於單次生成;CLAUDE.md 多半是情境式內容(「編輯遷移時,……」),因此 N=80 是最嚴苛載入方式的下限,不能說明一份有 200 條規則、每輪只有五條生效的檔案。可直接證偽:固定適用的子集,逐步增加不適用的其餘規則。
- 撰寫者/審查者模式與代理程式對代理程式審查(例如 OpenAI 的 Codex 工作流程)相比如何?撰寫者/審查者模式與代理程式對代理程式審查於 2026-09-02 部分解答:兩者是只在審查場域不同的同一架構,因此比較可拆成四個變數,而綜合分析已釐清其中一項。審查者血統是有數據支持的軸線——Greptile 配對的 500 個 PR 資料集顯示,在高嚴重度錯誤上,跨模型召回率為 61.0%,同模型為 52.1%,差異明確(資料集主效應 2.6 個百分點、審查者主效應 0.6 個百分點);來源自身的組合機制只能重現約 7%。這推翻本頁「擴展模式」項目中的括號說明:偏袒自己程式碼是模型家族的特性,而非工作階段的特性,因此「全新脈絡」不能帶來「不偏袒自己的程式碼」——審查者應改用不同家族的模型。資訊邊界軸線的設計論據更充分(Bun 的僅差異審查者、OpenCodeReview 的僅做反證的反思器、Cursor 未公開的 lens sweep),但沒有程式碼審查測量;關卡位置只有一項阻擋率(Ouroboros,63.5%),沒有結果測量;結果則分成兩種相差一個數量級的構念(AACR-Bench 的參考匹配精確率為 7.23–37.80%,對比 54,713 則真實代理程式審查評論中由開發者解決的比例 71.4%)。**仍待外部證據解答:**沒有研究讓兩種模式正面對比;最接近的研究是比較三種產品與一個落後 33 個版本的基準;審查場域本身(工作階段中的第二意見,或發布 PR 後的評論串)也沒有任何方向的測量。衍生頁面已規劃好定案實驗:在 AACR-Bench 上交叉測試資訊邊界 × 血統 × 關卡位置,預先登錄並固定評論量指示,並在相同評論上報告兩種結果構念。
- subagent 的額外負擔何時會超過脈絡隔離帶來的效益?Codex 從 0 到 1,000 萬名使用者:建置 ChatGPT Work — Akshay Nathan,OpenAI於 2026-08-03 部分解答(
practitioner-opinion,沒有測量)——依據的是任務形態,而非交叉點。Akshay Nathan(OpenAI)表示,多代理程式模式「最適合用在極度複雜的任務,例如開放式探索,或是可高度平行化的任務……但多數任務不屬於這兩類」,所以預設應使用單一代理程式。請注意,他實際點名的額外成本既不是脈絡也不是 token:速率限制消耗(Ultra「可能會用掉更多額度」,因此 OpenAI 在推出後將它移到進階設定中)和人類可理解性(sub-agent 逐字稿預設隱藏,避免讓使用者不堪負荷——參閱共用 harness、差異化介面)。同一集節目裡有一種方向相反的實務做法:Vibhu 說他會要求每個長時間執行的任務「盡可能使用 sub-agents」,以縮短實際時間並降低成本,透過分散工作交給較便宜的模型——Cost-per-Task Over Cost-per-Token指出這種做法不適合當作預設。第二個案例(2026-08-04):Willison採用了相同做法,甚至把模型等級選擇本身也交由代理程式決定:「所有編碼任務都請自行判斷適合的較低能力模型,並在 subagent 中執行」,但只表示他的 Fable 額度消耗得比較慢。如今有兩位實務工作者預設採用分散式擴展,但都沒有測量成效。目前仍待解答的是量化後的交叉點,語料中的來源沒有一份能提供答案。
已解答的問題#
- CLAUDE.md 多長最理想,才能避免指示開始遺失?是否存在可測量的門檻?Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models於 2026-08-04 作出解答(Eliav,arXiv 2607.19257,
empirical)——完整分析見 Instruction Compounding。確實存在門檻,而且它取決於指示數量,而非長度:在五個模型(包括 Claude Sonnet 5 和 Haiku 4.5)中,提示內每一條指示都獲遵守的比率,在同時可驗證規則約達 N≈40 時急遽下降,至 N≈80 時降為零;此結果在 N=160 仍持續,且 markdown、純文字、散文、表格呈現方式,以及 system prompt 與 user turn 的放置位置,都不影響結果。論文本身提出的建議,就是可直接採用的答案:40 條同時適用的指示是重新設計的時機,不是調校的時機——超過此數量後,只有拆分至不同輪次、工具或驗證步驟才有用,重新排版則沒用。這也釐清 2026-08-03 更新標籤後留下的疑問(逐行消融都通過後,整體規模是否還有獨立影響?):是——所有受測規則都互不相同、沒有重複,且各自都能滿足,因此逐行消融非劣性檢定會讓它們全數過關,卻仍無法發現遵循率崩落。答案還有兩項適用範圍限制:「完美回應」是嚴格的合取,因此部分下限來自將 N 項檢查以 AND 串接的算術結果,不代表模型完全漏掉整段內容;而且所有受測規則都是套用於單次生成的硬性輸出限制。真實 CLAUDE.md 具備的條件式政策情境,現在成了上方「尚待解答的問題」中的後續問題。
資料來源#
- Best Practices for Claude Code
- Auto mode for Claude Code — 權限模式的擴展
- Introducing Claude Opus 4.7 — tokenizer/預設 xhigh 對脈絡預算的影響
- Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models — Netanel Eliav,arXiv 2607.19257,2026-07-21(
empirical,唯一作者、單一實驗室、未經同儕審查):解答 CLAUDE.md 長度問題的指示數量上限,以及放置位置的影響。解析提醒:原始 markdown 中,論文表 1 的模型清單儲存格遭到合併,本文未引用該清單;請參閱 Instruction Compounding 中的來源說明 - Fable's judgement — Simon Willison,2026-07-03(
practitioner-opinion):委派編碼工作並讓 subagent 自行選擇較便宜模型的做法;僅用於討論上方的 subagent 額外負擔問題
Cited by 46
- Learning to Co-Work with AI: A Software Engineer's Field Guide×5
Build a CLAUDE.md / AGENTS.md for every project you own. Treat it like code: review when things go…
- Claude Code Auto Mode×4
Compared to OS-level sandboxing (mentioned in Claude Code Best Practices alongside auto mode),…
- Where Does Agent Harness Work Remain Durable as Models Improve?×3
Across Agent Harness Engineering, Claude Code Best Practices, and Hermes Agent, the stable work is:
- Open Questions Backlog×3
Claude Code Best Practices (62d) — Does the instruction-count ceiling hold for conditional policy,…
- Opus 4.6 → 4.7 Changes and Multi-Agent Coding Considerations×3
Claude Code Best Practices#Context Window as Primary Constraint applies to each agent in isolation.…
- Writer/Reviewer vs Agent-to-Agent Review×3
Codex, Claude Code, Claude Opus 4 7, Claude Code Best Practices — the two products' review surfaces…
- AI-Native Startup Lifecycle×2
vs. Claude Code Best Practices: the playbook recommends starting each Claude Code session with the…
- Opinions on Using AI Tools & the Future of the Software Engineering Role×2
This matches the official tooling guidance: Claude Code Best Practices and Agent Harness…
- App Server vs MCP, and the Claude-Side Equivalent: Three Boundaries for Driving Agents×2
The vault documents no Claude-side equivalent of the App Server protocol. The comparison the corpus…
- Claude Opus 4.7×2
Anthropic claims the net is favorable on their internal coding eval across effort levels, but…
- Code as Source of Truth×2
Checking the spec into the codebase isn't just freshness hygiene — it's what makes mechanical…
- Instruction Compounding×2
So the pruning obligation this page establishes is necessary but not sufficient: a context file…
- LLM-as-Compiler Knowledge Base×2
The schema layer (Karpathy's term) is the same artifact category as SPEC.md/WORKFLOW.md —…
- MCP and Computer Use×2
Hermes Agent — third-party agent product that consumes MCP (mentioned in cross-tool capability…
- Symphony×2
Symphony vs. Claude Code agents: parallel ecosystems. Symphony is daemon-first (always-on,…
- Ticket-Driven Agent Orchestration×2
Claude Code Best Practices — Claude Code's claude -p non-interactive mode is the building block for…
- When to Use Claude Opus 4.6 for Work×2
Claude Code Best Practices — context-window constraint, verification discipline, session management
- Agent Context Files
Claude Code Best Practices — the CLAUDE.md convention and the prune-ruthlessly discipline; the…
- Agent Control Plane Patterns: Tickets, Loops, Specs, and Memory Files
Layered agent control-plane synthesis: tickets as durable work graph, loops as execution primitive, specs/context files…
- Agent Harness Engineering
Claude Code Best Practices — practical application of many harness engineering principles in Claude…
- Agent Loop Pattern
Claude Code Best Practices — the best-practices guide treats /loop as a core workflow primitive
- Agentic Technical Debt
Claude Code Best Practices — official Anthropic guidance on CLAUDE.md; the playbook frames the same…
- AI Native Product Cadence
Claude Code Best Practices — the public artifacts of a team operating at this cadence
- Blast Radius (Agentic)
Claude Code Best Practices — sandboxed execution + write-access restrictions as a reference…
- Classifier Gates vs OS Sandboxing: The Defense-in-Depth Story for Auto Mode and Cowork
Unattended operation. AFK loops and fan-out are auto mode's raison d'être, and the human who would…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Claude Sonnet 5
Sonnet 5 uses an updated tokenizer — the same kind of change Opus 4.7 introduced — so the same…
- Client-Side Agent Optimization
Claude Code Best Practices — directly challenges the implicit "use the strongest model" default.…
- Codex App Server Protocol
Claude Code Best Practices — Claude's claude -p non-interactive mode plus the Claude Agent SDK are…
- Cost-per-Task Over Cost-per-Token
Claude Code Best Practices — where this guidance is applied per session (effort defaults, context…
- Deep Modules for Agents
Claude Code Best Practices — module map in CLAUDE.md sits in the same family
- Design Concept Grilling
Claude Code Best Practices — the explore→plan→code workflow has the same shape; grill-me is the…
- Engineer PM Convergence
Claude Code Best Practices — engineer-with-taste is the user persona Claude Code targets
- Evals as Product Spec
Claude Code Best Practices — verification-driven development; evals as the strict version
- Hermes Agent
Claude Code Best Practices — Hermes is the closest parallel ecosystem; many concepts map directly…
- Least Agency
Claude Code Best Practices — Claude Code's deny-by-default permissions and write-access…
- LLM-Driven Vulnerability Research
Claude Code Best Practices — Claude Code is the runtime used for all vulnerability research; the…
- Agent Systems & Harness Engineering
Claude Code Best Practices (hub) — Anthropic's guide to effective Claude Code usage: context…
- Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence
Bounded parallelism (Cat Wu's "simple setups work better"; Claude Code Best Practices explicitly…
- Scale-Dependent Prompt Sensitivity
Claude Code Best Practices — the context-window-as-primary-constraint framing pairs naturally with…
- Shared Harness, Differentiated Surfaces
Claude Code Best Practices — the configure-don't-abstract pole on sub-agents (.claude/agents/…
- The Verifiability Thesis
RL training rewards verified outcomes, so the gradient flows hardest toward domains where…
- Vertical Slice Tracer Bullets
Claude Code Best Practices — the task-decomposition pattern that fills the Kanban backlog the loop…
- Vibe Coding vs. Agentic Engineering
Claude Code Best Practices — concrete agentic-engineering practice (explore→plan→code,…
- Xiaohongshu
The company also names two internal agent surfaces in the paper: context-gc (interactive chat…
- Zero Trust for AI Agents
Claude Code Best Practices — Claude Code's deny-by-default permissions, sandboxing, managed…
Related articles
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Client-Side Agent Optimization
AgentOpt's framing of developer-controlled agent optimization (model-per-role, budget, routing) as distinct from server…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
