資料來源#
- Research: Why You Shouldn’t Treat AI Agents Like Employees
- The Speed Trap: 8 takeaways from our latest AI engineering research
摘要#
Kropp、Bedard、Wiles、Hsu、Krayer 在 HBR 2026/03 的〈When using AI leads to brain fry〉一文中提出此詞,指過度使用 AI 或監督 AI,超出認知能力所造成的精神疲勞。經歷 brain fry 的工作者表示,他們犯錯的頻率明顯高於同儕:輕微錯誤頻率高出 11%,重大錯誤頻率高出 39%。2026 年 5 月 HBR 的後續論文將其視為一種認知機制,可能會在 AI 員工框架下加劇。
機制#
員工監督 AI 產出時:
- 將 AI 視為工具的框架 — 審查的認知負擔仍由人類承擔。大量使用 → brain fry → 錯誤增加 11–39%。
- 將 AI 視為員工的框架 — 人類可能覺得不必全心投入審查負擔(「ALEX-3 已經做過了」)。短期內或許能減輕 brain fry 症狀,卻是藉由審查不足來達成,形成另一種失敗模式。實驗中錯誤捕捉率下降 18%,與此相符。
因此,兩種框架都有各自的成本面:工具框架讓審查者承擔負擔;員工框架則以投入不足取代負擔。
對 Human-AI Accountability Redesign 的啟示#
Brain fry 說明了為什麼「只要擴大管理幅度」行不通。若不重新設計審查方式,只增加每位人類審查者面對的 AI 產出量:
- 超過某個門檻後,brain fry 就會發作 → 錯誤率攀升。
- 即使尚未達到門檻,審查品質的邊際表現也會下滑。
論文所暗示的重新設計方案:
- 縮小審查範圍(以抽樣稽核取代逐項審查)
- 將審查集中在高風險決策點(以決策權限把關,見 Claude Code Auto Mode)
- 將人類角色從逐項審查轉為系統層級監督(編排品質、效能監測)
- 調整績效管理制度,獎勵編排能力,而非逐項找出錯誤
與程式開發工作流程研究的關聯#
- Context Window Smart Zone — 模型端的類比認知限制。模型超過約 100K 個 token 後,敏銳度會下降;人類則在超出監督能力後失去敏銳度。兩者都有一個智慧區間,超過後效能下滑的速度會比能力所顯示的更快。
- Harness Shrinkage as Models Improve — 更好的模型減少每項任務所需的審查,部分緩解 brain fry;但能更快產出更多內容的代理程式又會帶來產量壓力。
- Agent Loop Pattern — 迴圈會大幅放大產出;brain fry 是它們碰上的人類端限制。
相關文章#
- The Tragedy of the Cognitive Commons — 監督成本的另一面:brain fry 衡量驗證所造成的疲勞;Validation Tether 衡量驗證能力的流失
- Outsource Your Thinking, Not Your Understanding — 過度委派導致理解變淺,是監督疲勞在認知負荷上的近親
- The Automation–Optimism Link — 相反的訊號:Anthropic 的 AEI 調查發現,重度委派者表示自己沒有學習落差,且認為技能價值更高。採用的工具不同(自我回報與測量錯誤),機制也不同(委派感受與監督疲勞);本頁的錯誤資料進一步凸顯了感受與測量之間的張力
- Experimental Learning Impact of Generative AI — 將相同的「客觀測量勝過自我回報」方法用於學習任務:Contractor 與 Reyes 隨機分配 AI 使用權,發現自動化模式使用者在移除 AI 後,學習成果完全消失。這是本頁認知成本故事中技能退化的另一面,以因果方式測量,而非透過調查
- Verification as the New Bottleneck — 審查與驗證負擔是監督疲勞累積之處
- Psychological Costs of AI Adoption — 本頁衡量監督工作量的背後機制。 Kropp 等人測量過度監督造成的錯誤成本;該案例研究則訪談從業者,探討監督為何不可或缺,並發現驅動因素是責任仍由人類承擔,而非產出量——「我得逐行檢查結果,因為我要負責」,而最直接呈現此機制的說法是:「如果產出的程式碼不用由我們負責,速度會快得多。」兩者放在一起,呈現出一種值得正視的不安:從業者為了維持掌控而採取的每一種保留自主性的做法(逐行驗證、審查計畫與差異、接受變更前先做基準測試)都會增加監督工作量,因此,成功讓人類有意義地參與其中的組織,等於承擔了本頁所測量的錯誤率
- Loop Engineering — Osmani 所說的「你實際能跑幾個[迴圈],取決於你的審查頻寬,而不是工具」,正是把這道上限放在迴圈層級描述:工作樹消除了機械性的衝突,但 brain fry 仍是人類端的限制
- 相關概念:AI Employee Framing
- 重新設計目標:Human-AI Accountability Redesign
- 認知類比:Context Window Smart Zone(模型端)
- 產出放大器:Agent Loop Pattern
- 緩解方式:Claude Code Auto Mode(決策權限)、系統層級編排
- 監督品質風險:Compute Allocator — 「compute allocator」角色假定人類能做出良好決策;brain fry 是配置者照單全收的失敗模式
- 獨立創辦人放大效應:Founder as Agent Orchestrator — 同時執行多個平行代理程式工作階段,會比按人數擴編的組織更快讓監督負擔超出 brain fry 門檻
- Acceleration Whiplash — Faros AI 以組織規模遙測同一種疲勞:每位開發者每天處理的 PR 情境增加 67.4%、工作重啟增加 13.8%,且 31.3% 的 PR 未經任何審查便合併——這是在 4,000 個團隊中測得的投入不足失敗模式。Faros 於 2026 年 9 月發布的後續報告指出,壓力的型態正在改變,而非消退(The Speed Trap: 8 takeaways from our latest AI engineering research,
vendor-claim):平行作業/執行過多執行緒的壓力正在減輕,但工作重啟大幅升至 +66.7%——放棄進行中的工作並從頭重新處理;Faros 將此歸因於代理程式缺乏脈絡,而非人類負荷。兩項讀數都是在採用率已高的樣本群組中計算的期間成長率,而且兩份報告都沒有直接測量認知變數 - Parallel Agent Orchestration — 監督疲勞對並行作業形成的上限:OpenAI 使用者中位於第 99 百分位者,每天約執行 71 個代理程式小時,並同時使用許多代理程式;但代理程式執行時間總和不等於人類注意力——每個代理程式的審查負荷何時飽和,正是這道門檻
- Unknowns as the Agentic Bottleneck — 防止不理解內容就核准的對策:Thariq Shihipar 的測驗關卡(「我只有在完美通過測驗後才會合併」)讓合併取決於審查者是否理解,而非是否簽名
- Review as the Control Point — 以從業者論述而非對照實驗為來源的同一種疲勞機制:審查負荷增加會降低審查深度與動機,逐漸滑向照單全收(「審查者也許能撐過一個衝刺週期,但最後會精疲力竭或開始照單全收」)——這是 CMU 理論的 P2/P3
- Output Length Calibration — 同一種負荷的產量面,也是可調整的槓桿:Opus 5 預設每則訊息的代理程式敘述與書面交付成果都較長,因此除非在來源端提示調低節奏,否則每個工作階段的監督成本就會升高
- Security Debt of Agent-Generated Code — 有具體產物可佐證的投入不足失敗模式:在代理程式撰寫的 PR 中,真正外洩的憑證有 67.6% 是由人類提交,且其中 81.1% 沒有任何審查者留言;作者將此解讀為,在代理程式看似負責正確性的工作流程中,開發者警覺性降低/認知卸載
- Risk-Tiered Auto-Approval — 將「審查集中在高風險決策點」這項緩解方式機制化於合併關卡:PostHog 的 StampHog 在一個月內減少約 1.6K 次打斷;這些打斷原本都要求工程師離開工作流程,核准一項他們「幾乎完全沒有脈絡」的變更,而拒絕則會轉交給具名專家,而非送回佇列。本頁補充的但書是:被自動化的原本就是資訊不足的核准,因此疲勞確實得到緩解,但監督品質是否改善則尚未測量
- Configurable Human Participation — 參與成本面:HAS-Bench 的互動成本指標(回合數/人類步驟/token)及其「更多管道 ≠ 更好」的結果(在 6 種模式中,最佳單一管道有 5 種勝過所有管道;A4 過度介入會破壞原本已解決的任務)量化了過度詢問與不合時宜的介入所帶來的真實成本——更多人類參與並非免費
- Outsource Your Thinking, Not Your Understanding — 將「大腦是肌肉」的萎縮觀點應用於理解,而非疲勞:這種認知成本無法從已交付程式碼的指標中看出
衍生文章#
- Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence — 指出 brain fry 是創辦人手冊中「精實的 10 人獨角獸」主張未處理的成本面;提出以受限並行作業、抽樣審查與聚焦高風險事項,作為獨立創辦人的緩解方式
- Does the Augmentation/Automation Split Govern Skill at Work? — 將 brain fry 從「監督的認知成本」提升為第三種使用模式,有別於隨機學習實驗中的自動化組:自動化模式的使用者仍會讓產出經手自己,因此學習成果雖空洞卻真實;但本頁指出的投入不足(見 Acceleration Whiplash 中 31.3% 未經審查的 PR,以及 Security Debt of Agent-Generated Code 中 81.1% 未獲留言的憑證外洩)會讓人類完全退出迴圈——產量造成的影響,而 35 分鐘的監考時段能將其限制在一個單位
資料來源#
- Research: Why You Shouldn’t Treat AI Agents Like Employees(2026 年 5 月,提及 brain fry)
- 原始論文:Kropp 等人,When using AI leads to brain fry,HBR 2026/03
Cited by 35
- The Automation–Optimism Link×5
Ai Brain Fry — the direct tension: measured oversight fatigue and error increases vs. self-reported…
- Human-in-the-Loop Boundaries×5
Redesign the loop when the human is nominally accountable but cognitively overloaded; that is the…
- Does the Human-Facing Harness (HTML Artifacts) Hit Its Own Bloat Ceiling?×4
The binding constraint is "human attention and judgement, not generation cost" (Compute Allocator).…
- Does the Augmentation/Automation Split Govern Skill at Work?×4
> Does the same use-mode split govern workplace skill accumulation (the open question Automation…
- Opinions on Using AI Tools & the Future of the Software Engineering Role×3
"Brain fry" is real and measurable. Ai Brain Fry: mental fatigue from oversight beyond cognitive…
- Experimental Learning Impact of Generative AI×3
Ai Brain Fry — both put an objective, measured number on AI's cognitive effect (there, oversight…
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?×3
But the failure mode is a threshold, not a destiny — the countermeasures are also in evidence. What…
- Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence×3
The error surface is real. Playbook flags it via Agentic Technical Debt, Zero Friction Scope Creep,…
- Parallel Agent Orchestration×3
Ai Brain Fry — the cognitive cost of overseeing many parallel streams; the oversight-fatigue limit…
- Reviewer Habituation on Agent Pull Requests×3
Reviewer habituation is the hypothesis that a human who repeatedly reviews AI-agent pull requests…
- Acceleration Whiplash×2
Ai Brain Fry — the cognitive-load channel: context-switching and under-review are the human-side…
- AI Employee Framing×2
Brain-fry-adjacent disengagement. When output is "from an employee," reviewers may feel less need…
- Founder as Agent Orchestrator×2
How does the orchestration role change the founder's decision burden? Fewer hands-on tasks but more…
- Harness Patterns Under Scale and Domain Shift: Context Routing, Other Domains, Large Action Spaces, and the Overseer×2
Psychological Costs Of Ai Adoption, Ai Brain Fry — the verification tax; Faros's September shift.
- Loop Engineering×2
A fourth thread runs through the skills primitive: without skills the loop re-derives your whole…
- Outsource Your Thinking, Not Your Understanding×2
The framing Thawar gives it is physiological rather than economic — "the brain is a muscle; if you…
- Risk-Tiered Auto-Approval×2
A refusal doesn't dump the PR back into a queue; it routes to a subject-matter expert, selected by…
- Security Debt of Agent-Generated Code×2
Ai Brain Fry — the security register of oversight fatigue: humans committed 67.6% of the genuine…
- Unknowns as the Agentic Bottleneck×2
That is a direct, testable answer to a problem stated three ways across the wiki and solved in none…
- Agent Loop Pattern
Ai Brain Fry — the human-side limit on output multipliers: more loop output → more review → more…
- Claude Code
> Reading this as evidence — interpretation, flagged. Taken together the caps, the workflow-size…
- Claude Code Auto Mode
Ai Brain Fry — concentrating human review on high-stakes decision points rather than every action…
- Compute Allocator
Does treating humans as "compute allocators" risk the oversight-fatigue / accountability failure…
- Configurable Human Participation
Ai Brain Fry — the interaction-cost metrics (turns / human steps / tokens) and the "more channels ≠…
- Context Window Smart Zone
Ai Brain Fry — human-side analog of the smart zone: oversight has its own degradation curve past…
- Harness Shrinkage as Models Improve
Ai Brain Fry — partially mitigated by harness shrinkage (less to oversee), reintroduced by output…
- Human-AI Accountability Redesign
HBR five-pillar prescription: span-of-control redesign, role redesign, performance management reset, decision-rights/es…
- AI Economics & Labor
Ai Brain Fry — Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor…
- Open Questions Backlog
Experimental Learning Impact Of Ai: Does the same use-mode split govern workplace skill…
- The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales
The load that arrives is the error-prone kind. Oversight fatigue raises minor errors +11% and major…
- Output Length Calibration
Ai Brain Fry — narration volume is oversight load: more per-message output across more parallel…
- Psychological Costs of AI Adoption
Ai Brain Fry — the measured consequence of the strain described here. Kropp et al. find that…
- Review as the Control Point
Ai Brain Fry — review load → fatigue → rubber-stamping (P2/P3) is the oversight-fatigue mechanism,…
- The Tragedy of the Cognitive Commons
Ai Brain Fry — the other cost of oversight: brain fry measures the fatigue of validating, this…
- Verification as the New Bottleneck
Ai Brain Fry — the risk if verification stays manual: oversight fatigue increases errors as volume…
Related articles
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Human-AI Accountability Redesign
HBR five-pillar prescription: span-of-control redesign, role redesign, performance management reset, decision-rights/es…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Outsource Your Thinking, Not Your Understanding
"You can outsource your thinking but not your understanding"; understanding as the non-delegable human bottleneck; know…
- Agentic Technical Debt
Debt that *compounds* (not just accumulates) because each agentic-coding session re-derives architectural decisions wit…
