資料來源#
摘要#
Hermes Agent 是來自 Nous Research 的開放原始碼 CLI 程式碼與研究代理程式,定位為與 Claude Code 平行的生態系,並更著重於多平台訊息部署。CLI 提供可使用工具(終端機、檔案編輯、網頁搜尋、程式碼執行)的互動式 REPL;Hermes Gateway 則是長時間運作的 daemon,透過 Telegram、Discord、Slack 和 WhatsApp 提供同一個代理程式,並支援個別使用者工作階段、允許清單 + DM 配對授權、排程 cron 工作,以及可設定的容器後端。Hermes 不綁定特定供應商(可搭配 OpenAI、Anthropic 和其他 LLM API),採用較接近 OpenAI AGENTS.md、而非 Anthropic CLAUDE.md 的脈絡檔案慣例,並另以 SOUL.md 定義個性。
細節#
架構(兩種介面)#
| 介面 | 用途 | 生命週期 |
|---|---|---|
| Hermes CLI | 互動式 REPL,角色類似 Claude Code | 每次呼叫各自運作,可恢復(hermes -c、hermes -r "title") |
| Hermes Gateway | 長時間運作的 daemon,透過訊息平台提供代理程式 | systemd(Linux 使用者/系統服務)或 launchd(macOS);重開機後仍會持續運作 |
Gateway 是架構上最值得關注的部分——理念上與 Symphony 的常駐 daemon 模型平行,但租用單位是使用者而非 issue。
脈絡檔案#
Hermes 使用三層設定檔,並明確區分各自職責:
| 檔案 | 範圍 | 用途 |
|---|---|---|
AGENTS.md | 專案(cwd) | 專案脈絡、慣例、技術堆疊——每個工作階段自動載入 |
~/.hermes/SOUL.md(或 $HERMES_HOME/SOUL.md) | 全域個性 | 所有工作階段共用的穩定預設語氣與風格 |
.cursorrules / .cursor/rules/*.mdc | 專案 | 從 cwd 載入相容設定;不必重複維護 |
子目錄中的 AGENTS.md 檔案會在工具呼叫期間(透過 subdirectory_hints.py)延遲探索,並注入工具結果——不會在一開始就載入。這是刻意採用的脈絡預算策略:只有最上層專案脈絡會放進 system prompt;只有在相關時才會支付載入巢狀脈絡的成本。
AGENTS.md(專案)與 SOUL.md(個性)的分工,比 Anthropic 的 CLAUDE.md 慣例更明確,值得比較——實務上 CLAUDE.md 檔案也隱含了相同的職責分離,只是沒有分開存放。(後續會在 Agent Context Files 概念頁面進一步探討。)
記憶系統#
Hermes 實作了嚴格設限的記憶系統:
MEMORY.md:上限約 2,200 個字元。USER.md:上限約 1,375 個字元。- 記憶填滿時,代理程式會整併條目(壓縮較舊的筆記)。
文件特別提醒一個容易踩到的問題:
「記憶是凍結的快照——在工作階段中做的變更,不會出現在 system prompt,直到下一個工作階段開始。代理程式會立即寫入磁碟,但 prompt 快取不會在工作階段中途失效。」
這最清楚地說明了為什麼任何使用快取的代理程式都會讓人覺得記憶變更「延遲」——Claude Code 和其他代理程式也有這項限制,但很少明確記載。
記憶與技能的區分:
- 記憶 = 事實(環境、偏好、專案位置)。
- 技能 = 程序(多步驟工作流程、可重用的做法)。
- 「記憶記錄內容,技能記錄做法。」
Token 經濟工具#
直接提供給使用者的控制項——從操作層面來看,這些正是 AgentOpt 所正式化的調節手段,只是以 CLI 命令呈現:
| 命令 | 效果 |
|---|---|
/compress | 摘要整理對話歷史;保留關鍵脈絡並減少 token |
/usage | 查看 token 使用狀態 |
/insights | 查看近 30 天的使用模式 |
/model | 在工作階段中途切換模型(複雜推理用前沿模型,制式工作用快速模型) |
delegate_task | 以隔離脈絡啟動平行子代理程式;只傳回摘要 |
快取紀律的說明十分精準:
「大多數 LLM 供應商都會快取 system prompt 前綴。若讓 system prompt 保持穩定(使用相同的脈絡檔案與記憶),工作階段中後續訊息就能命中快取,大幅降低成本。避免在工作階段中途變更模型或 system prompt。」
CLI 操作便利性#
| 輸入 | 行為 |
|---|---|
Alt+Enter / Ctrl+J | 輸入多行內容但不送出 |
| 貼上多行內容 | 自動偵測並緩衝為單一訊息 |
Ctrl+C(按一次) | 中斷回應,並以新訊息重新導向 |
Ctrl+C(2 秒內按兩次) | 強制離開 |
Ctrl+V | 從剪貼簿貼上圖片(視覺能力) |
/ + Tab | 自動補全斜線命令 |
/verbose | 循環切換工具輸出模式:off → new → all → verbose |
Hermes Gateway:多使用者 daemon 模式#
Gateway 透過訊息平台提供代理程式:
- 個別使用者工作階段——每位獲授權的使用者都有自己的對話脈絡。
- 主要頻道(
/sethome)——指定接收 cron 輸出與主動訊息的聊天室。 - 兩種授權模式:
- 靜態允許清單:在
.env中設定TELEGRAM_ALLOWED_USERS=123,456。新增使用者需要重新啟動。 - DM 配對:未授權的 DM 會收到一次性代碼;管理員執行
hermes pairing approve telegram XKGH5N7P。不需重新啟動。
- 配對安全性:代碼 1 小時後過期,採用密碼學安全的隨機值,並設有速率限制(每位使用者每 10 分鐘 1 次請求、每個平台最多 3 筆待處理);若驗證失敗 5 次,該平台會鎖定 1 小時;所有資料都以
chmod 0600儲存。
服務安裝:
hermes gateway install # Default: user-level systemd (Linux) / launchd (macOS)
sudo hermes gateway install --system # Linux: boot-time system service
sudo loginctl enable-linger $USER # Linux: keep running after SSH logoutCron 工作#
排程工作會送到主要頻道:
- 在聊天室中建立:「每個平日上午 9 點,檢查 GitHub repo 是否有……」
- 設定儲存在
~/.hermes/cron/jobs.json;輸出位於~/.hermes/cron/output/{job_id}/{timestamp}.md。 - 重要限制:cron 提示會在完全全新的工作階段、沒有任何記憶的情況下執行。每個提示都必須包含所有必要脈絡——檔案路徑、URL、伺服器位址、指示。
這與 Symphony 由 daemon 驅動的派送模型相似——兩者都是由排程非同步派發工作給代理程式,而非以互動方式驅動。
容器安全模型(值得特別指出)#
Hermes 支援多種終端機後端:
TERMINAL_BACKEND=docker
TERMINAL_DOCKER_IMAGE=hermes-sandbox:latest支援項目:Docker、Singularity、Modal、Daytona。
重要的安全性轉變:使用容器後端時,危險命令檢查會停用——理由是「容器就是安全邊界」。這代表:
- 鎖定設定的容器映像品質成了安全關鍵。
- 安全性從核准提示 UX 轉向映像紀律 UX。
- 信任基礎從「每個命令都經過審查」轉為「容器無法逃逸」。
值得留意的是,這與 Claude Code 逐命令使用的 auto mode 分類器,在安全態勢上有實質差異。
與 Claude Code 比較#
| 能力 | Claude Code | Hermes |
|---|---|---|
| 專案脈絡 | CLAUDE.md | AGENTS.md(專案)+ SOUL.md(個性,分開存放) |
| Cursor 相容性 | n/a | .cursorrules / .cursor/rules/*.mdc 自動載入 |
| 工作階段精簡 | /compact | /compress |
| 工作階段中途切換模型 | /model | /model |
| 平行子代理程式 | .claude/agents/ 中的子代理程式 | delegate_task 工具 |
| 權限控管 | auto mode 分類器 | 依模式逐項核准(once/session/always/deny);容器中停用 |
| 記憶模型 | n/a(依賴對話 + CLAUDE.md) | 有界的 MEMORY.md + USER.md,並自動整併 |
| 多使用者部署 | 每位使用者/工作階段各自執行 claude -p | Hermes Gateway 搭配允許清單或 DM 配對 |
| Cron/排程工作 | n/a(依賴外部 cron) | 內建 cron,並傳送至主要頻道 |
最重要的架構差異是Gateway:Claude Code 以工作階段為核心;在團隊規模部署時,Hermes 則以 daemon 為核心。
從外部觀察開發速度(2026 年 7 月)#
OpenHands 對 GitHub 進行的十二個月分析(Coding Agents and Technical Debt,case-study,由競爭者撰寫)納入 NousResearch/hermes-agent,並指出它在兩個方向上都屬於異類。截至 2026-07-08 的一年期間:**合併了 7,736 個 PR——在比較的四個 harness 中最多;約有 300 萬行程式碼變更,目前程式碼約 175 萬行,是比較中規模最大的程式碼庫。**成長曲線也最陡:「從幾乎空白的 repo,到開始認真投入後幾個月內,每月約合併 2,000 個 PR。」
它有 68% 的修正錯誤 PR 占比(5,288 個 PR),名義上遠高於其他三者的 16–40%;但作者指出,這很可能是 Hermes 標籤慣例造成的結果,而非品質訊號,並表示四個專案的分類都採用啟發式方法。引用 68% 時務必附上這項限制。數據所處的脈絡及其支持的自建與外購論點:Harness Build-vs-Buy。
操作備註#
- VPS 規格:每月 $5 就足以運行 Gateway 本身——成本主要來自 LLM API 呼叫。
- macOS launchd 的 PATH 陷阱:plist 會在安裝時擷取 shell PATH。安裝新工具(Node、ffmpeg)後,請重新執行
hermes gateway install以更新。 - 工作階段自動重設:訊息工作階段會在閒置後重設(預設 24 小時),或每天凌晨 4 點重設。
- 自行更新:在聊天室中執行
/update會取得最新版本並重新啟動。
技能安裝掃描(2026 年 8 月)#
Nous Research 在 Hermes Agent 的技能安裝流程中試行了 SkillSpector(NVIDIA/SkillSpector),作為可選的諮詢式掃描:檢查 PII、Unicode 偷渡、指令碼 lint、授權條款與安全性,並在安裝前顯示檔案行號與檢查結果。29 項測試通過;每次技能掃描約需 1.4–1.5 秒。這項掃描只提供建議,不會阻擋安裝——但這是本 wiki 首次記錄技能登錄處在安裝流程中執行靜態分析,也與 Skill Lift 流程的 Tier 1 搭配,構成技能驗證閘門中的安全性部分。
相關連結#
- Claude Code Best Practices — Hermes 是最接近的平行生態系;許多概念可直接對應(
/compress↔/compact、delegate_task↔ 子代理程式、AGENTS.md↔CLAUDE.md);差異凸顯雙方各自的設計選擇 - Symphony — 兩者都是常駐代理程式 daemon;Symphony 以 issue 為租用單位,Hermes 則以使用者為單位。兩者都偏好使用容器後端來確保安全。Cron + 主要頻道與 Symphony 的輪詢 + tracker 寫入方式相呼應
- Client-Side Agent Optimization — Hermes 的
/model、/compress、delegate_task和 prompt 快取紀律,是 AgentOpt 正式化的各項調節手段直接呈現給使用者的介面 - Agent Harness Engineering —
AGENTS.md遵循 OpenAI「目錄而非百科全書」的原則;有界記憶檔案是 harness 層級的限制,類似 JSON 功能清單 - Scale-Dependent Prompt Sensitivity —
/verbose模式與有界記憶都會以隱含方式限制輸出長度;簡潔限制的研究結果預測這會帶來效益 - Claude Code Auto Mode — Hermes 依模式核准(
once/session/always/deny)以及容器停用核准的模型,可與以分類器為基礎的 auto mode 相互比較 - Agentic Misalignment (AM) — Hermes daemon 模式(長脈絡、使用工具、以容器作為安全邊界、對個別行動的人類監督薄弱)完全落在 AM 的威脅範圍內;容器隔離能縮小爆炸半徑,卻無法處理模型本身的失準
- Ticket-Driven Agent Orchestration — Hermes 的 cron 工作 + 主要頻道派送,是較輕量的對應做法:自行派送排程交付項目,再回報至聊天室,而非建立 ticket
- Harness Build-vs-Buy — Hermes 的維護負擔在實測中居於高端:十二個月內合併了 7,736 個 PR,程式碼約 175 萬行;原始資料本身也打了折扣的 68% 修正錯誤標籤異常值
- OpenHands — 發布這項比較的開放原始碼同業 harness
尚待解答的問題#
- 容器後端停用危險命令檢查,是站得住腳的設計,但也帶來顯著的安全模型轉變。實際經驗如何?熱門映像(Daytona、
nikolaik/python-nodejs)的鎖定設定失敗是否曾造成事故? - 有界記憶檔案(
MEMORY.md約 2,200 個字元)長期使用效果如何?文件提到自動整併,卻沒有說明細節——整併演算法是什麼?資訊損失有多嚴重? - Hermes 的 DM 配對流程是乾淨俐落的安全機制。為什麼 Claude Code 或 Cursor 尚未在共用/團隊部署中採用這種模式?
- Hermes 明確區分
AGENTS.md(專案)與SOUL.md(個性),Claude Code 的CLAUDE.md則只隱含這項區分。明確分工是否能實質改善結果,還是只是缺乏實證支持的文件編寫選擇? - Cron 工作在全新且沒有記憶的工作階段中執行——團隊如何安排「代理程式所需的脈絡」,又不讓每個 cron 提示都過度膨脹?是否有標準做法?
資料來源#
- Tips & Best Practices — CLI 實用操作方式、脈絡檔案、記憶、效能/成本調節手段、安全模式
- Tutorial: Team Telegram Assistant — Gateway 部署、BotFather 設定、允許清單與 DM 配對比較、cron 排程、正式環境操作
- Coding Agents and Technical Debt — Rajiv Shah、OpenHands、2026-07-28(
case-study,由競爭者撰寫):NousResearch/hermes-agent的第三方 GitHub 活動資料
Cited by 27
- Agent Control Plane Patterns: Tickets, Loops, Specs, and Memory Files×4
Backlog-draining loops answer some of this by using markdown issue files. Cron loops answer less:…
- Where Does Agent Harness Work Remain Durable as Models Improve?×3
Across Agent Harness Engineering, Claude Code Best Practices, and Hermes Agent, the stable work is:
- Symphony×3
Hermes Agent — parallel "always-on agent daemon" architecture, with per-user instead of per-issue…
- Agent Context Files×2
Hermes Agent — the sharpest role split (AGENTS.md project vs. SOUL.md personality) plus lazy…
- Agent Harness Engineering×2
Hermes Agent — parallel daemon-first agent ecosystem; per-user instead of per-issue isolation, with…
- Agentic Misalignment (AM)×2
This describes Cowork, Claude Code in agent mode (especially --dangerously-skip-permissions),…
- Claude Code Auto Mode×2
Compared to OS-level sandboxing (mentioned in Claude Code Best Practices alongside auto mode),…
- Claude Code Best Practices×2
Claude Code is one of several converging coding-agent ecosystems. Capability parallels with Hermes…
- Client-Side Agent Optimization×2
The client-side levers AgentOpt formalizes (model assignment, budget, caching, batching) appear as…
- Harness Build-vs-Buy×2
Every organization that decides it needs "our own coding agent" is making a make-or-buy decision,…
- Orchestration Sets Token Economics×2
Hermes Agent is credited as genuinely model-agnostic with isolated sub-agents, but the design-time…
- Ticket-Driven Agent Orchestration×2
Hermes Agent — Hermes's cron jobs and home-channel delivery are a lighter analogue: scheduled…
- Agent Identity Management System (AIMS)
Hermes Agent — a shipped contrast: Hermes secures multi-user agent access with an ad-hoc DM-pairing…
- Agent-Native Infrastructure
Hermes Agent — a concrete agent-native daemon (AGENTS.md context, gateway connectors) bridging chat…
- Classifier Gates vs OS Sandboxing: The Defense-in-Depth Story for Auto Mode and Cowork
Sandbox-only is a legitimate design point when the workload is fully containable: Hermes Agent…
- Continuous Self-Modification Under Review
Terminal-Bench 2.1 · Grok 4.5 · 84.94% audited · Cursor: 79.3%; Hermes: 77.53%
- Harness Shrinkage as Models Improve
Every measurement above is taken on the system prompt. OpenHands' July 2026 GitHub analysis…
- LLM-as-Compiler Knowledge Base
The schema layer (Karpathy's term) is the same artifact category as SPEC.md/WORKFLOW.md —…
- MCP and Computer Use
Hermes Agent — third-party agent product that consumes MCP (mentioned in cross-tool capability…
- Entities — People, Orgs, Tools & Projects
Hermes Agent — Nous Research's CLI agent + Gateway daemon (Telegram/Discord/Slack/WhatsApp);…
- NVIDIA
NVIDIA SkillEvaluator turns a with/without-skill ablation into a publication gate for its own 300+…
- Open Questions Backlog
Hermes Agent ×5 (oldest 160d) — The container backend disabling dangerous-command checks is a…
- OpenClaw
Hermes Agent — sibling personal-agent CLI (Nous Research); the two share the…
- OpenHands
Hermes Agent — the other open agent in the same comparison (7,736 merged PRs, ~1.75M lines)
- Prompt-Cache Economics
"Keep the prefix stable and you get cheap cache hits" is true only above the threshold. Agent…
- Scale-Dependent Prompt Sensitivity
Hermes Agent — /verbose modes and bounded memory implicitly cap output length; the…
- Skill Lift
Nous Research's Hermes Agent tested it as an optional advisory scan at install time — SkillSpector…
Related articles
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Claude Code Best Practices
Anthropic's guide to effective Claude Code usage: context management, verification-driven development, explore→plan→cod…
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- MCP and Computer Use
Anthropic's two complementary connector mechanisms: MCP for structured programmatic access (Salesforce/Drive/Gmail/Slac…
