資料來源#
- Model Spec Midtraining: Improving How Alignment Training Generalizes
- Verbalizable Representations Form a Global Workspace in Language Models
摘要#
一種對齊微調方法,使用(prompt, chain-of-thought, response)tuple 訓練模型——CoT 會根據規格或一組政策,推理該如何回應。由 Guan、Joglekar、Wallace、Jain、Bhalerao 等人(OpenAI,2025)提出,論文為 「Deliberative alignment: Reasoning enables safer language models」(arXiv 2412.16339)。此方法將規格內容蒸餾為監督式推理訊號。在 MSM 論文中,它被用作最強的非 MSM 基線。
機制#
對每個訓練提示:
- 將規格放入上下文,讓模型產生一段較長的 CoT,推理如何將規格套用於這個提示。
- 產生符合規格的回應。
- 以(prompt, CoT, response)進行 SFT——模型學會在部署時不依賴上下文中的規格,於內部完成審慎推理。
CoT 本身經常明確引用政策(「根據 SP2,我不能……」),藉此將以政策為依據的推理訓練進思維鏈中。
與 AFT 和 MSM 的比較#
在 MSM 論文中,此方法被視為 AFT(含 CoT),並與 AFT(不含 CoT) 和 MSM + AFT 變體比較。
| 方法 | 規格內容所在位置 | CoT 監督 |
|---|---|---|
| AFT(不含 CoT) | 隱含於回應示範中 | 無 |
| AFT(含 CoT,深思熟慮式對齊) | 蒸餾進 CoT | 有 |
| MSM + AFT(不含 CoT) | 以預訓練風格對規格文件進行中期訓練 | 無 |
| MSM + AFT(含 CoT) | 兩者皆有 | 有 |
實證比較(AM 評估)#
Qwen2.5-32B 基線 AM 比率為 68%:
- AFT(含 CoT)——也就是深思熟慮式對齊——降至 48%
- MSM + AFT(不含 CoT)降至 5%
Qwen3-32B 基線為 54%:
- AFT(含 CoT)降至 14%
- MSM + AFT(不含 CoT)降至 7%
在兩個模型上,MSM + AFT(不含 CoT)都勝過 AFT(含 CoT)。這項主要結果支持盡可能避免直接訓練 CoT。
CoT 可監控性的取捨#
Korbak et al. 2025 主張,過度以思維鏈訓練會損害它作為監控器的功能——一旦為了安全評分而最佳化 CoT,它就不再忠實揭露模型實際的推理。深思熟慮式對齊直接對 CoT 施加壓力。
MSM 提供另一條路徑:模型透過對談論規格的文件進行中期訓練,學會以規格為依據的推理,而非透過監督式 CoT。CoT 因此較少經過最佳化,作為監控器時也可能更忠實。MSM 是否確實保留 CoT 的可監控性,仍列為待解問題。
與 MSM 疊加#
MSM 論文第 5.1 節發現,價值增補式 MSM(說明規則為何存在的規格)與規則增補式 AFT-with-CoT(明確引用政策的深思熟慮式對齊)搭配效果良好。這表示以規則為基礎、類似深思熟慮式對齊的訓練,與價值說明式 MSM 彼此互補,而非重複。
高算力下的收斂#
在 Qwen3-32B 上使用 80k 個 AFT 樣本時,AFT(含 CoT)的表現收斂至 MSM+AFT(兩者的錯位程度都接近零,評估已飽和)。MSM 在低/中等 AFT 算力下優勢最大——它讓 AFT 的 token 效率提升 10–60 倍。
相關連結#
-
對比方法:Counterfactual Reflection Training——訓練模型產生一段評估時從不要求的反思式延續,因此在目標上下文中既不干預回應,也不干預推理軌跡;因此可避免此技術對 CoT 施加的直接壓力,而這種壓力會帶來 Chain-of-Thought Monitorability 風險
-
來源論文:Guan et al. 2025 (OpenAI)
-
研究/使用者:Anthropic(Anthropic 研究的對齊堆疊元件)
-
將規格作為上下文內 CoT 輸入:Claude's Constitution / Model Spec(將規格作為 CoT 生成時的上下文)
-
對比方法:Model Spec Science(教授規格內容的另一種方式)
資料來源#
- Model Spec Midtraining: Improving How Alignment Training Generalizes(將其作為基線)
- Guan et al. 2025 — Deliberative alignment: Reasoning enables safer language models(arXiv 2412.16339)
- Verbalizable Representations Form a Global Workspace in Language Models — 反事實反思訓練是對照案例:它監督一段從未被要求的反思式延續,在目標上下文中既不干預回應,也不干預推理軌跡
Cited by 11
- Alignment Fine-Tuning (AFT)×3
AFT (with CoT) — Deliberative Alignment-style. Each sample is (prompt, CoT, response) where CoT…
- Chain-of-Thought Monitorability×3
Outperforms AFT (with CoT) — i.e. Deliberative Alignment — at 14%
- Model Spec Midtraining (MSM)×3
New training phase between pretrain and AFT: train base model on synthetic docs discussing the Model Spec; controls AFT…
- OpenAI×3
On alignment training, deliberative alignment (OpenAI) is the direct-CoT-training baseline that…
- Agentic Misalignment (AM)×2
MSM + AFT with a Philosophy Spec (impermanence, self-preservation, goal-guarding, epistemic…
- Counterfactual Reflection Training×2
Deliberative Alignment — the closest rival technique; the contrast is where CRT's novelty lives
- Auditing the Misalignment-Measurement Instruments×2
Concept pages drawn on: Agentic Misalignment, Unsanctioned Action In Evaluations, Documented Agent…
- Anthropic
Alignment Fine Tuning, Deliberative Alignment, Synthetic Document Finetuning — alignment-stack…
- Claude's Constitution / Model Spec
Adjacent training method: Deliberative Alignment (treats the spec as in-context for CoT generation)
- Alignment & Safety
Deliberative Alignment — Guan et al. 2025 (OpenAI): SFT on (prompt, CoT, response) tuples with…
- Model Spec Science
Compares to: Deliberative Alignment as a different way to teach spec content
Related articles
- Model Spec Midtraining (MSM)
New training phase between pretrain and AFT: train base model on synthetic docs discussing the Model Spec; controls AFT…
- Alignment Fine-Tuning (AFT)
Standard post-pretraining stage (SFT + RLHF) for installing values; shallow-alignment failure mode motivates [[model-sp…
- Claude's Constitution / Model Spec
Anthropic Model Spec / Constitution by Askell et al.; document specifying Claude's values + hard constraints (SP1–3, GP…
- Agentic Misalignment (AM)
Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
