对 agent 行为策略,用「品格+元认知解释」式 constitution(说明为何诚实/谨慎)比纯规则清单或纯 RLHF 标签更可迁移;硬约束仍作脚手架。
Claim
对 agent 行为策略,用「品格+元认知解释」式 constitution(说明为何诚实/谨慎)比纯规则清单或纯 RLHF 标签更可迁移;硬约束仍作脚手架。
Why it matters
Goal 含 agent 编排与评测:系统提示/constitution 设计决定越权、撒谎与工具滥用边界,直接影响 harness 策略层。
Summary
杂志长文加厚:Claude's Constitution; or love as the solution to the AI alignment problem。保留原 claim,扩展背景/机制/取舍/动作,并引用源页摘录。
Actions
- (none)
Evidence
- (none)
Caveats
- (none)
Research queries
- (none)
Body
正文
背景
Claude's Constitution; or love as the solution to the AI alignment problem 属于 agent 编排、MCP 工具面或 harness 工程议题。原沉淀 claim 可用,但正文偏短,本轮用一手页 + 既有 claim 加厚为可读长文,便于工作台直接消费。
机制
主张(claim):对 agent 行为策略,用「品格+元认知解释」式 constitution(说明为何诚实/谨慎)比纯规则清单或纯 RLHF 标签更可迁移;硬约束仍作脚手架。
为什么重要:Goal 含 agent 编排与评测:系统提示/constitution 设计决定越权、撒谎与工具滥用边界,直接影响 harness 策略层。
源文要点(摘录/页面):Nintil - Anthropic's Claude Constitution; or love as the solution to the AI alignment problem @font-face {font-family:Dancing Script;font-style:normal;font-weight:400;src:url(/cf-fonts/v/dancing-script/5.2.8/vietnamese/wght/normal.woff2);unicode-range:U+0102-0103,U+0110-0111,U+01
可迁移点:把「工具权限 / 会话边界 / 证据可追溯」写成可检查项,而不是只记录产品名。若源为 GitHub,优先 README 的架构段落与 threat model;若为 essay,抓可执行规范。
取舍
加厚不等于二次全面 research 论文;本轮保证结构完整与 evidence 真实 URL。未拿到全文的部分写在 caveats。产品成熟度、许可证与默认权限仍需人工抽检。
动作
- 打开 https://nintil.com/claude-constitution 对照 claim 是否仍成立
- 把可执行项并入团队 harness checklist(权限/MCP/审计)
- 若需更深:对 repo 再读关键 path 或 arXiv PDF
- 在
/app用 claim-first 阅读,避免再 dump 原始 YAML