Search agent 的 RL 表现不仅取决于算法与数据,还强依赖 live web 检索后端:用 Exa 作训练期 search tool 可比 Google/SERP 基线以更少 compute 达到更高任务表现。
Claim
Search agent 的 RL 表现不仅取决于算法与数据,还强依赖 live web 检索后端:用 Exa 作训练期 search tool 可比 Google/SERP 基线以更少 compute 达到更高任务表现。
Why it matters
评测与训练 agent 时常默认 SERP/Google;检索质量进入 reward 回路会改变策略学习轨迹,直接影响 agent 编排中 tool choice 与评测基准公平性。
Summary
Mode B 规范化重写:围栏 frontmatter + 正文四段。 Search agent 的 RL 表现不仅取决于算法与数据,还强依赖 live web 检索后端:用 Exa 作训练期 search tool 可比 Google/SERP 基线以更少 compute 达到更高任务表现。
Actions
- (none)
Evidence
- (none)
Caveats
- (none)
Research queries
- (none)
Body
正文
背景
Search agent RL 文献多锁死单一检索后端,忽略 tool 质量对学习动力学的影响。
机制
检索命中质量 → rollout 轨迹多样性与正确性 → reward 信号 SNR → 策略更新效率。
取舍
换后端可抬升样本效率,但绑定供应商与定价;评测需披露后端以免虚高。
动作
在 agent 评测协议中把 search backend 列为一等公民超参。