← 返回快讯

快讯

研究称 agent 框架差异主要源于 token 消耗:精简提示词与工具最高可省 3 倍成本

dair_ai

作者称在 SWE-bench Verified 上用同一模型运行 Claude Code、mini-SWE-agent 和 OpenCode:Claude Code 与 mini-SWE-agent 在 447 个任务上得分相差 5 分以内,高难度 45 任务子集上各有 13% 结果翻转。成本差异来自每步重复发送的系统提示词与工具 schema,作者建议精简两者,称可在不降准确率的前提下将成本削减最高 3 倍。

查看原文
Useful paper on what an agent harness changes when the model stays the same. One takeaway: keep your harness prompt and tools small. In this study, that can cut costs by up to 3x without lowering accuracy. The authors ran Claude Code, mini-SWE-agent, and OpenCode with the same model on SWE-bench Verified. Claude Code and mini-SWE-agent scored within 5 points of each other on 447 tasks. Swapping the harness changed results about as much as rerunning it. On a 45-task hard set, both flipped 13% of tasks. The cost difference comes from the system prompt and tool schemas, which each harness sends again at every step. The more steps the agent takes, the more times you pay for them. Paper: https://academy.dair.ai/papers/what-does-a-harness-buy-tokens-mostly-2610.04433

前后快讯