← 返回快讯

快讯

研究:模型不变时,agent harness 可带来最高 3 倍成本差异

omarsar0

作者称在 SWE-bench Verified 上用同一模型运行了 Claude Code、mini-SWE-agent 和 OpenCode:前两者在 447 个任务上得分相差不超过 5 分,更换 harness 的影响约等于重跑一次(45 题困难子集上各有 13% 结果翻转)。成本差异主要来自每一步重复发送的 system prompt 和 tool schema;精简 harness 提示词和工具可在不降低准确率的情况下最多降本 3 倍。

查看原文
RT DAIR.AI Useful paper on what an agent harness changes when the model stays the same. One takeaway: keep your harness prompt and tools small. In this study, that can cut costs by up to 3x without lowering accuracy. The authors ran Claude Code, mini-SWE-agent, and OpenCode with the same model on SWE-bench Verified. Claude Code and mini-SWE-agent scored within 5 points of each other on 447 tasks. Swapping the harness changed results about as much as rerunning it. On a 45-task hard set, both flipped 13% of tasks. The cost difference comes from the system prompt and tool schemas, which each harness sends again at every step. The more steps the agent takes, the more times you pay for them. Paper: https://academy.dair.ai/papers/what-does-a-harness-buy-tokens-mostly-2610.04433

前后快讯