快讯
NVIDIA论文:用“决定性编辑”预测基座模型后训练后的编码智能体表现
NVIDIA一篇论文提出在为编码智能体做后训练之前,对基座模型checkpoint进行排名的新方法。研究者指出基座模型难以直接在智能体编码任务上评估:六个基座模型在SWE-bench Verified上运行,五个解决的任务数为零。新方法改为回放强后训练智能体的代码编辑步骤,找到让测试通过的“决定性编辑”,再考察基座模型能否写出或识别该修复。三种评分方式在十组基座与后训练模型对上,与后训练的SWE-bench Verified得分高度吻合。
事件来源
查看原文
Super interesting NVIDIA paper on choosing base models for coding agents. It's actually a clever way to rank base checkpoints by how well each one is likely to do as a coding agent after post-training. They document that base models are really hard to evaluate on agentic coding tasks. They ran six base models on SWE-bench Verified, and five of them solved zero tasks. So instead, they look at the one step in a coding run that actually fixes the task. In other words, they take tasks that a strong post-trained agent already solved, replay its code edits one by one, and run the tests after each edit. The first edit that makes the tests pass is the decisive edit. Then they give the base model everything that happened before that edit and check whether it can come up with that fix. They score this in three ways. They check how likely the base model is to write the fix, whether it can pick the fix out of a set of rejected patches, and whether any fix it writes on its own passes the tests. All three rankings closely match post-trained SWE-bench Verified scores across ten base and post-trained model pairs. Why is this useful? If you pick checkpoints for agentic post-training, this method can give you a signal before you spend the training budget. Paper: https://academy.dair.ai/papers/before-they-can-solve-predicting-post-training-coding-agent-performance-from-bas-2610.10478