快讯
NVIDIA 论文 VERA:基准轨迹转为 9000+ 可重启沙箱,协同训练智能体模型与 harness
作者推荐 NVIDIA 论文 VERA,称其将基准轨迹转化为 9000 多个可重启沙箱并附评分 rubric,只保留可运行且能从可观察证据打分的环境,同时更新模型权重与 harness:harness 修改须通过自测且开发集提升至少 5 分才保留,模型 checkpoint 分数降超 20% 则拒绝。作者称 9B 协同演化智能体在 AutoCoWorkBench、AutoMedBench 上分别比最强单维基线高 10.3、13.0 分;27B 在 AutoCoWorkBench 得 71.6 分,高于 Claude Opus 4.8;环境语料已开源。以上为作者声称,未经独立核实。
事件来源
查看原文
Banger paper from NVIDIA. One exciting trend I am seeing is building verifiable environments for your agents and training the harness alongside the model. Co-evolving the harness and the model is a big part of owning the intelligence stack. And many frontier AI companies have started doing that. This paper shows how this works: VERA turns benchmark trajectories into more than 9,000 restartable sandboxes with rubric scoring, and keeps only environments that run and can be scored from observable evidence. It then updates both the model weights and the harness. A harness edit is kept only if it passes self-tests and adds at least 5 points on the development set. The system rejects a model checkpoint if its score drops by more than 20%. At 9B, the co-evolved agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench. At 27B, it scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8. The environment corpus is open-sourced. Paper: https://arxiv.org/abs/2610.05923 Chat with Paper: https://academy.dair.ai/papers/vera-scaling-verifiable-environments-for-agentic-co-evolution-2610.05923