快讯
NVIDIA 论文 VERA:构建可重启沙箱环境,让模型与 harness 协同进化
作者介绍 NVIDIA 论文 VERA:将基准轨迹转化为 9,000 多个可重启沙箱,用 rubric 评分,仅保留能运行且可评分的环境。系统同时更新模型权重与 harness:harness 修改需通过自测且开发集提升至少 5 分,模型检查点得分下降超 20% 则被拒。作者称 9B agent 在 AutoCoWorkBench 和 AutoMedBench 上分别比最强单基线高 10.3 和 13.0 分,27B 版本达 71.6 分,高于 Claude Opus 4.8。环境语料已开源。
查看原文
RT elvis Banger paper from NVIDIA. One exciting trend I am seeing is building verifiable environments for your agents and training the harness alongside the model. Co-evolving the harness and the model is a big part of owning the intelligence stack. And many frontier AI companies have started doing that. This paper shows how this works: VERA turns benchmark trajectories into more than 9,000 restartable sandboxes with rubric scoring, and keeps only environments that run and can be scored from observable evidence. It then updates both the model weights and the harness. A harness edit is kept only if it passes self-tests and adds at least 5 points on the development set. The system rejects a model checkpoint if its score drops by more than 20%. At 9B, the co-evolved agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench. At 27B, it scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8. The environment corpus is open-sourced. Paper: https://arxiv.org/abs/2610.05923 Chat with Paper: https://academy.dair.ai/papers/vera-scaling-verifiable-environments-for-agentic-co-evolution-2610.05923