快讯
Laxman发布Step-Jev:无需标准答案的agent RL逐步奖励方法
Laxman宣布发布Step-Jev,为agent强化学习提供逐步奖励,即使没有标准答案也能使用。该方法用通用验证器(JEV)对过程给予奖励;发帖者CompleteSkeptic对此表示期待,同时希望其在对抗性攻击下足够稳健。正文在介绍实验结果处截断。
事件来源
查看原文
super sick! process rewards using a universal verifier (jev) - hopefully it's adversarially robust enough laxman: Excited to release Step-Jev 🔥 step-by-step rewards for agent RL, even when there’s no answer key After DeepSeek R1, everyone said step rewards get hacked, so just reward the final answer. We found the opposite. With an AI judge grading only the final answer, our agent learned