快讯
RSIGym为自改进研究智能体提供服务化环境,Opus 4/5研究系统在SWE-bench Verified上从17.67%提升至50.33%
Fanqing提出RSI(递归自我改进)本质上是系统工程问题:RSIGym为研究型智能体提供训练、推理、评估和沙箱作为可调用的服务,让智能体把预算花在实验上而非重建基础设施。发帖者omarsar0赞同这一观点,并称以Opus 5作为研究者,改进后的系统在SWE-bench Verified上的成绩从17.67%提升至50.33%。
事件来源
查看原文
Recommended read. And I agree that the RSI is also a systems engineering problem. Self-improving agents need better research environments. RSIGym gives a research agent training, inference, evals, and sandboxes as services it can call. The agent spends its budget on experiments instead of rebuilding infra. With Opus 5 as the researcher, the improved system went from 17.67% to 50.33% on SWE-bench Verified. Also cool to see a way to measure the quality of co-evolution between harnesses and models, which is how full-stack AI companies stay on the frontier. Fanqing: 💡Our view: RSI is a systems engineering problem, not just a model problem. Progress depends on the environment a research agent works in: what resources it can call, what it can change, and how it runs experiments. That environment should reflect real production workflows and