← 返回快讯

快讯

Meta超级智能实验室提出"智能体可塑性":按每美元学习收益衡量自我改进模型

omarsar0

Meta Superintelligence Labs的研究人员提出"agent plasticity"指标,在模型权重冻结、每次运行从全新上下文开始的条件下,衡量智能体在留出任务上每美元学习投入的得分增益。研究发现,表现最好的模型往往与学习效率最高的模型不同:在国际象棋、围棋和Hex中,Claude Fable 5取得最高最终得分,而GPT-5.6 Sol的单位美元增益最大;在NetHack中,仅Claude Opus 5.5显著提升66个归一化点,学习成本约1073美元。研究还发现,学习较慢的智能体常忽略自己已写下的产物,而学习更快的智能体会复用这些产物,但产物质量低时仍会失败。

所属事件 →

事件来源

查看原文
Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to know exactly what drives self-improvement, since so many variables are at play (notes, skills, tool calls between runs, etc.). Meta researchers explore and discuss a way to measure whether self-improvement pays off. They call it agent plasticity. It is the gain on held-out tasks per dollar spent on learning, with model weights frozen and every run starting from a fresh context. They find that the model that performs the best is often a different model from the one that learns most efficiently. In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar. In NetHack, only Claude Opus 5.5 improves significantly, by 66 normalized points for about $1,073 of learning. Another interesting finding is that slow learners often ignore artifacts they already wrote. Faster learners reuse their artifacts and still fail when an artifact is low quality. Paper: https://arxiv.org/abs/2610.08902 Chat with Paper: https://academy.dair.ai/papers/agent-plasticity-measuring-self-improvement-through-experience-2610.08902

前后事件