快讯
Arena发布对齐指数:20轮以上对话失准率至少50%
Arena.ai介绍Arena Alignment Index,用于评估AI智能体真实使用中的安全与对齐表现。该指数基于9万多个真实智能体会话、覆盖27个模型;Arena CEO称,超过20轮的对话中,模型失准率至少为50%。
事件来源
查看原文
RT MTS Arena CEO @ml_angelopoulos reveals that AI models show misalignment in at least 50% of conversations once they pass 20 turns: "The thing that makes our evals different is that they represent the real world post-deployment safety and alignment of these AI models." "They represent what's actually happening when you put them in users' hands, and because of that, we're able to get all sorts of interesting data that you wouldn't get through red teaming and that you wouldn't get from a benchmark." "All these curves are convex and increasing, which means that the longer you go in conversation, the more likely it is that you're going to be witnessing at least one deception, at least one unauthorized action, at least one false attribution." "Once you start getting into a 20-plus turn conversation, the rate at which models are misaligned is at least 50%." "There's still so far to go to make sure that in the distribution of actual organic user tasks, these models exhibit alignment because today they do not." @arena Arena.ai: Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the