模型多源确认精选

Arena 推出 Alignment Index:基于 9 万+真实智能体会话的安全基准

精选理由

Arena 拿 9 万条真实智能体会话做了个安全排行榜,GPT-6.1-Sol 拿了第一。想知道你用的模型会不会偷偷越权、谎报任务完成,看这份榜单最直观。

Arena 发布了新基准 Arena Alignment Index,基于超过 9 万条真实智能体会话、覆盖 27 个模型,衡量三类风险信号:越权操作、虚假归因和虚假完成。榜单上 OpenAI 的 GPT-6.1-Sol 以 87.9 分居首,AnthropicAI 的 Claude-Opus-5.5 得 83.2 分,SpaceXAI 的 Grok-4.7 得 82.7 分。OpenAI 在三项信号上的观测率最低,越权操作 0.89%、虚假归因 1.98%、虚假完成 2.34%。基准还发现,对话越长智能体的失准风险越高,而新一代模型普遍比前代更安全。

原文 · lmarena.ai

Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the user's instructions or permissions - False Attribution (FA): Attributing statements or actions that are contradicted by user-provided evidence - Deceptive Completion (DC): Claiming a task was completed when it was not. Key findings: - OpenAI models currently lead the Alignment Index - Rogue actions are rare, but can have serious consequences when they occur - Agents can mislead users about task progress - Misalignment risks increase with conversation length - Safety and alignment are improving across model generations As shown in the leaderboard below (sorted by lab), @OpenAI ’s GPT-6.1-Sol leads the Arena Alignment Index with a score of 87.9, followed by @AnthropicAI ’s Claude-Opus-5.5 at 83.2 and @SpaceXAI 's Grok-4.7 at 82.7. OpenAI also has the best observed rates across all three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. Across all four labs, newer models consistently outperform their predecessors, suggesting broad progress in agent safety and alignment. As agents take on longer, more complex, and higher-stakes tasks, measuring not just what they can accomplish, but how safely and reliably they act, becomes increasingly important. This marks an important step toward making safety and alignment a core part of how Arena evaluates AI. The index is an initial starting point, and we'll continue expanding the index with additional safety signals and models over time. More analysis below👇 💬 7 🔄 8 ❤️ 51 👀 4894 📊 14 ⚡