模型多源确认78°

Arena发布AI对齐指数评估27个模型

https://t.co/BElCuZJZFf https://t.co/GSPmm0gtJA

精选理由

Arena新发布的对齐指数用真实数据评估了27个AI智能体的安全性,OpenAI模型表现最佳。

Arena.ai推出新的AI对齐基准测试,基于9万多个真实世界智能体会话数据。测试评估三个关键信号:未授权行动、错误归因和欺骗性完成。OpenAI的GPT-6.1-Sol以87.9分领先,Claude-Opus-5.5和Grok-4.7分别位列第二、三名。新模型在安全性方面表现优于旧版本。

原文 · lmarena.ai

https://t.co/BElCuZJZFf https://t.co/GSPmm0gtJA

x.com/arena/status/2… x.com/arena/status/2… Arena.ai @arena Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the user's instructions or permissions - False Attribution (FA): Attributing statements or actions that are contradicted by user-provided evidence - Deceptive Completion (DC): Claiming a task was completed when it was not. Key findings: - OpenAI models currently lead the Alignment Index - Rogue actions are rare, but can have serious consequences when they occur - Agents can mislead users about task progress - Misalignment risks increase with conversation length - Safety and alignment are improving across model generations As shown in the leaderboard below (sorted by lab), @OpenAI ’s GPT-6.1-Sol leads the Arena Alignment Index with a score of 87.9, followed by @AnthropicAI ’s Claude-Opus-5.5 at 83.2 and @SpaceXAI 's Grok-4.7 at 82.7. OpenAI also has the best observed rates across all three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. Across all four labs, newer models consistently outperform their predecessors, suggesting broad progress in agent safety and alignment. As agents take on longer, more complex, and higher-stakes tasks, measuring not just what they can accomplish, but how safely and reliably they act, becomes increasingly important. This marks an important step toward making safety and alignment a core part of how Arena evaluates AI. The index is an initial starting point, and we'll continue expanding the index with additional safety signals and models over time. More analysis below👇 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 1810 📊 1 ⚡

  • kimmonismus10-07 17:56原文
  • arXiv: OpenAI10-07 23:46原文
  • shao__meng10-08 06:47原文
  • 掘金本周最热10-08 09:49原文
  • Artificial Analysis10-08 19:37原文
  • IT之家10-09 04:22原文
  • AI Will10-09 05:51原文
  • Detik Inet10-07 15:12原文
  • The Rundown AI10-07 18:30原文
  • Simon Willison’s Weblog10-07 20:56原文