Artificial Analysis 推出 Cyber Index,评估 AI 模型的企业网络防御能力
网络安全方向的新基准,Grok 4.7 和 MiMo-V2.6-Pro 并列第一,但 GPT-6 Sol 在端到端任务上全部拒绝执行,这个对比挺有意思。
Artificial Analysis 发布 Cyber Index,联合 CollinearAI、IBM、NVIDIA 和 Vercel 组成评估联盟。该指数包含三个基准:CWE-Bench-AA 覆盖 OWASP Top 10 的 120 个审计与修补任务,DeepsecBench-AA 测试漏洞发现,CyberGym-E2E-AA 要求完整复现并修补内存安全漏洞。Grok 4.7 (xhigh) 和 MiMo-V2.6-Pro 以 56 分并列第一,GPT-6 Luna (max) 得 53 分。GPT-6 Sol、Claude Opus 5.5 等多个前沿模型因安全拒绝放弃 32-38% 的任务,在 CyberGym-E2E-AA 上拒绝率最高达 99%,落后榜首 19 至 31 分。
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense
The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best.
We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners.
Benchmarks in the Artificial Analysis Cyber Index:
➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust.
➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized.
➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces.
Key results:
➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44).
➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.