前沿模型儿童安全新基准测试:非显式风险失败率2%-34%

This benchmark tests how frontier models handle ch…

精选理由

想看看AI模型在安全上有多不靠谱?这个基准测了5个前沿模型在12类儿童非显式风险上的表现,失败率最高34%。

AI 摘要

该基准测试评估前沿模型处理儿童安全风险的能力,覆盖12个风险类别,包括诱骗、冒充、描述未成年人和情感依赖。测试了5个前沿模型,失败率在2%到34%之间。这是首个针对非显式虐待风险的评估。现有评测仅捕捉显式滥用,完全忽略此类问题。论文链接附于推文中。

原文 · Ate-a-Pi

This benchmark tests how frontier models handle ch…

This benchmark tests how frontier models handle child-safety risks that aren't explicit abuse material: grooming, impersonation, profiling minors, and emotional dependency on AI.

12 risk categories, 5 frontier models, failure rates from 2% to 34%.

This is the first time we have had something like this.

Existing evals catch explicit abuse, but completely miss non-obvious problems.

Here is the paper: https://t.co/VlCSnsTLqk