7月22日
12:30
12:30官方账号arXiv cs.AI@Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Amanda Darling, Joshua Stallings, David Stern, Shawn Higdon, Claire Duvallet, Bryan Tegomoh, Kenny Workman
新基准BioSecBench-Surveillance包含100个评估任务,覆盖7个类别,从分类到基因工程检测。16个模型-工具组合在3962次尝试中,最佳配置Opus 4.8 with PI仅达50.2%准确率,与GPT-5.5 with Codex并列。Opus 4.7 with PI为49.6%,Sonnet 4.6 with PI为48.6%。即使调用正确工作流,错误仍源于参考序列、阈值、过滤器等选择失误。该基准为衡量下次疫情爆发时AI代理的可信度提供了标准。
事件专题

推荐理由:想看看AI在基因组监测上多不靠谱?新基准BioSecBench-Surveillance测了16个模型,最好才一半正确,细节值得细读。