模型

GeneralityLabs Exploit Bench测试结果

精选理由

GeneralityLabs公布Exploit Bench测试结果,V4.1 Flash性能下降,Claude Code超越Codex。

GeneralityLabs在Exploit Bench基准测试中表现出对harness参数的高度敏感性。V4.1 Flash版本性能弱于5.3 Flash版本,在Claude Code中表现最佳,Codex排名第二。Exploit Bench默认基准使用类似Python while循环的原始实现。

原文 · Teortaxes

Another big result from @GeneralityLabs: Performance on Exploit Bench is VERY harness-sensitive. sad news: V4.1 Flash is much weaker than 5.3 Flash, and does best in Claude Code (Codex is 2nd). btw, "original" is not ZCode&DSH, it's ExploitBench default (≈a Python while loop). https://t.co/xIPnOEIez7

  • DeepLearning.AI10-04 02:59原文