FrontierHarness Eval 发布模型性能评测

Related tweet: https://t.co/x5CnypfIL8

精选理由

Guanlan 测了 9 个模型,通过率差 17%,成本差 17 倍,帮你选最划算的。

AI 摘要

Guanlan Dai 发布 FrontierHarness Eval,对 Pi、Exo、Claude Code、Codex、DeepSeek Harness 等 9 个模型进行评测。评测包含 360 次运行和 20 亿 token 的处理。通过率从 50% 到 67% 不等,单次通过成本从 1.05 美元到 18.34 美元。

原文 · elvis

Related tweet: https://t.co/x5CnypfIL8

Related tweet: x.com/guanlan/status… Guanlan Dai @guanlan A year ago the question was which model. Now it's which harness. Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime. 360 runs, 2 billion tokens. Pass rates: 50% to 67%. Cost per pass: $1.05 to $18.34. Introducing FrontierHarness Eval. 🧵 🔗 View Quoted Tweet 💬 1 🔄 2 ❤️ 3 👀 1088 📊 2 ⚡