精选理由
OpenRouter 的评测跑得更快了,同一轮能比推理参数,出错也不再一头雾水。适合经常跑模型对比的人。
OpenRouter 将模型对比评测改为并行运行,普通对比速度提升约 5 倍。新版本支持在单次运行中比较不同的推理努力(reasoning effort)参数,而不只限于模型。检查失败时现在会输出具体错误原因,便于定位问题。该更新面向在 OpenRouter 上运行评估的用户。
原文 · OpenRouter
5/ Eval runs were slow, and failures were unclear Models comparisons now run in parallel instead of...
5/ Eval runs were slow, and failures were unclear Models comparisons now run in parallel instead of one after another, ~5x faster on a normal comparison. A run can also compare reasoning effort, not just models. A failed check also now prints the exact failure. 💬 2 🔄 0 ❤️ 0 👀 323 📊 2 ⚡