想知道哪些开放权重模型在 GPQA 和 TAU-Bench 上表现最好?OpenRouter 每周跑分、公开排名,还能帮你选最靠谱的推理服务。
OpenRouter 持续在 GPQA 和 TAU-Bench 上测试多数开放权重模型,并公开发布结果。这些数据用于其 AutoExacto 元基准,该基准默认用于路由工具调用。在最新排名中,Parasail.io 和 Zai_org 位列第一。
Open-weight provider meta-benchmarking, run and published continuously by OpenRouter. Powers our to...
Open-weight provider meta-benchmarking, run and published continuously by OpenRouter. Powers our tool-call routing and ensures you receive the highest-quality inference: x.com/OpenRouter/sta… OpenRouter @OpenRouter Tip: OpenRouter continuously runs GPQA and TAU-Bench on most open-weight models and publishes the results publicly. This informs our AutoExacto meta-benchmark, used by default when routing tool calls. Here, @Parasail_io and @Zai_org rank first: openrouter.ai/z-ai/glm-5.2#p… 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 2 👀 1123 ⚡