Composio 评测 6 款模型跑 30 项智能体任务,Pareto 26.9 以 1/3 成本并列第一
有人把多个模型混着调用,30 个智能体任务测下来追平 GPT-6 Astra,成本只要 1/3,混合路由这事有数据了。
Composio 对 GPT-6 Astra、Opus 5.5、GPT-6 Sol、Pareto 26.9、DeepSeek V4 Pro 和 GLM 5.3 Flash 共 6 款模型进行了 30 项智能体任务评测。Pareto 26.9 由 TheUnbiasedCo 推出,会把请求同时发给多个前沿和开源模型并保留最佳答案,最终与 GPT-6 Astra 并列第一,每个成功任务的成本约为其 1/3。GPT-6 Sol 得分与 Opus 5.5 持平,速度更快,成本约为其 1/4。Pareto 26.9 的完成速度也超过 DeepSeek V4 Pro 和 GLM 5.3 Flash。
Interesting results here. This is why I expect more agent workloads to run on blended models. Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer. In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task. It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash. Composio @composio We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 11 👀 1869 📊 3 ⚡