flow-1 发布:用 RL 训练的智能体轨迹纠错模型,成本仅 GPT-6-sol 的 1/23
有个叫 flow-1 的新模型专查智能体运行日志里的错误,性能对标 GPT-6-sol 但便宜 23 倍,做 Agent 监控的可以看看。
flow-1 是一个用强化学习训练的模型,专门在智能体运行轨迹中查找错误。官方称其在轨迹智能上对标 GPT-6-sol,但成本低 23 倍,运行成本也比 GPT-6-luna 低 25%。它让开发者可以不靠抽样、对每一次智能体运行做全量监控与失败排查。作者认为 RL 与 RSI 不只适用于通用智能,也会加速智能体运维这类垂直能力的成本下降。
Like Jev, I believe RL will unlock several more like this, slashing the cost of critical agent operations. flow-1 is competitive in performance, but a huge cost-saver for finding failures in agent traces. RSI doesn't only apply to general intelligence. It will equally accelerate specialized intelligence. Robert @skull8888888888 Introducing flow-1, our new model trained with RL to find errors in agent traces. It matches GPT-6-sol in trace intelligence while being 23x cheaper. It also costs 25% less to run than GPT-6-luna. flow-1 finally makes it possible to monitor and understand every agent run, without sampling. 1/6 Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 3 👀 524 📊 1 ⚡