Jev 模型用于构建自定义验证器,提升代理系统效率
One of the craziest use cases I’ve found for Jev: verifiers. I am so excited about this that I at l...
朋友发现 Jev 模型能用来做自定义验证器,能降低代理系统验证成本,让代理更高效工作。
作者使用 Jev 模型构建了 /goal 特性的自定义验证器,该验证器在每一步后检查目标是否完成,使持续验证成本降低并实现可扩展。这允许更频繁地运行这些验证器(此前由另一个昂贵的推理模型处理),以保持代理系统按计划运行。System One 模型非常适合此验证任务。
One of the craziest use cases I’ve found for Jev: verifiers. I am so excited about this that I at l...
One of the craziest use cases I’ve found for Jev: verifiers. I am so excited about this that I at least wanted to share the high-level idea. I used Jev to build a custom verifier for the /goal feature in my agent harness. It checks whether the goal is actually complete after every turn, making continuous verification cheap enough to scale. This means I can run more of these verifiers (previously handled by another expensive reasoning model) more frequently to keep the agents on track. System One models are perfect for verification. I think of this as scaling harnesses further by cleverly combining System One and System Two models. I have a feeling this will enable a new wave of scalable test-time compute methods. Watch this space closely. I've just started to experiment with this and am already seeing really good results. I need to explore and figure out a way to benchmark it. I will share more once I have more results. This is an insane unlock for long-horizon agents. You heard it here first. And you can expect to see more harnesses embracing this new pattern. Full guide dropping in the next couple of days. Your browser does not support the video tag. 🔗 View on Twitter 💬 15 🔄 3 ❤️ 47 👀 3462 📊 25 ⚡