LangChain 发布 Tuned Evaluators,自动评分智能体行为,成本降 82%,比前沿模型准。
LangSmith 推出 Tuned Evaluators,自动对生产环境中的智能体行为进行评分。该功能首发指标为 Perceived Error,用于判断智能体是否提供有效帮助。在基准测试中,LangChain 的专用模型击败了所有测试过的前沿模型。该模型将评估成本降低了 82%。
Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, start...
Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that your agent is giving users a helpful experience. In our benchmark, our specialized model outperformed every frontier model we tested and reduced evaluation cost by 82%. Learn more: langchain.com/blog/introduci… 💬 1 🔄 1 ❤️ 7 👀 1516 📊 3 ⚡