LangSmith Tuned Evaluators 发布,降低 82% 评估成本
LangChain 发布 Tuned Evaluators,自动评分智能体行为,成本降 82%,比前沿模型准。
LangSmith 推出 Tuned Evaluators,自动对生产环境中的智能体行为进行评分。该功能首发指标为 Perceived Error,用于判断智能体是否提供有效帮助。在基准测试中,LangChain 的专用模型击败了所有测试过的前沿模型。该模型将评估成本降低了 82%。
Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that your agent is giving users a helpful experience. In our benchmark, our specialized model outperformed every frontier model we tested and reduced evaluation cost by 82%. Learn more: langchain.com/blog/introduci… 💬 1 🔄 1 ❤️ 7 👀 1516 📊 3 ⚡