LangChain创始人分享自动化构建Evals的流程

building evals is hard! we're working on some skills to try to automate as much as possible. still ...

精选理由

LangChain创始人分享了他们用Harbor和编码代理自动化评估构建的实操流程,适合做LLM应用评测的人参考。

AI 摘要

Harrison Chase在推文中指出构建评估(evals)难度高,团队正开发技能以自动化流程。其方案包括:用编码代理处理代码库和实际追踪(traces),与用户迭代评估方向,使用Harbor构建评估并运行,再根据结果迭代。流程仍需人类参与,但能加速启动。

原文 · Harrison Chase

building evals is hard! we're working on some skills to try to automate as much as possible. still ...

building evals is hard! we're working on some skills to try to automate as much as possible. still requires human in the loop, but should help bootstrap overall flow is: - give coding agent the codebase + actual traces - iterate on eval direction with user - build evals (using harbor) - run evals - look at results, iterate with user again, repeat Viv @Vtrivedy10 x.com/i/article/2079… 🔗 View Quoted Tweet 💬 1 🔄 1 ❤️ 4 👀 378 📊 1 ⚡