精选理由
Harrison Chase讲了他们内部怎么测智能体,就是统一用Harbor,还能把日志变成评测任务。做智能体评测的可以参考这套流程。
LangChain创始人Harrison Chase在X上分享了内部智能体评估的通用方法。他们以Harbor为统一评估基准,并通过专用skill将traces或原始数据转换为Harbor任务。该评估思路计划沿用到LangChain后续的基准测试中。相关帖子获得了1332次浏览。
原文 · Harrison Chase
sharing more about how we evaluate different agents we build internally common themes (from this an...
sharing more about how we evaluate different agents we build internally common themes (from this and future benchmarks): - standardize on harbor - skill for going from traces/raw data to harbor tasks LangChain @LangChain x.com/i/article/2083… 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 12 👀 1332 📊 4 ⚡