DolphinBench 发布:以动作正确性评测智能体记忆系统
Launching DolphinBench - agent memory benchmark that grades action after long history and measures a...
做智能体的朋友看看这个新基准 DolphinBench,考的是长对话后动作对不对,还把准确率、成本、延迟放一张榜上,选记忆系统时挺实用。
Deshraj Yadav 团队发布 DolphinBench,一个针对智能体记忆的基准测试。它不再用问答形式打分,而是在长历史对话后检验智能体实际执行的动作是否正确。榜单同时记录准确率、成本和延迟三项指标,用于绘制记忆系统的 Pareto 前沿。测试页面在 dolphinbench.ai 开放。
Launching DolphinBench - agent memory benchmark that grades action after long history and measures a...
Launching DolphinBench - agent memory benchmark that grades action after long history and measures accuracy, cost and latency on one board. Deshraj Yadav @deshrajdry Today we are releasing DolphinBench: Mapping the Pareto frontier of agent memory. Coding, tool use, and long-context retrieval already have strong public benchmarks. Memory still mostly gets graded as a quiz: ask a question, judge the answer, maybe report precision and recall. A lot of those boards are getting saturated. They also skip the two questions that matter when you pick a memory system to ship: At what cost? At what latency? Agents are not taking memory quizzes but they take actions. The fact that matters is usually missing from the request. Either it shows up in the tool call, or the agent does the wrong thing while sounding fluent. DolphinBench grades that action after long history, and puts accuracy, cost, and latency on one board. dolphinbench.ai 🔗 View Quoted Tweet 💬 3 🔄 4 ❤️ 23 👀 1174 📊 6 ⚡