产品多源确认

dair_ai 创始人用 Slack 里的 AI 员工 Viktor 自动分析 agent 评测结果

Reading eval results is now the slowest part of building agents. I'm Elvis, founder of @dair_ai. I ...

精选理由

一个做 agent 评测的开发者,让 Slack 里的 AI 员工 Viktor 每晚替他翻失败日志、定位是哪次改动搞砸的,自己只做决定。

dair_ai 创始人 Elvis 每晚跑 harness 实验,涵盖 memory、tool use、context compaction 等改动,但逐个翻失败日志占了大量早晨时间。他尝试过 dashboard,只能看到 pass rate 下降,无法定位原因。Slack 里的 AI 员工 Viktor 会通宵复查评测结果,例如在 23 个任务从前一日通过变为失败时,检查全部 23 份日志,追溯到引发失败的那一处改动并建议回滚。Elvis 仍保留最终决策权,Viktor 也会主动标记问题。

原文 · elvis

Reading eval results is now the slowest part of building agents. I'm Elvis, founder of @dair_ai. I ...

Reading eval results is now the slowest part of building agents. I'm Elvis, founder of @dair_ai . I lead research, build, and teach about AI agents. I run harness experiments every night, but reading the results was eating my mornings. Every change to my harness gets evaluated overnight, whether it touches memory, tool use, or context compaction. The morning after is the hard part. I check which tasks my agent got right yesterday but wrong today. Then I open the logs for each failure, one by one, to figure out which of my changes caused it. I tried a dashboard first. It showed the pass rate dropped. It couldn't tell me why. That is the job Viktor, an AI employee in Slack, is built for. He reviews the results overnight. Here is how that plays out. Say 23 tasks that passed yesterday fail today. Viktor checks all 23 logs, traces them to the one change that caused them, and suggests undoing it. I check the logs and make the call. Viktor does the digging. I decide what goes into the harness. He is also proactive. He flags problems before you ask, which helps you stay on track with complex eval runs and other research tasks. Harness engineers, do you check every eval run, or only when the pass rate drops? Try free at @viktor_com . $100 in credits, no card. Full link in my first reply. Thanks to the team for partnering with me on this post 💬 10 🔄 1 ❤️ 16 👀 2279 📊 11 ⚡