想知道智能体在多轮对话里会不会越改越烂?EvoCode 用 26 个任务 227 轮测试,专门抓这种翻车现场。
EvoCode 是一个新评测基准,包含 26 个任务和 227 轮次(每任务 5-15 轮),测试智能体在持续对话中跟随进化需求并保持已有行为的能力。评测采用单一持久容器,每轮后累积测试。常见失败模式不是无法构建新功能,而是实现或更新时破坏了已工作的部分。该评测由 Philipp Schmid 发布。
I hardly write a blog about an eval, but this one felt interesting. EvoCode is a new eval that tests...
I hardly write a blog about an eval, but this one felt interesting. EvoCode is a new eval that tests whether agents can follow evolving requirements and instructions over multiple turns without breaking existing behavior. 26 tasks across 227 sequential rounds (5-15 per task). One persistent container. Cumulative tests after every turn. Common failure isn't "can't build the feature". Agents tend to implement or update something and break something that already worked. Read more about it. ⬇️ Philipp Schmid @_philschmid x.com/i/article/2081… 🔗 View Quoted Tweet 💬 2 🔄 3 ❤️ 34 👀 2626 📊 6 ⚡