Surge 发布 sudo L7 基准:衡量 AI 能否达到 Staff 级工程师水平
Surge 出了个新基准 sudo L7,用 60 个真实生产仓库任务测 AI 编程,最好的 agent 也只过 45%,想看差距在哪可以读读。
Surge 推出名为 sudo L7 的新基准,用来评估 AI 编程智能体能否胜任 Staff 级软件工程师的工作。基准包含 60 个任务,大多取自真实公司的私有生产代码仓库,由实际维护过这些系统的工程师编写,评分维度涵盖功能正确性、工程工艺、架构判断与是否引入不必要的复杂度。测试结果显示,表现最好的智能体仅完成约 45% 的任务,而给定明确工单写代码这类 L3 级工作它们已经做得很好。
Sweet new benchmark from Surge looking at how close AI is getting to a staff-level software engineering. We're half-way there. echen @echen introducing sudo L7, our new benchmark for staff-level software engineering. coding agents are getting very good at being L3s. give them a well-defined ticket and they can write excellent code. but a staff engineer is valuable for everything that happens around the code. what should we build? what already exists? what's going to break in prod? what did the tests miss? should we even be doing this? sudo L7 has 60 tasks, mostly in private production repos from real companies, written by engineers who've actually owned these systems. expert rubrics grade everything from functional correctness to engineering craft, architectural judgment, thought partnership, and unnecessary complexity. the best agents succeed on only ~45% of tasks. coding is only part of the job. for now, we're keeping them at idiot genius L3. 🔗 View Quoted Tweet 💬 5 🔄 1 ❤️ 5 👀 1846 📊 4 ⚡