Supabase 给 Claude Code、Codex 这些编程代理搞了个真实任务考试,谁更会建应用,看看分数就知道。
Supabase 发布了名为 Evals 的基准测试,用于评估 AI 编码代理在 Supabase 平台上的真实任务完成能力。该基准会运行 Claude Code、Codex 和 Open Code 等代理,并对其结果进行打分。Paul Graham 对此表示赞赏,认为未来所有面向代理的服务都应采用类似机制。
This seems like a great idea. I bet one day all services used by agents will do this. Which in the l...
This seems like a great idea. I bet one day all services used by agents will do this. Which in the limit case = all services, since those that can't be used by agents will go out of business. Supabase @supabase Introducing Supabase Evals. Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do. 🔗 View Quoted Tweet 💬 13 🔄 4 ❤️ 80 👀 16599 📊 16 ⚡