Supabase开源Evals,用真实任务评估Claude Code等编码智能体,比普通基准更实战
Supabase开源了Apache-2.0许可的基准框架supabase/evals。该框架让Claude Code、Codex和OpenCode在容器化环境中执行真实Supabase任务,例如构建schema、调试Edge Functions、修复RLS策略。评分结合确定性检查与LLM-as-a-judge。
Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions, fixing RLS policies — inside containerized stacks, then scores them with deterministic checks and LLM-as-a-judge. The post Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks appeared first on MarkTechPost .
- Paul Graham07-31 20:32原文