LangChain 做了个叫 ReviewBench 的基准,专门测代码审查能力,数据来自真实 PR 评论,还能用 Harbor 复现。
LangChain 推出 ReviewBench,一个专注于代码审查场景的新基准。该基准从真实 PR 审查意见出发,将审查中发现的问题整理成具体任务,并转化为基于 Harbor 框架的可复现实验。其目标是模拟开发者在实际代码审查中遇到的典型问题,帮助评估模型在审查场景中的表现。
We wanted a benchmark tied to the kinds of issues our reviewers catch in real PRs, so we built Revie...
We wanted a benchmark tied to the kinds of issues our reviewers catch in real PRs, so we built ReviewBench. 1️⃣Start from real reviews 2️⃣Curate them into concrete review issues 3️⃣Turn them into reproducible @harborframework tasks An inside look from @nickhollon10 . LangChain @LangChain x.com/i/article/2083… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 724 📊 1 ⚡