论文精选

Sakana AI 提出辅助同行评审框架:测试大模型能否发现论文中的植入错误

精选理由

Sakana AI 这篇 TMLR 论文把错误直接种进论文里,再让 AI 审稿去找,还按严重程度分级,比单纯模仿人类审稿的评测靠谱。

Sakana AI 的论文被 TMLR 接收,主题是用 LLM 辅助同行评审而非替代审稿人。团队构建了自动化管线,向论文中故意植入与正文其他内容矛盾的陈述,作为测试 AI 审稿能力的明确目标。为区分错误严重程度,他们将论文的声明、方法与证据构建成知识图谱来估算每个矛盾的严重性。系统方面提出了 Multi-Layered Review,借鉴三遍阅读法,先梳理主旨、再检查细节与弱点、最后汇总成审稿意见。在评估中,该系统检测出的错误数量超过所测试的其他审稿系统,包括在因真实错误而撤回的论文上的表现。

原文 · Sakana AI

Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review

Peer review needs support, not substitutes. Accepted at TMLR: our new paper on using AI to help reviewers catch errors in research papers.

https://t.co/KqDLUPkwib

As research submissions grow, so does the workload for the experts who evaluate them. AI review systems are emerging as a potential solution, but much of their development and evaluation focuses on how closely they imitate human reviews.

In our work “Beyond Imitation: A Framework and Benchmark for LLM Assisted Peer Review”, we explore how AI can support the review process without losing the human touch. We focus on one demanding but essential task: catching errors in research papers.

We build an automated pipeline that deliberately introduces contradictions into papers by adding statements that conflict with information elsewhere in the manuscript. These planted errors give us clear targets for testing whether AI reviewers can spot and explain what is wrong.

Not all errors carry the same weight. Some undermine a paper’s central findings; others affect smaller details. We map the connections between each paper’s claims, methods, and evidence in a knowledge graph to estimate the severity of each contradiction and better assess what different systems can catch.

We also introduce Multi-Layered Review, an AI review system inspired by the Three-Pass Approach to reading research papers. It first outlines the main ideas, then examines the details and potential weaknesses, and finally brings its observations together into a review. The idea is simple: understand the paper before judging it.

In our evaluations, the system detected more errors than the other review systems we tested, including on papers withdrawn because of real mistakes. Its feedback emphasized different aspects of the work from human reviews, offering a complementary perspective, while its assessments of paper quality remained broadly consistent with human judgments.

Our goal is to give reviewers useful support in checking research, with human expertise and judgment at the center.

Open Review: https://t.co/ZaFdyOfEwZ