让GPT-5.4、Gemini 3.1和Claude 4.6评审300篇ICLR论文,发现它们能区分录用和拒稿,但分不清oral和poster,Gemini给分偏高。用AI审稿前先看看这个。
研究者让OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview和Anthropic Claude Opus 4.6评审300篇ICLR 2026投稿(oral、poster、rejected各100篇),与人类评审和最终决定比较。三个模型都能区分被接收和拒稿的论文,但都无法复现人类评分中oral与poster的区分。Gemini给出的分数系统性偏高,OpenAI和Claude对被拒和poster论文更接近人类,但对oral论文更苛刻。人类评审更常提出计算效率问题,LLM则更常指出缺少基线对比。
How Closely Do LLM Reviews Align with Human Peer Review?
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.