研究测试 6 个开源 LLM 给科研基金申请书打分的准确性
Scoring Grant Applications with Large Language Models
拿 2267 份真实英国基金申请,让 Gemma 3 27B 和 DeepSeek R1 等六个开源模型打分,跟人类审稿人比相关系数,结论是初筛能用、终审不行。
一项研究用 Gemma 3 1B/4B/12B/27B、DeepSeek R1 32B 和 Qwen 3 32B 这六个开源 LLM,对 2267 份英国 ESRC 和 EPSRC 的基金申请书打分,并与原审稿人及评审小组成员的分数对比。表现最好的 Gemma 3 27B 在 10 次不同提示词迭代后,与审稿人均分的秩相关为 rho=0.26,但与评审小组成员的相关仅 rho=0.14,远低于成员之间的 rho=0.38。研究者认为相关强度不足以替代终审阶段的人类评审,但可用于初筛环节,例如快速识别最弱申请做桌面拒稿、替代一名人类审稿人或用于偏差三角验证。
Scoring Grant Applications with Large Language Models
Purpose: Assessing grant applications is time-consuming and difficult, adding to the overall burden of academic peer review. Whilst funders are exploring whether AI can help, there is no published research into the accuracy of Large Language Models (LLMs) for scoring contemporary grants. Design/methodology/approach: This study investigates whether six open-weight LLMs (Gemma 3 1B/4B/12B/27B, DeepSeek R1 32B, Qwen 3 32B) can give useful scores for 2267 recent UK Economic and Social Research Council (ESRC), and Engineering and Physical Sciences Research Council (EPSRC) grant applications, comparing them with scores from the original reviewers and funding panel members. Findings: Although the LLM scores are individually inaccurate, when averaged and converted to ranks they correlate positively with expert average scores. The best performing LLM, Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26). Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24, suggesting that it scores are slightly weaker than individual reviewer scores. Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14), which is substantially lower than the inter-panellist correlation (mean rho=0.38). Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state, such as by helping identify the weakest proposals for fast-track desk rejections, to replace one human reviewer, or for triangulation to check for bias.