研究发现 LLM 文献检索准确性受提示词类型影响:摘要优于全文
Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts
六个主流模型实测文献检索,结果挺反直觉:给摘要比给全文更准,列表越靠后越容易编,写文献综述前值得看看
一项研究量化比较了 Claude、ChatGPT、Grok、DeepSeek、Perplexity、Gemini 六个平台在环境科学文献检索中的错误率。测试对象是 Energy and Environmental Science、Nature Sustainability、Nature Climate Change 等 5 本期刊 2024 至 2025 年的 50 篇原创文章,每篇让模型检索 10 条参考文献。用摘要作提示词的检索准确性显著高于用全文,这一结论经多水平混合效应回归在调整期刊、平台和输出顺序后依然成立。参考文献在输出列表中位置越靠后,准确性越低,且存在完全编造的文献条目。
Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts
Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.