这篇论文实打实比较了GPT-4o、Mistral和DSIT-Taxonomies在提取基金提案实体上的能力,Mistral准确率90.5%碾压对手,做科研数据挖掘的可以看看。
这篇论文比较了GPT-4o、Mistral和DSIT-Taxonomies算法从42份UKRI基金提案摘要中提取研究实体的效果。Mistral实现了90.5%的主题分类准确率,远超DSIT-Taxonomies的71.4%。Mistral与GPT-4o的实体集质量相当且语义重叠度高,但Mistral在操作效率和安全性上更优。研究依托OpenAlex Topics分类体系,为大规模敏感数据分析提供参考。
Research Entity Extraction and Topic Detection from UKRI Grant Proposals
This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke algorithm, DSIT-Taxonomies, for extracting and classifying research entities from funding proposals. Our project "Tracking Stars and Unicorns" aims to identify early signals of emerging research areas to inform public investment. Our methodology employed a three-stage pipeline, leveraging Mistral for primary entity extraction and mapping against the OpenAlex Topics taxonomy. We evaluated our approach across 42 proposals' abstracts from different areas and observed that Mistral and GPT-4o produce comparable, high-quality entity sets with significant semantic overlap, outperforming the fragmented DSIT-Taxonomies approach. Crucially, the Mistral-based approach achieved superior topic classification accuracy (90.5%) compared to the full DSIT-Taxonomies pipeline (71.4%). We conclude that Mistral offers a high-performance, operationally efficient, and secure solution for large-scale analysis of sensitive grant data.