超越准确率:大语言模型统计推理的多维评估

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

精选理由

这篇论文用多维方法测了15个大模型的统计推理,不只是比正确率,还看解释风格,发现Anthropic和OpenAI的模型各自有相似的话风。

AI 摘要

一项研究提出结合准确率、响应行为、主题建模和词汇相似性的多维评估框架,用于分析大语言模型的统计推理。该框架应用于15款当前大语言模型对90道统计考题的解答,题目覆盖高中、本科和研究生四个考试。模型准确率介于55%到78%之间。结构主题建模显示所有模型在概念组织上高度一致,词汇相似性分析则发现Anthropic、OpenAI等同一厂商的模型解释风格更为接近。

原文 · arXiv: Anthropic

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\% to 78\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.