论文

大模型在分子属性基准中存在数字记忆现象

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

精选理由

最新研究揭示大模型在分子属性预测中存在数字记忆现象,影响模型评估准确性。

研究审计了22个前沿模型在12个回归基准上的verbatim检索行为。在5个数据集上,超过50%的LLM表现出verbatim检索。同一实验在不同推理水平下,高推理水平比最低水平被标记的检索行为多89%。研究还发现,抑制检索会使不同模型的预测误差在相对意义上更接近。

原文 · arXiv cs.AI

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cases, and find that the strongest models in some cases still recognise a combination of transformed SMILES strings and original labels. Furthermore, suppressing retrieval moves the prediction errors of the different models closer together in relative terms, while their differing use of verbatim retrieval spreads them apart. This indicates that the general predictive capability of an LLM is not determined solely by the amount of memorised values. This work provides an overview of the amount and depth of verbatim retrieval in molecular regression benchmarks using LLMs.