论文精选73°

大语言模型通过辅助视图提升知识获取

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

精选理由

这篇论文揭示了大模型预训练中知识获取的关键机制,解释了为什么数据多样性重要。

AI 摘要

研究人员通过控制实验发现,在预训练过程中,辅助视图(知识的重新表述)能显著提升大语言模型的学习效果。在固定token预算的情况下,将重复文档的token分配给辅助视图,即使对事实回忆也有帮助。研究表明,辅助视图的有效性不依赖于生成它们的教师模型的强度,且在存在先验知识差距时,情境性和基础性知识形式能促进学习。

原文 · arXiv cs.AI

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.