大模型通过持续预训练适应瑞典新闻业

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

精选理由

瑞典学者用新闻数据微调大模型,发现经验回放能防止遗忘,提升新闻写作质量。

AI 摘要

研究人员使用数百万篇新闻文章构建高质量数据集,对大模型进行持续预训练以适应瑞典新闻领域。他们创建了涵盖六项编辑任务的专业基准,评估两种模型尺寸的完全和参数高效微调效果。研究发现,仅当结合经验回放以减轻遗忘时,持续预训练才能在目标领域带来收益,提升模型生成质量和事实知识,但不提高判别任务能力。

原文 · arXiv cs.AI

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional training on target-domain corpora. In this work, we investigate such continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that we curate from millions of news articles. To evaluate the adaptation efficacy, we also construct a novel domain-specific benchmark that covers six editorial tasks. Through full and parameter-efficient fine-tuning across two model sizes, we find that continued pre-training yields benefits in the target domain, but only when paired with experience replay to mitigate forgetting. We observe consistent enhancements in the models' generation quality and factual knowledge, but not their proficiency in discriminative tasks. Exploring a training-free method to facilitate instruction following, we see further improvements, but exclusively for models trained with low-rank adaptation. Crucially, we demonstrate the importance of targeted evaluation in the adaptation process, as an existing Swedish benchmark largely fails to capture the models' in-domain performance gains.