预训练数据去重总留尾巴,斯坦福算了一笔账:最坏情况浪费三分之一的算力,模型越大规律越清楚。训练前值得看看。
斯坦福AI实验室的研究(论文编号2606.24998)量化了预训练语料中残留重复数据对计算效率的影响。最坏情况下,重复结构可导致高达33%的FLOPs被浪费。该研究还发现最坏情况重复结构与模型大小存在可预测关系。论文将作为口头报告在ICML 2026深度生成模型基础研讨会上展示。
Deduplication is standard practice but never perfect --this work measures what the residue costs in ...
Deduplication is standard practice but never perfect --this work measures what the residue costs in compute-equivalent terms, and shows the worst-case repetition structure is predictable from model size. The wrong combination can waste as much as 33% of compute! Jessica Chudnovsky ✈️ ICML 2026 @jchudnov Flying to #ICML2026 to present Internal Data Repetition Destroys Language Models, an Oral at Foundations of Deep Gen Models Workshop! Paper: arxiv.org/abs/2606.24998 You might be curious to know what we mean by “destroys”! Pretraining is now data-constrained, and even aggressively deduplicated corpora keep some repetition. We measured what that repetition actually costs in the currency practitioners care about: compute. The answer, in the worst case, is a third of your FLOPs. 🔗 View Quoted Tweet 💬 3 🔄 5 ❤️ 21 👀 3423 📊 5 ⚡