AI代币价值研究:网络文本中的缩放定律
新研究揭示AI代币在模型训练中的价值变化,提出能预测不同规模模型效果的缩放定律。
研究人员分析了2026年6月至8月网络数据,发现经过FineWeb质量过滤后,AI生成文本占比从27.5%上升至31.1%。研究团队预训练了800个语言模型,通过调整AI代币与人类代币的比例,发现数据匮乏模型添加AI代币初期会降低人类文本损失,但添加过多会转为有害。对于预算充足的人类文本训练模型,AI代币几乎立即增加损失,而同等数量的人类代币则持续降低损失。
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
"After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August."
"How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, whil"e the same number of fresh human tokens keeps lowering it."
"We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios."
link: https://t.co/HqVuEu3DEw