Meta新论文提出Skaling law:耦合容量与数据的新缩放定律

Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act...

精选理由

Meta新论文搞了个Skaling law,比Chinchilla和Kaplan那套准1.5到3倍,而且小规模跑就能算准,预算算力能省不少。

AI 摘要

Meta新论文提出Skaling law,通过单一交互指数耦合模型容量与训练数据。该定律将平均绝对百分比误差在插值和外推中降低1.5倍至3倍。在数据稀缺和重度过训练区域,Chinchilla与Kaplan传统定律的预测会漂移,Skaling law的修正幅度最大。配合稀疏网格,只需约十分之一计算量即可外推完整训练网格。论文预印本编号为arXiv:2608.07222。

原文 · elvis

Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act...

Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act on loss independently. This work introduces Skaling law, which couples capacity and data through a single interaction exponent. The extra term cuts mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation. The largest corrections land in the data-scarce and heavy-overtraining regimes where the standard Chinchilla and Kaplan forms drift. Paired with a sparse grid restricted to low-compute runs, it extrapolates the full grid using roughly 10x less compute than a uniform sweep. Why does it matter? Deployment now happens well past compute optimal. A law that stays accurate there, and that can be fit from small runs, changes how a pretraining budget gets planned. Paper: arxiv.org/abs/2608.07222 Track more trending AI papers in our academy: academy.dair.ai 💬 12 🔄 24 ❤️ 197 👀 11277 📊 65 ⚡

Meta新论文提出Skaling law:耦合容量与数据的新缩放定律 · AI 热点