ScAn-Bench 基准发布:系统评估缩放定律研究方法
ScAn-Bench: Evaluating Scaling Analysis Methodology
研究者用四千多个模型检查点做了两套基准,专门测缩放定律的推导方法靠不靠谱,做 scaling 实验前可以先看看。
arXiv 论文提出 ScAn-Bench-LLM 和 ScAn-Bench-VLM 两个替代基准,分别基于 4524 和 8024 个语言模型与视觉语言模型的训练检查点构建。该工作首次对缩放分析中数据获取与外推方法在不同数据模态下进行系统评估。论文指出,此前没有系统性研究检验缩放定律及其最优配置(架构、数据、超参数)的推导方法本身。
ScAn-Bench: Evaluating Scaling Analysis Methodology
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.