论文

小型语言模型在抽象推理任务上的系统性研究

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

精选理由

研究者跑了 1000 多次实验,测小模型学 ARC 网格推理到底靠不靠谱,结论是换个网格大小就崩,做微调的朋友可以看看避坑。

一项研究在 ARC-TGI 基准上系统考察小型语言模型的抽象推理能力,涵盖超过 1000 次运行,比较 decoder-only、encoder-decoder 和 mixture-of-experts 三类模型家族在监督微调下的表现。结果显示模型在训练分布内可以达到较高准确率,但一旦网格尺度变化等分布外偏移出现,性能急剧下降。研究发现技能习得对优化过程敏感,且在不同任务家族间分布不均,额外 in-context 示例的效果也因模型家族而异。注意力诊断显示不同行为特征对应不同的注意力集中模式,但未确立一般性因果机制。

原文 · arXiv cs.LG

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.