Tapered Language Models：沿深度锥形分配MLP参数提升性能

精选理由

这篇论文发现了一个简单技巧：同等算力下，把更多参数分给前几层、少给后几层，模型效果就能更好，试了多种架构都管用。

AI 摘要

当前语言模型在深度上均匀分配参数，但研究表明各层贡献不同。该论文在固定预算下实验发现，将更多参数分配给前层、减少后层可以改进困惑度。提出Tapered Language Models（TLMs），通过余弦调度平滑锥形化MLP宽度。在Transformer、Gated Attention、Hope-attention和Titans四种架构上，三个模型尺度均一致提升困惑度和下游基准性能，且不增加参数或计算量。

AI 翻译 · 中文

arXiv cs.AIModern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This is a default inherit…

阅读原文