变换乘法结构保持参数量:Transformer的关联代数层
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
MIT团队用新乘法结构提升Transformer速度,参数不变但生成更快,适合关注模型效率的研究者。
研究人员提出了一种基于关联代数的Transformer新架构,使用稀疏交互表替代普通矩阵乘法。实验训练了两个约1.1亿参数的解码器模型,仅在前馈层使用不同乘法方法。在四个提示域测试中,代数模型实现了6.2-7.8%的端到端生成吞吐量提升,但在三个下游指标上得分较低。
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.