Ask Self, Ask Others: Relation Is All You Need

精选理由

Relation模型通过关系组织提供更高效的注意力机制,比MHA表现更优,值得一看。

AI 摘要

Relation模型通过将成对证据组织成显式的Self和Exchange关系,然后推导信息流,实现更高效的注意力机制。Full Relation在10M、30M和100M参数的匹配解码器模型中,最终验证NLL低于MHA。FlashRelation在固定上下文基准测试中比Full Relation快3.60-4.41倍。Hybrid Relation使用75%的Linear Relation层,达到强大的语言建模质量。

原文 · arXiv cs.LG

Attention directly derives normalized information flow from pairwise scores. We introduce Relation, an alternative token-mixing primitive that first organizes pairwise evidence into explicit Self and Exchange relations and derives information flow afterward. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and a KV-style Relation Cache. Across matched decoder-only models at approximately 10M, 30M, and 100M parameters, Full Relation achieves lower final validation NLL than MHA at all three scales. In a fixed-context reference benchmark, FlashRelation is 3.60-4.41x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. Hybrid Relation uses 75% Linear Relation layers and achieves strong language-modeling quality. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.