DeepSeek稀疏注意力有个瓶颈,LongCat这套新方案用流式感知和跨层索引把它治了,跑长文本不慢,还开源了个69B模型,可以试试。
LongCat Sparse Attention(LSA)提出流式感知索引、跨层索引和分层索引三种策略,解决DeepSeek Sparse Attention的O(L^2)计算开销和内存访问不连续问题。LSA在69B-A3B到560B-A27B规模模型上,与全注意力性能持平。LSA支持原生训练最长100万token的上下文,并支撑LongCat-2.0(1.6T-A48B)的开发。团队开源了LongCat-Flash-Lite-Sparse(69B-A3B)。
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.