ATFlash 在 1M token 上下文把 Qwen2.5-7B 提速 1.31 倍,精度不掉,而且能直接叠加现有稀疏方法。
ATFlash 提出按 RoPE 每个频率分量的波长设定注意力距离窗口,对 Qwen2.5-0.5B 和 Llama-3.2-3B 可剪掉 37%–48% 的 query-key 内积项。与滑窗不同,低频分量仍能访问所有 key,剪枝率与输入无关,闭式复杂度随序列长度 N 对数增长。在 LongBench-v2 上 top-1 匹配率保持 96%–98%,输出分布 KL 在 10^-3 nat 量级;RULER、OpenAI-MRCR、LongCodeQA、∞Bench 分数基本不变。移植到 FlashAttention-4 和 FlashInfer 后,RTX PRO 6000 上 Llama 推理最高提速 1.29 倍(128K),Qwen2.5-7B-1M 在 1M token 上下文剪掉 57% 内积项,端到端提速 1.31 倍。
ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference
The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. Aligned with this structure, we propose the per-RoPE-wavelength distance window: it prunes the query--key inner-product terms beyond a wavelength-proportional distance. Unlike a sliding window, every key remains reachable, at least through the low-frequency pairs. The reduction rate is input-independent, with a closed form logarithmic in the sequence length $N$, in contrast to dynamic-sparse methods like MInference. Such token-level selection is orthogonal to our frequency-level pruning. The window can therefore be applied on top of those methods. On Qwen2.5-0.5B and Llama-3.2-3B, the window prunes 37--48\% of the query--key inner-product terms within each model's native context length. Relative to full attention, the top-1 match rate stays at 96--98\% and the mean output-distribution KL at the $10^{-3}$-nat level on LongBench-v2 contexts. We examine absolute scores on long-context benchmarks such as RULER, OpenAI-MRCR, LongCodeQA, and $\infty$Bench: they are broadly preserved. We implement the window as a slice of the query--key contraction axis, leaving the online-softmax recurrences untouched, and port it with minimal diffs into the released FlashAttention-4 prefill and FlashInfer decode. On RTX PRO 6000 with Llama, both ports outpace stock with gains growing with context length, up to $1.29\times$ at 128K. End to end on Qwen2.5-7B-1M, with 57\% of the inner-product terms pruned, the speedup reaches $1.31\times$ at a 1M-token context.