SAP 发布表格基础模型注意力基准测试
Benchmarking Attention for Tabular Foundation Models
SAP 把 TabPFN 这类表格模型的注意力掰开测了一遍,FlashAttention、cuDNN、SageAttention 谁快谁慢、在哪张卡上快都列清楚了,做表格模型推理的可以直接抄作业。
arXiv 论文针对 TabPFN、Mitra、ConTextTab 等表格 in-context 学习模型,系统测试了行注意力与列注意力在 2D 序列上的不同特性。基准覆盖 Torch SDPA、FlashAttention-2/3/4、cuDNN、vLLM、SageAttention 等后端,在 A100、H100、B200 三代 GPU 上测量前向与反向吞吐。结果显示最优后端因注意力方向、硬件和模型而异:FlashAttention 总体最佳,但长序列列注意力下 cuDNN 有时反超,交叉点取决于 head 维度;推理侧 SageAttention 在超过 16k 行的大序列行注意力上表现好。基准与评测代码已在 GitHub 开源。
Benchmarking Attention for Tabular Foundation Models
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark
- shao__meng09-24 01:40原文