WeightPE 把压缩器嵌进训练,让神经网络权重学会更小的语法
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
有人把压缩器直接塞进训练循环,让 ViT 权重自己长出可压缩语法,压缩率比 int8 QAT 好一半以上,思路挺新的。
论文提出 Weight Pair Encoding(WeightPE),把有损 Re-Pair 压缩器放进 straight-through estimator,使 int8 权重在全局 L2 预算内形成可被语法压缩的重复模式。在 CIFAR-10 上微调的 ViT-B/16 和 ViT-L/16 MLP 权重中,WeightPE 生成的 Re-Pair 语法分别为等效 int8 QAT 的 0.43x 和 0.38x,代价是 1.9 和 1.1 个精度点。与固定大小的 codebook 相比,语法形式支持变长模式并能在更大模式中层级复用。作者还观察到该趋势可迁移到 LZ78、SEQUITUR 等网络未针对性训练的语法压缩器上。据作者称,这是首次把语法大小作为网络权重的显式训练目标。
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.