NeuralZip:可复用设置加速模型权重无损压缩
NeuralZip: Reusable Setup for Fast Lossless Compression
这篇论文搞了个无损压缩模型权重的方法,压缩后比基线快最多 21 倍,还能省 27.5% 显存,做部署的可以看看。
arXiv 论文提出无损压缩框架 NeuralZip,将指数分布相近的分块分组并共享 Huffman 码,对重复出现的指数元组采用打包指数表示。在浮点模型 checkpoint 上,完成设置后的压缩速度比基线快 1.81-21.33 倍,且能做到逐比特完全一致的重建。该设置可预先计算并迁移到兼容架构,训练 checkpoint 在权重演化过程中持续复用。GPU 实验中最多降低 27.5% 的显存占用,同时精确复现 logits。
NeuralZip: Reusable Setup for Fast Lossless Compression
Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.