分层误差归因方法让混合精度量化提速最高2570倍
Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
一篇量化压缩论文:给每层算局部误差打分来分配比特,不用求解器,分配速度最高快2570倍,校准数据脏了也不崩。
一篇arXiv论文提出面向混合精度训练后量化的分层概率误差分析方法,把传播误差与单层局部扰动分开建模。该方法据此构建可分离打分,无需外部求解器即可逐层分配比特。在DRUNet去噪任务、平均每权重4比特预算下,方法在干净校准数据下达到或超过现有基线,在校准数据被污染时PSNR最高提升7.5 dB。比特分配速度较所测基线快28倍到2570倍,直接应用于量化扩散模型也改善了现有最优结果。
Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no external solver. The probabilistic nature of our approach brings robustness to corrupted data. On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions. Experiments show bit-allocation speed-ups from 28x to 2,570x over the studied baselines. For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.