两层量化神经网络STE训练稳定性研究
Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks
这篇论文从统计学习理论角度解释了STE训练的稳定性与泛化能力,给出了明确的误差边界和收敛速率。
该研究从统计学习理论角度分析了两层二值激活网络的身份直通估计器(STE)。研究显示在饱和输出区域,零初始化的样本级STE递归等价于凸潜在损失上的随机次梯度下降。研究团队推导了两个耦合更新的精确距离恒等式,证明了公共示例映射的近似非扩张性。在可分离边界条件下,研究获得了最优阶O(R²/(γ²n))的期望超额分类错误率。
Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks
We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(γ^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.