论文

贝叶斯神经网络无需全随机:按先验尺度学习划分随机参数

Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales

精选理由

这篇论文把贝叶斯网络里该随机化哪些参数交给算法自己学,只用一半随机参数就能达到全随机网络的效果,做不确定性估计的可以看看。

论文提出对先验尺度做 deep weight factorization,自动学习贝叶斯神经网络中哪些参数应为随机、哪些为确定。先验尺度低于阈值的参数转为确定性参数并在推理阶段优化,正则项因此稀疏化的是随机性而非模型容量。作者给出可在 O(n) 时间检验的通用条件密度逼近证书及其失败时的最小修复方案,并证明常见的混合采样-优化方案本质是对 type-II MAP 目标的随机逼近。在 UCI 基准上,该方法与全随机网络表现相当,但让约一半参数保持确定性;在双峰目标上,随机掩码的对比方案最多差两个数量级。

原文 · arXiv cs.LG

Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales

Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.