arXiv 论文解析 Direct Feedback Alignment 训练停滞的共模崩塌机制
Common-Mode Collapse and Recovery in Direct Feedback Alignment
用 DFA 训网络总卡在常量预测器?这篇论文把根因拆到共模误差和 tanh 饱和,还给了减批量均值这种直接可用的修复办法。
论文研究 Direct Feedback Alignment(DFA)训练中,tanh 隐藏层配合独立 sigmoid 输出时随机梯度下降停滞在类别频率常量预测器损失附近的现象。作者用均值-协方差分解定位出误差的共模分量,发现其主导的秩一更新会把 tanh 单元推向饱和。一个从网络初始化、无拟合参数的简化模型在 48 种设置下预测了激活敏感度的集中程度。在 MNIST 上,校准读出层到类别先验可抑制崩塌并加速学习,减去批量均值也能防止持续崩塌,相关效应在 CIFAR-10 和更深、卷积网络中同样出现。
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.