研究人员发现,在共享语言模型中,限制模块间的证据可见性,能让模型在组合任务上表现更好,比全局可见的模型平均高出20个百分点,而且学到的接口更可复用。
研究测试了限制证据可见性对梯度训练发现解决方案的影响。在共享基因组语言模型社会中,通过固定中继器仅用两个模型宽度连续向量进行通信。在自然语言函数组合任务上,限制可见性的社会在9/10对中优于全局可见的对照组,深度二和深度三的中位数优势分别为0.7648和0.6050。切断通信会使限制社会降至随机水平,但深度三的优势在复合函数未出现在训练中的程序上仍为0.558。限制可见性并非组合的必要条件,但在此协议下显著增加了泛化中继的概率。
What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies
Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs sharing initialization bytes, training order, token layout, parameters, and computation; only the attention mask differs. Restricted societies outperform their globally visible twins by at least 20 points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and the depth-three advantage remains 0.558 on programs whose composite function never appeared in training. Across six audited restricted societies, same-value packet transplants preserve behavior at 0.94-1.00 across all tested interfaces; destructive interventions collapse performance; and counterfactual packets redirect outputs toward the mathematically predicted answer. The sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Restricted visibility is thus not necessary for composition; under this protocol it substantially increases the probability of a generalizing relay and favors a reusable, value-indexed interface. The complete preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor. An earlier qualification cohort likewise yielded 0/10 complete passes: one model met every task-performance gate, but all ten failed ordinary-language preservation, confining the system to explicitly task-gated use.