别只看参数数量,量子电路里的纠缠才是泛化关键。这篇论文用PAC-Bayesian理论加实验验证,帮你理解怎么设计泛化更好的量子策略。
该论文通过PAC-Bayesian框架分析了参数化量子电路(PQC)作为策略和价值函数时的泛化机制。研究发现,泛化由Fisher几何有效维度主导,而非参数数量,且纠缠会增大该维度,成为独立的复杂性轴。在固定参数数量的实验中,Fisher有效维度更大的电路显示出更大的训练-测试差距,纠缠电路比非纠缠电路泛化更差。分类、上下文强盗和值函数实验中,效果显著,且在IBM Heron量子处理器上真实噪声下依然成立。
Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions
Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize. We give a PAC-Bayesian account in which generalization is governed not by the raw number of circuit parameters, but by the effective dimension of the Fisher geometry induced by the circuit. This quantity is inflated by entanglement, making entangling connectivity an independent axis of complexity.In controlled experiments that fix the number of trainable rotations and vary only entanglement, we find that circuits with larger Fisher effective dimension exhibit larger train-test gaps, while parameter count is a weak predictor. The resulting bound acts primarily as a ranking certificate: it correctly orders circuits with identical parameter count, which parameter-counting bounds cannot do. We validate this mechanism across supervised classification, quantum contextual bandits, and value-function generalization, where entangled circuits consistently generalize worse than non-entangled circuits of equal parameter count, with gaps shrinking as sample size increases.Our strongest evidence comes from low-variance decision models, including single-observable classifiers, value heads, and one-step policies. In end-to-end multi-step policy learning, entanglement effects remain statistically significant but high return variance leaves the full ordering only partially resolved. Partial-correlation analysis shows that Fisher effective dimension screens off entangling pattern, and controls for training accuracy, readout, and optimizer rule out major optimization confounders. The effect also persists on an IBM Heron quantum processor under real noise. Overall, our results reframe quantum policy design around an entanglement--generalization trade-off rather than expressivity alone.