核脊回归在幂律各向异性下的渐近分析

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

精选理由

这篇论文揭示了数据各向异性如何影响核脊回归的学习曲线,对理解高维机器学习理论有重要价值。

AI 摘要

该研究分析了各向异性高斯数据下的核脊回归,其中输入协方差以指数α≥0的幂律衰减。研究推导出核谱和泛化误差的渐近精确表达式,揭示了各向异性如何重塑学习曲线。对于弱各向异性(0<α<1),问题保持高维特性,方差在整数样本复杂度κ∈N处达到峰值,但随α增长而逐渐减弱。对于强各向异性(α>1),问题有效维度恒定,方差不再依赖样本量。

原文 · arXiv cs.LG

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<α<1$), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities $κ\in\mathbb{N}$, but these peaks are progressively damped as $α$ grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy ($α> 1$), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.