想省离散扩散采样算力?这篇用UGC度量自动排去掩码节奏,常数个块逼近理论最优误差,给√d维数增益例子。
论文提出路径解析的数据几何测度“去掩码增长复杂度”(UGC),用于分析掩码扩散离散采样。UGC局部增量直接控制KL离散化误差,统一了伯努利子集和固定基数去掩码方案。在log-reveal-odds坐标下,该方法可生成单块与多块优化调度,并量化按数据几何分配计算量的收益。作者证明UGC增量可从样本中估计,从而得到达到给定KL误差的认证最优采样器,迭代复杂度在常数因子内逼近oracle过程。示例显示自适应块相比粗粒度调度可带来√d量级的维度相关改进。
The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity
We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}). Its local increments directly control Kullback--Leibler (KL) discretization error, yielding a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes. In log-reveal-odds coordinates, this structure yields optimized single-block and multi-block schedules, and quantifies the gains from adapting computational effort to data geometry. Crucially, we show how {\textsf{UGC}\xspace} increments can be estimated from samples via KL increments along coupled reveal trajectories. This leads to \emph{certified-optimal} samplers that achieve a prescribed KL error with high probability and iteration complexity within a constant factor of the corresponding oracle procedure. Collapsing the \ugc path yields the aggregate {\textsf{UGC}\xspace} mass, which connects to classical multivariate dependence measures and complexity measures from previous analyses of discrete diffusion. In the fine-partition limit, the squared integral of the square-root {\textsf{UGC}\xspace} density determines the sharp leading-order optimal Euler discretization error. Examples exhibit substantial dimension-dependent gains over coarse schedules, including $\widetildeΩ(\sqrt{d})$ improvements achievable with a constant number of adaptively placed blocks.