论文

LFRM:用连续潜在扩散做推理的新训练方案

Reasoning with Continuous Latent Diffusion

精选理由

一篇把扩散模型搬进推理赛道的新论文,LFRM 用 638M 参数在 GSM8K 拿到 63.74%,比同规模的连续扩散基线都强,搞生成范式研究的可以看看。

论文提出 Latent Flow Reasoning Models(LFRM),基于 ELF 训练推理方案,让连续扩散模型在潜在空间通过迭代精化生成完整推理过程。方法从自回归教师模型的多层中学习紧凑表征,并用分阶段课程训练出可在推理时替代教师 Transformer 的紧凑 prompt 编码器。638M 参数去噪骨干的 post-NFT LFRM-L 在 GSM8K 上达到 63.74% pass@1、MATH500 上 24.6%(64 步去噪),HumanEval 上 32.85%、HumanEval+ 上 30.18%(128 步去噪)。作者还在数学推理和代码生成任务上超过了同等骨干规模下近期连续扩散基线的报告成绩。

原文 · arXiv cs.AI

Reasoning with Continuous Latent Diffusion

Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT LFRM-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/LFRM