嵌入预测提升图像生成质量
Embedding Prediction Helps Image Generation
清华团队提出NEPA方法,用预测嵌入替代传统固定条件,在相同计算资源下生成质量更高。
研究提出NEPA方法,通过Transformer预测连续嵌入序列。在ImageNet 256×256数据集上测试,NEPA-DiT-XL模型达到1.32的FID分数,仅使用REPA模型约三分之一的训练计算量。该方法通过多嵌入预测和嵌入条件生成,使条件信号能适应当前噪声状态。
Embedding Prediction Helps Image Generation
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.