Na-Li66团队发布了ConvergeFlow,这是一种基于嵌入空间的流语言模型,它通过改进训练方法,在性能上与现有模型相当,值得一试。
ConvergeFlow 是一种基于嵌入空间的流语言模型,通过约束数据预测器到标记嵌入的凸包,并仅用流匹配引起的均方误差目标进行训练。实验表明,ConvergeFlow 在 OpenWebText 上的性能与现有连续和离散扩散语言模型相当。
ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and trains it solely with the mean squared error objective induced by flow matching. Under suitable regularity conditions, we prove that the resulting flow converges to valid token embeddings despite errors in the data predictor, enabling direct token prediction without a CE-supervised decoder. We further develop three sampling mechanisms for controlling the trade-off between the generative perplexity and entropy. Experiments on OpenWebText demonstrate that ConvergeFlow achieves performance competitive with existing continuous and discrete diffusion LMs. These findings demonstrate the potential of the flow-based paradigm for language modeling. Our code is available at https://github.com/Na-Li66/ConvergeFlow.