这篇论文提出了一种新数据设计方法,训练出能同时处理多种图像生成任务的大模型,比传统方法更高效。
研究人员提出能力驱动的数据基础设施,构建了440M图像的T2I语料库和120M编辑对。该框架包含三个专门数据引擎,为文本-图像接地、图像间转换和图像-知识关联提供监督。基于此,团队训练了3B和6B两种规模的多模态扩散模型,在CPI-Bench评估中展现广泛视觉覆盖和多功能渲染能力。
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.