ORCA:对齐文本与视觉表示,修复扩散模型组合性失败
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
文生图经常画错“左边的红球”这类提示,这篇论文用低秩对齐把 DiT 训练成本砍半,还顺手超过 REPA 基线。
论文提出 ORCA(Orthogonal Residual Compositional Alignment),针对文生图扩散模型在属性绑定、空间关系、多物体计数上的失败。方法用一个辅助损失将扩散 transformer 的潜变量对齐到冻结视觉编码器导出的低秩目标,正交基由 T5 与 CLIP 嵌入的学习残差参数化。在 DiT-B/2、DiT-L/2、U-ViT-L 三个骨干上,ORCA 的 FID 和 GenEval 均优于 vanilla 与 REPA 基线:DiT-L/2 在 200K 步达到 FID 16.65、GenEval 0.291,超过 400K 步的最强基线。推理阶段零额外开销,增益集中在属性绑定与空间关系类提示上。
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.