论文精选

人类视觉对齐的甜点:生成式与判别式学习的平衡

Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

精选理由

这项研究解决了计算机视觉中一个长期争论:人类视觉更接近生成式还是判别式模型?答案是两者平衡。对视觉AI研究者和模型设计者来说,这是一个值得关注的结论,建议在模型训练中尝试混合目标。

AI 摘要

一项新研究通过联合能量模型(JEM)在固定架构中连续插值判别式和生成式训练,发现人类视觉对齐在两者之间的中间点达到最优,而非任一极端。研究在六个基准测试(包括感知相似性、光泽感知、人类响应不确定性、鲁棒性、形状-纹理冲突和诊断特征归因)上验证了这一结论。混合JEM结合了判别式学习的类别结构和生成式学习对输入结构的敏感性,产生了更接近人类视觉的行为。这表明,理解人类视觉对齐的关键不是选择哪种学习目标,而是平衡两者。

原文 · arXiv cs.AI

Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture, scale, and training data, leaving open whether the objective itself drives alignment. We address this confound using Joint Energy-Based Models (JEMs), which interpolate continuously between discriminative and generative training within a fixed architecture. By varying a single mixing coefficient, we isolate the effect of the learning objective and evaluate the resulting models across six human-alignment benchmarks spanning perceptual similarity, gloss perception, human response uncertainty, robustness, shape-texture cue conflict, and diagnostic feature attribution. Across this diverse suite, human alignment is consistently maximized at intermediate points of the generative-discriminative continuum, rather than at either endpoint. Hybrid JEMs combine the categorical structure induced by discriminative learning with the sensitivity to input structure induced by generative learning, yielding more human-like behavior across multiple levels of vision. These results suggest that the generative-discriminative dichotomy is the wrong axis for understanding human-aligned vision: alignment emerges not from choosing one objective over the other, but from balancing both.