研究揭示两类多模态架构的信息融合路径差异
Architecture-Dependent Fusion Pathways in MLLMs
想搞懂多模态模型内部到底怎么把图和文字融合的,这篇把拼接式和原生架构两条路径讲清楚了,做架构选型可以参考。
一篇 arXiv 论文研究了拼接式与原生多模态两类 MLLM 架构的内部融合机制。作者通过三层递进分析——模态对齐解耦、注意力路由与熵、特征空间内在维度——并结合因果干预实验进行验证。结果显示拼接式模型遵循先文本后视觉的融合路径,原生多模态模型则在更早阶段发生视觉与文本的协同适应和特征空间重组。论文还用视觉 CKA 检验了 Platonic Representation Hypothesis 作为补充分析。
Architecture-Dependent Fusion Pathways in MLLMs
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.