8月27日
10:48
10:48官方账号arXiv cs.AI@Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu
This paper introduces a new framework for Multimodal Unsupervised Continual Post-Training (MU-CPT), addressing the issue of visual dependence (VD) in MLLMs. The proposed Visual Dependence-Aware (VDA) framework uses Visually Constrained Optimal Transport (VC-OT) and Visually Modulated Adaptation (VMA) to enhance cross-modal learning and maintain task stability. Experiments validate the effectiveness of VDA in MU-CPT settings.
推荐理由:Read this if you're interested in the latest advancements in MU-CPT and how to improve cross-modal learning in MLLMs without supervision.