A Visual Dependence-Aware Framework for MU-CPT

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

精选理由

Read this if you're interested in the latest advancements in MU-CPT and how to improve cross-modal learning in MLLMs without supervision.

AI 摘要

This paper introduces a new framework for Multimodal Unsupervised Continual Post-Training (MU-CPT), addressing the issue of visual dependence (VD) in MLLMs. The proposed Visual Dependence-Aware (VDA) framework uses Visually Constrained Optimal Transport (VC-OT) and Visually Modulated Adaptation (VMA) to enhance cross-modal learning and maintain task stability. Experiments validate the effectiveness of VDA in MU-CPT settings.

原文 · arXiv cs.AI

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting, and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.