研究微调过程中视觉语言模型语义能力的获取与退化轨迹
Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning
微调 VLM 的朋友看这个:不同能力峰值时间不同,挑 checkpoint 的标准可能让你丢掉可迁移能力,SSR 方法能改善 name-free 检索。
论文提出基于轨迹的框架,追踪微调过程中身份类与属性类语义能力的获取、峰值与后续特化。研究覆盖 DFN、MetaCLIP 和 OpenAI CLIP 六个预训练骨干。结果显示微调可以在预训练状态之外获得语义能力,且在留出评估上也有提升。作者进一步提出 Structured Semantic Routing(SSR),配合随机名称分支 dropout,可显著增强 name-free 属性检索。不同能力的峰值阶段不同,按目标类名检索挑选的 checkpoint 不一定是可迁移语义的最优点,持续训练可能在保留类名检索的同时削弱先前获得的可迁移能力。
Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning
Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training. We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available.