X$^3$-OPD:用策略对齐为音频大模型蒸馏推理能力

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

精选理由

想提升音频模型的逻辑推理?这个框架用文本教师蒸馏,在MMSU等基准上效果拔群,比纯音频训练强很多。

AI 摘要

X$^3$-OPD是一个跨模态策略蒸馏框架,将文本教师模型的推理能力迁移到音频语言学生模型上。训练时学生基于自身声学感知生成推理轨迹,教师用匹配文本和验证答案提供词级指导。框架构建了三层对称语料:文本推理合成语音、复杂声场音频事件推理、含副语言线索的会话推理。在MMSU、MMAU、BIG Bench Audio和MMAR基准上,X$^3$-OPD显著提升了音频推理和思维链质量,且保持了模型在领域偏移下的现有能力。

原文 · arXiv cs.LG

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched textual inputs and verified answers. We further construct a three-tier symmetric corpus covering textual reasoning rendered into speech, audio-event reasoning grounded in complex acoustic scenes, and spoken-dialogue reasoning involving paralinguistic cues. This design extends cross-modal distillation beyond textually recoverable content to reasoning grounded in non-linguistic events, prosody, and conversational context. Experiments on MMSU, MMAU, BIG Bench Audio, and MMAR demonstrate that X$^3$-OPD substantially improves audio-grounded reasoning and chain-of-thought quality while largely preserving the model's existing capabilities under domain shift.