想评估音频大模型的信息还原能力?MMAC有五千多段音频和十五个细粒度维度,比只看生成质量更靠谱。
MMAC基准包含5,638个音频片段,来自20多个数据源,覆盖6个能力类别和15个评估维度。它检查模型生成字幕是否提及目标维度信息以及内容与参考标签的一致性。评估了开源和商业音频大模型,结果显示不同维度下的信息覆盖度和描述可靠性差异明显。
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.