音频编辑是 AI 落地的重要场景,MMAE 基准揭示了现有模型的巨大短板,做音频 AI 或语音交互的开发者值得关注这个评估工具。
腾讯混元与上海交大、南洋理工等机构合作推出 MMAE,这是首个针对语音和音频编辑的综合评估基准。与单纯生成音频不同,MMAE 要求 AI 理解现有音频并根据自然语言指令精确修改,保留无关部分。基准包含 2000 个真实场景样本、17741 个细粒度评估项,覆盖声音、音乐、语音及其混合的 7 种模态设置。当前模型在精确匹配率(EMR)上低于 5%,揭示了可靠音频编辑的巨大差距。该基准已开源,包含论文、代码和演示。
Can AI truly edit audio, not just generate it? 🎧 Tencent Hy, in collaboration with SJTU, SII, NTU,...
Can AI truly edit audio, not just generate it? 🎧 Tencent Hy, in collaboration with SJTU, SII, NTU, TJU, ZODA, PKU, FDU, and other collaborators, introduces MMAE. MMAE--A Massive Multitask Audio Editing Benchmark, is the first comprehensive evaluation benchmark for speech and audio "Banana🍌" Instead of simply requiring the AI to "generate" audio, it demands that the AI understand an existing audio clip and precisely modify it according to natural language instructions—altering what needs to be changed while leaving the rest untouched. Current models show an Exact Match Rate (EMR) below 5%, revealing a major gap in reliable audio editing. MMAE includes: ✅ 2,000 high-fidelity samples from real-world scenarios ✅ 17,741 fine-grained rubric evaluation items ✅ 7 modality settings across sound, music, speech and their mixtures ✅ 6 task complexity from basic modifications to multi-hop reasoning and multi-round editing ✅ 8 operation types across local and global granularities How to use: arXiv arxiv.org/abs/2606.07229 PZ GitHub: github.com/ddlBoJack/MMAE MD HuggingFace huggingface.co/datasets/BoJac… Jn Demo youtu.be/6At5nTWhlXI k8 Your browser does not support the video tag. 🔗 View on Twitter 💬 3 🔄 15 ❤️ 69 👀 3120 📊 23 ⚡