Voice Arena 发布 Monsoon 语音数据集:50 种语言、10 万小时
Monsoon 是个 10 万小时的多语言语音数据集,微调 Whisper 后孟加拉语词错率从 85% 降到 7.65%,做语音方向的可以试试 API 测测效果。
Voice Arena 在悉尼 Interspeech 上推出 Monsoon ASR 语料库,覆盖 50 种语言共 10 万小时,主打英语之外的长尾语言。在孟加拉语 FLEURS 基准上,用 Monsoon 微调 Whisper Medium 后,LLM 词错率从 85.27% 降到 7.65%。上线 7 天内已有 80 多家机构申请授权,实验室可先通过其模型 API 在自己的内部基准上测试效果。计划到 2026 年 2 月扩展到 100 种语言,2027 年底达到 1,000 种。
Most voice AI is great in English and falls apart in most other languages. Voice Arena is building the data to fix that: 50 languages now, 1,000 by end of 2027. 80+ labs have already asked to license it. Shobhit Banga @shobhitbanga Crazy week for Voice Arena at Interspeech in Sydney. 80+ organisations have asked to license Monsoon ASR corpus since we launched it seven days ago. The most common reason why labs are interested: Monsoon promises results. When we decided to build datasets at Voice Arena, we set one rule. Either the dataset promises results, or we don't build it. Monsoon promises results. On Bengali FLEURS, fine-tuning Whisper Medium on Monsoon took its LLM word error rate from 85.27% down to 7.65%. The other thing labs like is that we give them access to our model API. They can test it on their own internal benchmarks and see for themselves whether it will improve their models. You can see the detailed results and get API access here: voicearena.com/datasets/monso… Monsoon is 100,000 hours across 50 languages from around the world, and most of them are long-tail languages. 100 languages by February, 1,000 by the end of 2027. It's only the first dataset Voice Arena has launched. We are excited to keep working on the key problems that get us closer to the dream of machines that talk like humans. 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 1272 📊 1 ⚡