产品多源确认

Voice Arena 语音数据集 Monsoon 发布,孟加拉语 WER 微调后降至 7.65%

精选理由

微调 Whisper Medium 就把孟加拉语 WER 从 85% 打到 7.65%,80 多家机构一周内排队要授权,低资源语音方向的值得关注。

Voice Arena 在 Interspeech 悉尼大会上宣布,其 Monsoon ASR 语料库上线 7 天已有 80+ 机构申请授权。在孟加拉语 FLEURS 上,用 Monsoon 微调 Whisper Medium 后,LLM 词错率从 85.27% 降到 7.65%。Monsoon 覆盖 50 种语言的 10 万小时数据,以长尾语言为主,计划到 2026 年 2 月扩至 100 种语言,2027 年底达 1000 种。申请方还可通过其模型 API 在自己的内部基准上验证效果。

原文 · elvis

WER dropped from 85% to 7.65% on Bengali with one fine-tune. Numbers like that are why 80 labs asked to license Monsoon in a week. What I respect most is the rule behind it: @voicearena_ai only builds a dataset if it moves the needle. Most data vendors can't say that. Shobhit Banga @shobhitbanga Crazy week for Voice Arena at Interspeech in Sydney. 80+ organisations have asked to license Monsoon ASR corpus since we launched it seven days ago. The most common reason why labs are interested: Monsoon promises results. When we decided to build datasets at Voice Arena, we set one rule. Either the dataset promises results, or we don't build it. Monsoon promises results. On Bengali FLEURS, fine-tuning Whisper Medium on Monsoon took its LLM word error rate from 85.27% down to 7.65%. The other thing labs like is that we give them access to our model API. They can test it on their own internal benchmarks and see for themselves whether it will improve their models. You can see the detailed results and get API access here: voicearena.com/datasets/monso… Monsoon is 100,000 hours across 50 languages from around the world, and most of them are long-tail languages. 100 languages by February, 1,000 by the end of 2027. It's only the first dataset Voice Arena has launched. We are excited to keep working on the key problems that get us closer to the dream of machines that talk like humans. 🔗 View Quoted Tweet 💬 3 🔄 1 ❤️ 7 👀 1234 📊 3 ⚡