阿里出了个TTS新模型,能用标签控制笑声耳语,还能自然语言调语速,排行榜第一,支持16语言。
阿里发布Qwen-Audio-3.0-TTS文本转语音模型,提供Flash(实时交互)和Plus(高质量生成)两个版本。新功能包括细粒度内联标签(如whisper、angry、breaths、laughs)和自由自然语言控制(如“读慢点像睡前故事”)。支持16种语言,能从嘈杂参考音频中生成干净输出。支持一次生成长达3分钟的语音。在Artificial Analysis TTS排行榜上位列第一。
Introducing the Qwen-Audio-3.0-TTS. Our latest text-to-speech model, in two flavors: • Flash: real-...
Introducing the Qwen-Audio-3.0-TTS. Our latest text-to-speech model, in two flavors: • Flash: real-time interaction • Plus: high-quality generation What's new: • Fine-grained inline tags-steer [whisper], [angry], [breaths] & [laughs] • Free-style natural-language control-“read this slowly, like a bedtime story” • 16 languages • Clean output even from noisy reference audio • One-pass long-form up to 3 min now #1 on the Artificial Analysis TTS Leaderboard. Blog: funaudiollm.github.io/qwen-audio-3.0… API: alibabacloud.com/help/en/model-… 💬 38 🔄 38 ❤️ 436 👀 17459 📊 95 ⚡