模型83°

Qwen-Audio-3.1 发布:五个音频模型全面升级并大幅降价

⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next fo...

精选理由

阿里 Qwen 把语音识别、合成、实时对话打包成五个模型一起发,ASR 最高打 0.5 折,做播客或有声书的可以试试 TTS-Next。

Alibaba Qwen 发布 Qwen-Audio-3.1 音频系列,包含 ASR、TTS、Realtime 及新增的 TTS-Next、ASR-Next 共五个模型。ASR 支持多语言与方言识别,自动去除语气词;ASR-Next 支持多说话人标注、时间戳及环境声、情绪理解。TTS-Next 采用 LM+diffusion 统一框架,一次生成语音、音效和背景音频,可用于有声书、播客和游戏。价格方面 TTS 降价约 70%,Realtime 降价约 85%,ASR 最高降价 95%。

原文 · 阿里通义 Qwen

⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next fo...

⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off. Highlights: 🥳 - ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts. - ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning. - TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions. - TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads. - Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood. Unlock the full potential of Qwen-Audio-3.1! 👇 - Blog …shanghai.oss-cn-shanghai.aliyuncs.com/cuijiayan.cjy/… Kg - Qwen-Audio-3.1-ASR qwencloud.com/models/qwen-au… yz - Qwen-Audio-3.1-Realtime qwencloud.com/models/qwen-au… AC - More APIs: coming soon @qwen_cloud 💬 34 🔄 63 ❤️ 608 👀 21429 📊 143 ⚡