实时语音助手开发者终于有了速度最快的 STT 模型——每秒处理 303 秒音频,成本还低,做语音交互的团队可以直接在 Together AI 上试试。
Together AI 的语音转文本(STT)模型在 Artificial Analysis 排行榜上包揽了转写速度的前两名。其中 NVIDIA Parakeet TDT 0.6B V3 排名第一,每秒可处理 303 秒音频,速度最快。该模型每 1000 分钟音频仅需 1.50 美元,在三个真实数据集上的平均词错误率为 4.6%。对于构建实时语音助手的 AI 开发者来说,快速 STT 是核心基础设施,Together AI 的云服务能帮助团队降低转录、推理和响应的整体延迟。
Together AI STT models now hold the top two spots …
Together AI STT models now hold the top two spots for transcription speed on the @ArtificialAnlys Speech to Text leaderboard.
NVIDIA Parakeet TDT 0.6B V3 on Together AI ranks #1, transcribing 303 seconds of audio per second of processing time.
→ Fastest STT model measured by Artificial Analysis → $1.50 per 1,000 minutes of audio → 4.6% AA-WER across 3 real-world datasets
For AI natives building real-time voice agents, fast STT is core infrastructure. Running leading speech models on the AI Native Cloud gives teams more room to keep latency low across transcription, reasoning, and response.
Full leaderboard: https://t.co/kU5d39gArF