AI模型精选78°

Cartesia 发布 Sonic-3.6 流式 TTS 模型登顶榜单

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

精选理由

Cartesia 新出的 Sonic-3.6 语音模型延迟很低,还拿了两个第一。

AI 摘要

Cartesia 发布了基于状态空间模型构建的流式 TTS 模型 Sonic-3.6。该模型在 Artificial Analysis 的 Provider Voice 榜单上以 1,283 Elo 排名第一。它还在 Controlled Voice 榜单上取得 1,123 Elo 的成绩,该基准将所有模型克隆到相同的 8 个参考音色上。Sonic-3.6 实现了低于 90ms 的首字音频生成时间。

图片来源 · marktechpost
原文 · marktechpost

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to isolate the synthesis engine. Cartesia states sub-90ms time-to-first-audio. The model is available in beta on Cartesia's own API The post Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas appeared first on MarkTechPost .