Inworld收购Ultravox整合语音与AI代理平台
Inworld收购了语音代理平台Ultravox,整合了语音合成和代理平台,TTS-2模型支持200多种语言且响应极快。
Inworld AI收购了实时语音代理平台Ultravox,现在一家公司同时拥有AI代理的语音和运行平台。Ultravox帮助开发者构建实时语音代理,处理理解、推理、工具调用和回复等多任务。收购后,Ultravox平台上的所有内置语音升级为Inworld的Realtime TTS-2模型,该模型在Artificial Analysis排行榜上名列前茅,支持200多种语言,响应时间低于100毫秒。
. @Inworld has acquired @ultravox_dot_ai , so one company now owns both the voice an AI agent speaks with and the platform the agent runs on.
Ultravox is a platform developers use to build real-time voice agents, which are AIs people talk to out loud, like a support line, a tutor, or a companion app.
An agent like that has to do several things at once: understand what the person said and what they meant, reason about it, call whatever tools it needs to actually get the task done, and reply. Ultravox handles all of that.
It also handles the part that makes these agents feel robotic when it goes wrong, which is timing. The agent has to tell the difference between someone pausing to think and someone actually being done, and if the user cuts in mid-sentence it has to stop talking and listen instead of plowing on.
Inworld is an AI research lab and inference provider, meaning it trains its own speech models and runs LLMs for other companies through APIs used by some of the largest consumer AI apps.
Ultravox customers get the first benefit right away, because every built-in Inworld voice on the platform is now upgraded to Realtime TTS-2, Inworld's text-to-speech model.
TTS-2 is top-ranked on the Artificial Analysis leaderboard, an independent ranking where listeners compare voices blind and vote for the better one.
It supports 200+ languages and starts speaking in under 100 milliseconds, meaning roughly a tenth of a second passes between the text being ready and the first sound coming out.
It also takes natural-language direction, so developers can write plain instructions inside the script, like "yell at the top of your lungs with anger" or "slow down and speak with a warm whisper" and the model changes its pacing, tone, and emotion to match.
The longer-term goal is speech-to-speech: one system that listens, reasons, and speaks, rather than three separate steps handing off to each other.