NVIDIA 开源 Nemotron 3 Diarization 说话人分离模型
Our new Nemotron 3 Diarization model + @pollenrobotics Reachi Mini + DGX Spark Great work Andi @hu...
NVIDIA 开源了个说话人追踪模型,1秒语音就能分清谁在说话,还能让 Reachy Mini 机器人认人记名字,商用许可也友好。
NVIDIA 发布并开源 Nemotron 3 Diarization 模型,可在实时对话中追踪不同说话人,采用商业友好许可证。据 Hugging Face 员工 Andi Marafioti 测试,模型在 1 秒语音片段下表现良好,可用于语音智能体。测试中还配合 Pollen Robotics 的 Reachy Mini 机器人和 DGX Spark,实现识别新声音、询问并记住对方姓名。模型发布即支持 Transformers 集成。
Our new Nemotron 3 Diarization model + @pollenrobotics Reachi Mini + DGX Spark Great work Andi @hu...
Our new Nemotron 3 Diarization model + @pollenrobotics Reachi Mini + DGX Spark Great work Andi @huggingface 🙌 Andi Marafioti @andimarafioti Voice agents still don’t understand who’s speaking to them. That’s a huge gap compared with humans, hidden by all the “phone-call” demos. But that changes today! NVIDIA is open-sourcing Nemotron 3 Diarization: a model that can reliably track speakers in live conversations, under a commercial-friendly license! In my tests, the quality is really good with one-second speech chunks. So we can use it for voice agents! I tested it with Reachy Mini and speech-to-speech running on a DGX Spark. It’s super fun to see the robot notice a new voice, ask for a name, and remember it. The model has day-zero integration with Transformers! Kudos to the NVIDIA team for shipping useful tools for the whole community! Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 4 🔄 13 ❤️ 57 👀 5715 📊 11 ⚡