模型多源确认精选

NVIDIA 发布 Nemotron 3 Diarization 说话人分离模型

精选理由

NVIDIA 出了个 100M 参数的小模型,能把重叠的多人对话分清谁在何时说话,做会议转写和播客处理可以试试。

NVIDIA 在 Hugging Face 上发布了 Nemotron 3 Diarization 模型,用于区分音频中不同说话人。该模型参数量为 100M,可处理最多 8 个说话人,并在多人同时说话、声音重叠的场景下识别谁在什么时间发言。模型上线后曾登上 Hugging Face 趋势榜,团队在视频中回答了社区的部分问题。

图片来源 · NVIDIA AI
原文 · NVIDIA AI

Huge thank you to everyone who downloaded Nemotron 3 Diarization and helped it trend on @huggingface ! And we appreciate all the comments. @sabbassi_11 answered a few of your questions: Your browser does not support the video tag. 🔗 View on Twitter NVIDIA AI @NVIDIAAI When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗 Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 12 🔄 12 ❤️ 88 👀 12341 📊 19 ⚡