Meta把语音识别、说话人分离和端点检测三个模型合成了一个,延迟更低,故障点更少。
Meta Superintelligence Labs本周发布Muse Voice Transcribe模型,将语音识别、说话人分离和端点检测三个功能整合为一个自回归模型。该模型专为实时流式语音处理设计,解决了传统语音系统中多个模型交接带来的延迟和故障点问题。Muse Voice Transcribe通过单一模型完成三项任务,提高了语音处理的效率和可靠性。
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode. Muse Voice Transcribe, announced by Meta Superintelligence Labs this week, collapses those three jobs into a single autoregressive model. Meta calls […] The post Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing appeared first on MarkTechPost .