论文

SepRQ 开源:无掩码多尺度语音分离 SSL 框架

SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

精选理由

做语音分离或说话人日志的可以看看,SepRQ 开源了,参数才 85.68M 却在 SUPERB 上超过 WavLM,还支持三人混合语音分离。

SepRQ 是一个开源的自监督语音表征学习框架,用冻结随机投影码本上的伪源分离目标替代掩码预测。在 SUPERB 基准上,SepRQ 在说话人日志和语音分离任务中超过 WavLM 等模型,Base 和 Large 规模均领先,推理参数仅 85.68M。它在 DIHARD 3 多域说话人日志数据集和目标说话人 ASR 任务上表现稳定,并在三说话人混合语音 WSJ0-3Mix 上展示了当前 SSL 文献难以企及的分离能力。目前同类 cocktail-party SSL 只有 C-HuBERT 和 SA-WavLM,且均为闭源。

原文 · arXiv cs.LG

SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective over frozen random-projection codebooks. By adopting a novel mask-free, multiresolution approach, SepRQ achieves state-of-the-art performance in Speaker Diarization and Speech Separation on the SUPERB benchmark, surpassing WavLM and other cocktail-party derived SSLs at both Base and Large scales, while requiring only 85.68M inference parameters. SepRQ also demonstrates strong performance across target-speaker tasks requiring enrollment (such as Target-Speaker Automatic Speech Recognition), and on the challenging multi-domain DIHARD 3 diarization dataset. Notably, we report strong separation capabilities on three-speaker mixtures (WSJ0-3Mix), where current SSL literature struggles. While cocktail-party SSLs remain scarce and closed-source, limited to C-HuBERT and the enrollment-based SA-WavLM, we open-source SepRQ to the community.

  • Apple ML Research10-02 00:00原文