高保真房间声学模拟使多通道语音增强词错误率降低38%

Improving multichannel speech enhancement through accurate room-acoustic simulations

精选理由

这篇论文用实验告诉你,声学模拟精度越高,语音增强效果越好,SpatialNet在实测数据上词错误率降了38%,训练数据质量是关键。

AI 摘要

该论文对比了不同精度的房间声学模拟方法对多通道语音增强模型的影响。作者使用SpatialNet在基于几何声学、波动声学及其混合模拟的数据集上训练,并在实测数据上评估。高保真模拟数据集训练出的模型,中位数词错误率相对降低了38%。结果表明,高保真房间声学模拟能直接提升多通道语音增强性能。

原文 · arXiv cs.LG

Improving multichannel speech enhancement through accurate room-acoustic simulations

Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on simplified geometrical acoustics, wave-based approaches offer greater physical accuracy. In this work, we examine how simulation fidelity affects multichannel speech enhancement performance. To this end, we train SpatialNet on datasets augmented with different room-acoustic simulation methods and evaluate the resulting models on measured data. We compare lower-fidelity datasets based on geometrical acoustics with a high-fidelity dataset using advanced acoustic modelling and a hybrid combination of wave-based and geometrical acoustics simulations. Training on the high-fidelity dataset results in an up to 38 % relative reduction in median word error rate compared to the lower-fidelity alternatives. These results show that augmentation with high-fidelity room-acoustic simulations directly translates into improved multichannel speech enhancement performance.