XSQ-AST:无需重训练即可定位合成语音伪影的可解释框架
XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts
合成语音听着怪却说不清哪里怪?XSQ-AST 用显著性图把伪影定位到具体音素,40 人听测验证有效。
arXiv 论文提出 XSQ-AST,将 SQ-AST 语音质量模型与 WhisperX 音素对齐、多种显著性方法结合,无需重训练模型即可产出带时间定位的伪影诊断。框架通过核密度估计把显著性图转成连续分布,再按音素边界离散化。40 人参与的听测在 5 个感知维度上验证,Attention Rollout、Attention Flow 和改编版 GradCAM 生成的分布均与听众标记相关,但不同伪影类型各有最适合的方法。AUC-ROC 分析确认其区分度高于随机水平。
XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts
Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.