TTS评估新基准:10维度拆解自然度,MOS与Audio-LLM各有短板

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

精选理由

做TTS评测的可以看看,这篇把自然度拆成10个维度,实测发现MOS和Audio-LLM都测不全,还公开了数据集和代码。

AI 摘要

arXiv论文提出首个维度级TTS元评估基准,将自然度拆解为10个语言学感知维度,包含860条由语言学家标注的语音样本。基准测试4个MOS预测器和4个Audio-LLM评判器后发现,MOS预测器仅反映声学信号质量,Audio-LLM评判器对特定提示有选择性检测且无法跨维度泛化。两类方法均不能可靠捕捉语言学结构化的语音错误。数据集、标注模式和评估代码已公开。

原文 · arXiv cs.AI

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.