隐喻解释评估新框架:六维度分解人类判断结构

A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

精选理由

这篇论文把隐喻解释的评估从单一打分拆成六个维度,还发现人类分歧是有规律的,做NLP评估的可以看看。

AI 摘要

该研究提出一个认知驱动的多维框架,将隐喻解释质量分解为六个理论维度。在包含11,200条评分的密集标注研究中,发现解释质量确实具有多维性,且标注者分歧是系统性的而非随机。六维度可归为一个共享簇和两个独立判断轴。探索性可行性研究表明,标准自动评估流程能恢复部分结构,对最具区分度的维度预测良好,其误差与人类(不)一致性相关。

原文 · arXiv cs.AI

A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.