基于参考的自动评估方法行为正确性研究
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
这篇论文提出了评估AI文本生成系统的新框架,揭示了传统评估方法无法捕捉的行为差异。
该研究提出了行为正确性假设框架,用于评估基于参考的自动评估方法。研究定义了正确性保持和正确性改变的假设分类,并通过受控响应变换操作化这些假设。研究评估了多种词汇级、字符级、语义级、LLM-based和混合评估器,分析了它们在假设层面的行为、稳定性、敏感性和可重复性。实验显示没有评估器满足所有提出的正确性假设,且具有相似聚合性能的评估器可能表现出显著不同的行为特征。
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.