做AI评估和模型安全测试的团队,终于有了量化标注者偏差的方法论——多级建模直接告诉你需要多少标注才能得到可靠结论,建议做实验设计的点开看看。
生成式AI模型(如LLM)的普及使系统安全性和可信度评估变得至关重要,但当前AI领域面临可重复性危机,主要源于不可靠的评估和不可重复的实验结果。人类评估者引入的偏见和主观意见加剧了这一问题,而现有评估实践通常每个项目仅使用3-5个标注,且缺乏持久评估者标识。该研究提出一种多级自助法(bootstrapping)来建模标注者行为,利用大量标注数据和持久评估者标识,分析项目数量(N)与每个项目响应数(K)之间的权衡,以达成统计显著性。这项工作为改进评估可重复性提供了方法论基础。
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. However, AI is currently facing a reproducibility crisis driven by unreliable evaluations and unrepeatable experimental results. While human raters are often used to assess models for utility and safety, they introduce divergent biases and subjective opinions into their annotations. Overcoming this variance is exceptionally challenging because very little data exists to study how experimental repeatability actually improves as the annotator pool grows. Standard evaluation practices typically rely on a small number of annotations per item (often 3 to 5) and lack the persistent rater identifiers necessary to model individual variance across items. In this work, we introduce a multi-level bootstrapping approach to realistically model annotator behavior. Leveraging datasets with a large number of ratings and persistent rater identifiers, we analyze the tradeoffs between the number of items ($N$) and the number of responses per item ($K$) required to achieve statistical significance.