论文多源确认

评估LLM生成评分标准的干预转移方法

Evaluating Rubric Generation with Interventional Transfer

精选理由

论文提出了一种创新的评分标准评估方法,揭示了不同LLM生成评分标准之间的不对称性,对AI性能监控有重要启示。

研究人员提出了一种名为干预转移(IT)的新方法,用于评估AI评分标准的生成质量。该方法基于一个核心观点:当回答被扰动以通过/失败某个评分标准时,如果两个评分标准表现出一致性,则它们是相似的。研究团队在HealthBench基准测试中应用了这一方法,分析了Qwen3.8-27B、Deepseek-V4-Flash和Opus-5生成的评分标准对GPT-5.6-Terra回答的评估效果。

原文 · arXiv: DeepSeek

Evaluating Rubric Generation with Interventional Transfer

Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.