论文精选

CAREBench:评估LLM情绪理解的新基准,聚焦认知评价推理

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

精选理由

做AI情感计算或人机交互的团队,这个基准能帮你发现模型在情绪理解上的真实短板——别被下游指标骗了,建议点开看看评价推理链的设计。

AI 摘要

现有LLM情绪理解评估依赖离散标签预测,忽略了情绪产生的认知过程。研究者基于评价理论提出CAREBench,首个包含完整推理链注释的基准,涵盖评价推理、评价评分和多标签情绪标注,从第一和第三人称视角分析真实叙事。实验发现,强模型在某些任务上达到或超越人类,但在评价推理和积极情绪识别上仍有不足;模型在推理链步骤和评价干预敏感性上表现出分离现象,且未内化人类主观异质性的机制。这表明下游情绪预测指标可能高估了LLM的真实情绪理解能力,CAREBench为更诊断性的情感认知评估提供了基础。

原文 · arXiv cs.AI

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion label prediction and fail to capture the cognitive processes underlying emotion generation. Grounded in appraisal theory, we introduce CAREBench, the first benchmark with complete inferential chain annotations from both first- and third-person perspectives on real-world narratives, spanning appraisal reasoning, appraisal ratings, and multi-label emotion annotation. We propose a process-level evaluation framework and conduct systematic experiments across six LLMs organized around four research questions. We find that stronger models match or surpass human observers on certain tasks, yet fall short on appraisal reasoning and positive emotion recognition; performance across chain steps and sensitivity to appraisal interventions exhibit dissociations across models; and current models have not internalized the mechanisms needed to capture human subjective heterogeneity. These findings suggest that downstream emotion prediction metrics may overestimate LLMs' true emotion understanding, and CAREBench provides a foundation for more diagnostically informative evaluation of LLMs' affective cognitive capabilities.