新基准SPLIT专门测LLM跨语言共情,发现DeepSeek-V3在乌克兰语上最稳,文化基础评价人和AI有分歧。
SPLIT基准测试包含500个提示,覆盖压力、恐慌、孤独、内部流离失所和紧张五类。评估了Gemini-2.5-Flash、LLaMA-3.3-70B-Instruct和DeepSeek-V3三个模型。结果显示DeepSeek-V3在转向乌克兰语时保持稳定,而前两个模型性能下降。人类和AI评估者在共情和自然度上弱一致,但在文化基础上分歧明显。研究强调生成乌克兰语文本不等于提供乌克兰语情感支持。
SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses
Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade when transitioning to Ukrainian, while DeepSeek-V3 remains comparatively stable within our benchmark. We additionally find that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding. We further argue that producing Ukrainian text is not equivalent to producing Ukrainian emotional support. Our findings may assist in the future development of more culturally tailored benchmark designs, as well as encourage a stronger emphasis on human-centered evaluation.