这篇论文用《红楼梦》测了LLM翻译文化内容的能力,结果发现GPT-4也不好使,连人工打分都吵起来了。适合关注翻译质量的人看。
本研究系统评估了大型语言模型(LLM)在文化负载翻译中的表现。使用《红楼梦》构建了500条中-日双语数据集,覆盖多种文化类别。发现前沿LLM对文化负载内容存在显著性能差距。人工评估中,不同文化背景的评估者导致明显分歧。常用自动评估指标如BLEU无法可靠评估翻译质量。
On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges: (1) task challenges, where frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content; (2) human evaluation challenges, where evaluator backgrounds lead to substantial disagreement in translation judgments; and (3) automatic evaluation challenges, where widely used metrics fail to reliably assess translation quality for this task. These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.