Cultivar:面向污染与本地化鲁棒性的对比翻译基准

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

精选理由

想测翻译模型有没有背答案?Cultivar 用本地化对照来抓数据污染,32 个开源模型跑下来,发现 MT 专用模型反而更脆,美国内容翻译得最好。

AI 摘要

Cultivar 是一个基于 FLORES 的本地化翻译基准,用于评估模型在特定地区文化语境下的翻译能力。该基准通过源语言对比设计,将未本地化版本与本地化版本配对,以检测数据污染和本地化鲁棒性。研究对 32 个开源权重模型进行测试,发现机器翻译专用模型鲁棒性较差,部分模型可能过拟合 FLORES,且模型普遍更擅长翻译美国内容而非其他地区内容。

原文 · arXiv cs.AI

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.