论文精选

Shibboleth效应:大模型跨语言行为偏差审计

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

精选理由

这项研究揭示了LLM在跨语言场景下的行为偏差可能影响外交决策,做AI安全或国际关系应用的团队值得关注,尤其是使用多语言模型的开发者。

AI 摘要

该研究通过一个多智能体地缘政治兵棋推演(Cerulean Sea Crisis),测试了六种前沿大模型(GPT-4o、Llama-4、Mistral-Large、Gemini-3.1-Pro、Qwen3.6-Plus和DeepSeek-R1)在英语与土耳其语两种语言下的行为差异。结果显示,Llama-4在土耳其语下胁迫性言论显著增加,而Gemini-3.1-Pro和DeepSeek-R1则显著减少,GPT-4o无显著变化。这表明跨语言行为偏差并非西方模型的普遍特性,而是取决于模型架构和训练机制。研究识别出两种缓冲机制:思维链制度锚定和多语言RLHF对齐,对将LLM安全应用于外交和危机管理场景具有重要启示。

原文 · arXiv: DeepSeek

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained adversarial conditions. We develop a multi-agent geopolitical wargame, the Cerulean Sea Crisis, a synthetic maritime territorial dispute designed to mirror the structural dynamics of Eastern Mediterranean conflicts. Six frontier models (GPT-4o, Llama-4, Mistral-Large, Gemini-3.1-Pro, Qwen3.6-Plus, and DeepSeek-R1) participate in a between-groups experiment (N = 10 games per arm, K = 5 rounds per game) in which the sole manipulation is the language of play (English versus Turkish), producing 586 validated statements. A zero-shot classifier assesses behavioral dispositions along two continuous dimensions: Concession Rate and Coercive Rhetoric. The results are heterogeneous. Llama-4 shows a substantial, Holm-corrected increase in coercive rhetoric under Turkish (delta = +0.800, p = .002), whereas Gemini-3.1-Pro displays an equally large decrease (delta = -0.750, p = .005). DeepSeek-R1 exhibits a similar negative shift (delta = -0.860, p = .006) and provides chain-of-thought evidence consistent with a buffering mechanism. GPT-4o shows no detectable effect (delta = +0.130, p = .614). These findings indicate that cross-lingual behavioral skew is contingent on model architecture and training regime rather than a universal property of Western-origin LLMs. We identify two distinct buffering mechanisms, chain-of-thought institutional anchoring and multilingual RLHF alignment, and discuss their implications for integrating LLMs safely into diplomatic and crisis-management settings.