这篇论文揭示了大型语言模型在不同语言中的技能差异,对于理解多语言模型的发展具有重要意义。
本研究通过多语言自我对弈,量化了大型语言模型在不同语言中技能的不一致性。测试了三个模型在八种语言和六种游戏中的表现,发现同一模型在不同语言中的表现差异显著,包括胜负差距、无效动作和策略倾向。结果表明,技能差异是真正多语言模型发展中的主要障碍之一。
Skill Issue: Are Skills Language-Invariant in LLMs?
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.