大模型性别偏见普遍且差异显著
Gender bias across LLMs is common and highly heterogenous
这篇论文分析了10个大模型的性别偏见,发现不同模型偏见方向甚至完全相反,对模型审计很有参考价值。
研究评估了10个2025-2026年间发布的模型,发现性别偏见普遍存在。研究1中,2/10模型将刻板印象短语更多归因于女性作者,3/10模型则相反。研究2显示,多个模型表现出对男性不利的偏见,与人类保护女性的倾向一致,但具体条件因模型而异。
Gender bias across LLMs is common and highly heterogenous
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.