这篇论文用70道政治题+12种人格提示,测了GPT-5、Claude、Grok等7个模型,发现提示词比模型本身更能左右回答方向。想了解AI怎么被‘带节奏’的可以看看。
该研究对GPT-5、Claude、Grok、Gemini、DeepSeek、Kimi、Qwen共7个主流模型进行了基于政治轴的可控性压力测试,使用12种意识形态角色提示和70个政治指南针项目,共收集63,700条回答。结果发现,上下文框架解释了经济和社会轴上约88%-93%的方差,而模型自身身份的影响不到3%。在威权提示下,不同模型在相同问题上表现出相似偏移。研究强调需要进行可操控性审计,报告分散性、对称性、饱和度和拒绝底线。数据集和代码已开源。
Auditing Alignment Controllability in LLMs via Political Axes
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.