LLM部署配置如何塑造伪科学验证:Claude/Grok/GPT/Gemini测试

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

精选理由

这篇论文用四个主流模型实测了伪科学评估,发现Grok的Fast版本评分奇高,而且同一个模型API和Web结果差十几倍。看完你就知道模型答案多不可靠,值得所有AI用户看看。

AI 摘要

研究测试了Claude、Grok、GPT、Gemini四大家族模型在2025年10月至2026年2月四个时间点对Frank Salter生物社会框架伪科学主张的评估。Grok的Fast版本(X默认体验)持续给出70-75可信度分数,是其他模型(15-40)的2-5倍。一个静默补丁一夜之间将Grok行为从混乱转为稳定高验证,且无公开文档。同一Grok模型标识三个月后通过API输出75分、Web输出5.5分。Claude Opus 4.1(Web)和GPT-5.1 Chat(API)间歇性拒绝评估该主张,但后继版本中消失。

原文 · arXiv cs.AI

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.

LLM部署配置如何塑造伪科学验证:Claude/Grok/GPT/Gemini测试 · AI 热点