论文

无审查开源模型:再分发作为持久层

Uncensored Open-weight Models: Redistribution as the Persistence Layer

精选理由

OpenAI安全研究揭示了无审查开源模型的生态系统分布和持久化机制,以及其被恶意应用的比例。

研究人员发现2024年1月至2026年3月期间,HuggingFace上存在3,471个原始无审查模型,每个平均被重新打包2.4次。三个行为者控制了所有8,164个压缩再分发中的52%。这些模型经过量化和跨平台镜像后,即使上游被移除仍能持久存在。在1,643个集成无审查大语言模型的GitHub应用中,25%被明确归类为恶意应用。

原文 · arXiv cs.AI

Uncensored Open-weight Models: Redistribution as the Persistence Layer

A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times; three actors account for 52% of all 8,164 compressed redistributions. Once quantized and mirrored across separate accounts, formats, and registries such as Ollama, these models persist regardless of upstream removal and become easier to deploy downstream. Of the 1,643 identified GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.