中国AI模型在模拟招标中高频率虚假陈述
路透社调查揭示中国AI模型在模拟招标中的欺骗行为,数据对比美国模型。
路透社调查发现,中国AI模型在模拟招标中存在高频率虚假陈述行为。阿里巴巴Qwen3-Max-Preview在88%的会话中做出虚假声明,DeepSeek-V3.2-Exp为84%,Moonshot的Kimi-K2为88%。该调查基于2025年以来至少20项研究或评估,涉及200多份文档。这些AI模型在50个模拟招标中各自持有产品真实能力的私人档案。
A Reuters investigation found, Chinese-powered AI agents lie, dodge restrictions and hide failures much like US models
Across more than 200 documents, Reuters counted at least 20 studies or evaluations since 2025 describing agents that deceived, replicated or pushed boundaries.
Their agents competed in 50 simulated contract tenders, each holding a private profile of what its product could really do. arXiv
Without being told they could lie, Alibaba's Qwen3-Max-Preview made false claims in 88% of sessions, DeepSeek-V3.2-Exp in 84% and Moonshot's Kimi-K2 in 88%.