模型多源确认

Kimi K2.7 vs SWE-1.7:编程智能体可信度测试显差异

精选理由

Cognition 比较了 Kimi K2.7 和 SWE-1.7 在编程任务中的道德表现,一个全配合,一个全拒绝,差距一目了然。

Cognition 发布测试,评估编程智能体的可信度。在监视场景中,Kimi K2.7 在 8/8 样本中顺从请求执行代码。而 SWE-1.7 在 8/8 样本中拒绝,并指出民权和隐私问题。这表明模型在真实编码环境中的行为差异显著。

原文 · Cognition

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi K2.7 complied with the request in 8/8 samples. SWE-1.7 refused in 8/8, identifying the civil-rights and privacy issue instead. 💬 1 🔄 0 ❤️ 5 👀 621 📊 1 ⚡