论文多源确认83°

长期AI决策能力研究:模型表现远低于人类

精选理由

最新研究揭示,即使是顶尖AI模型在长期复杂任务中也仅达人类四分之一表现。

研究测试了八款领先模型,包括GPT-5.6 Sol和Claude Opus 4.8。最佳配置Qwen3.7-Max搭配Hermes仅达到人类平均表现的27.3%。在需要一年期互联决策和延迟反馈的任务中,AI系统表现远低于人类。

原文 · rohanpaul_ai

This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.

The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.

A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.