精选理由
OpenAI聊了基准测试的细节:得分高低不只靠模型,还跟你用的工具和设置有关,长期智能体还能靠压缩上下文学得更久。
OpenAI指出基准得分不仅反映模型本身,还依赖于测试工具和设置。对于长期运行智能体,保留推理和压缩上下文可帮助模型基于已有内容持续学习。该观点来自OpenAI官方博文,涉及评估方法论。
原文 · OpenAI
A benchmark score reflects the model as well as the harness and settings used to run it. For long-r...
A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already learned. openai.com/index/how-two-… 💬 5 🔄 7 ❤️ 101 👀 14450 📊 14 ⚡