技巧精选

OpenAI 提醒:评估结果受 API 设置、提示设计等影响

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they als...

精选理由

想用好 API 的别光看模型分数,OpenAI 告诉你设置和提示更重要,直接抄他们的作业就行。

AI 摘要

OpenAI 表示评估很少孤立测量模型,而是包含 API 设置、评估设计和提示等隐性因素。建议开发者使用 OpenAI 产品中的相同设置:使用 Responses API 而非旧版 Chat Completions API,保留推理(Retain reasoning),并启用压缩(Use compaction)。这些选择会显著影响模型输出表现。开发者可前往 arcprize.org/tasks 参与公开测试。

图片来源 · OpenAI
原文 · OpenAI

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they als...

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting. If you’re an API developer trying to maximize performance, we recommend using the same settings that we deploy in our own products: - Use our Responses API, not our legacy Chat - Completions API - Retain reasoning - Use compaction If you want to test your own mettle against frontier models, try the public games yourself at arcprize.org/tasks 💬 7 🔄 3 ❤️ 91 👀 13409 📊 15 ⚡

OpenAI 提醒:评估结果受 API 设置、提示设计等影响 · AI 热点