精选理由
ARC Prize实测了降价后的GPT-5.6 Luna,性能没缩水,跑一次ARC-AGI-2只要0.18美元,性价比确实猛。
ARC Prize团队重新测试了OpenAI的GPT-5.6 Luna,在ARC-AGI-2上取得59.6%准确率,每任务成本0.18美元。在ARC-AGI-1上取得90.7%准确率,每任务成本0.07美元。降价后新结果与原始性能一致,而成本降低了80%。这一价格性能比在推理模型中相当少见。
原文 · Greg Brockman
luna is such a special model, incredible price performance
luna is such a special model, incredible price performance ARC Prize @arcprize We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, $0.18/task - ARC-AGI-1: 90.7%, $0.07/task The new results match Luna's original performance at a much lower cost. 🔗 View Quoted Tweet 💬 3 🔄 2 ❤️ 36 👀 2988 📊 5 ⚡