AI模型精选73°

GPT-6 Astra 展示符号世界建模能力

Marcus, 2020, arXiv: We need symbolic, world models Almost everyone but @fChollet: Marcus is wrong ...

精选理由

GPT-6 Astra 能自己创建符号世界模型,在推理任务上接近满分,比人类还高效。

AI 摘要

GPT-6 Astra 在 ARC-AGI-3 基准测试中得分为 66%,使用连续对话测试时接近 100%。该模型能够高效地进行即时符号世界建模,甚至开发自己的特定领域 DSL 表示游戏情况。Astra 在行动效率上几乎超越所有级别的人类基准线,标志着模型智能的重大突破。

原文 · Gary Marcus

Marcus, 2020, arXiv: We need symbolic, world models Almost everyone but @fChollet: Marcus is wrong ...

Marcus, 2020, arXiv: We need symbolic, world models Almost everyone but @fChollet : Marcus is wrong Today: GPT-6 Astra does very well by creating and using symbolic world models. Open questions: 1. how general is Astra’s ability to induce such models? 2. how does it manage to do so? 3. can it follow explicit instructions reliably? François Chollet @fchollet GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: arcprize.org/blog/astra 🔗 View Quoted Tweet 💬 9 🔄 10 ❤️ 75 👀 16344 📊 18 ⚡