AI模型精选73°

OpenAI GPT-6 Astra 引发争议

Hot take on OpenAI GPT-6 Astra*, with a challenge to @gdb’s claims about it being AGI toward the end...

精选理由

OpenAI 的 GPT-6 Astra 在 ARC-AGI 测试中表现亮眼,但专家对其是否真正实现 AGI 提出质疑。

AI 摘要

OpenAI 发布的 GPT-6 Astra 在 ARC-AGI-3 测试中达到 63% 分数,通过新提供商适配器达到 99%。该模型在 96% 的 ARC-AGI-3 关卡上超越人类表现。Astra 被认为构建了迄今为止最精确的新环境符号模型。Gary Marcus 对其是否为 AGI 表示质疑,指出其在开放现实世界任务中可能存在问题。

图片来源 · Gary Marcus
原文 · Gary Marcus

Hot take on OpenAI GPT-6 Astra*, with a challenge to @gdb’s claims about it being AGI toward the end...

Hot take on OpenAI GPT-6 Astra*, with a challenge to @gdb ’s claims about it being AGI toward the end: • Looks to be pretty impressive. Multiple reports suggest it is a genuine advance. • As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its computations. • What we don’t know is how robust that capability is. That is THE key question. • Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains. • And as a scientist, it’s disappointing that we don’t (yet) know much about how the system actually works. • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well. • The new system appears to be *less* monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK; link: garymarcus.substack.com/p/where-will-a… ) ——————- *This hot take is VERY tentative, pending more information about how it works and what its limitations are. ARC Prize @arcprize GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: 🔗 View Quoted Tweet 💬 17 🔄 8 ❤️ 131 👀 41822 📊 32 ⚡