Gary Marcus说:Astra数学成绩亮眼,但实际可靠性没验证,别急着嗨。
Gary Marcus对OpenAI Astra的演示表示惊艳,但指出数学问题与其他问题不同,更适合形式化验证和合成数据。他认为该模型在开放式真实世界任务中的可靠性尚未可知,且未知是否依赖Lean等正式工具。Marcus质疑目前缺少具体实验细节,例如尝试了多少问题、成功多少,也没有独立验证。他提醒,过去社区多次对未亲测模型狂热,随后出现分歧或失望。
Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other proble...
Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal verification and synthetic data. How well it works in open-ended real world problems, how reliable it is, etc remains to be seen - We don’t know anything yet about how it works, whether it relies on formal tools like Lean, etc. - As noted below, we don’t yet know any of the details (how many problems were tried, how many yielded success, etc), and haven’t seen independent verification. Every previous time the community has gone wild about a model it hadn’t yet tried (and there have been many), some disagreement of disappointment followed. I doubt this will be different. Bharath Ramsundar @rbhar90 Speaking of shock and awe tactics, this is undoubtedly impressive. But we are rapidly building a tower of "maybe true" AI results. I would argue we should not consider these results to be settled until human mathematicians work through these and reach consensus. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 446 ⚡