Chollet确认纯扩展失败,测试时计算才是关键

pure scaling didn’t work. @fchollet confirms exactly what I said in 2022.

精选理由

Chollet 承认以前低估了 LLM,现在说纯 scaling 不行,o3 的 TTC 才是关键。

AI 摘要

François Chollet 在 X 上回应 Gary Marcus 的帖子,承认自己 2023 年至 2024 年初低估了大语言模型的长期重要性。他表示,2024 年 12 月 o3 模型在测试时计算(TTC)上的突破让他改变了观点。Chollet 指出,2023 年初“仅靠扩展基础 LLM 就能解决 AGI”的叙事并未实现。当前基础 LLM 在 ARC 1 上表现不佳,甚至无法可靠完成简单数学运算。

原文 · Gary Marcus

pure scaling didn’t work. @fchollet confirms exactly what I said in 2022.

pure scaling didn’t work. @fchollet confirms exactly what I said in 2022. François Chollet @fchollet One thing I want to make perfectly clear: back in 2023 and early 2024, I was wrong about the role that LLMs would come to play. I underestimated their long-term importance. I have acknowledged this many times. This was the moment I changed my mind, in December 2024, following the o3 test-time compute breakthrough: arcprize.org/blog/oai-o3-pu… I did not initially see that LLMs could work as a base to build systems actually capable of fluid intelligence. Then in late 2024 I updated my views. And here's what did *not* happen: the early 2023 narrative that all we needed to solve AGI was scaling up base LLMs did not pan out. To this day, current base LLMs (considerably scaled up compared to the models from that time) still do not perform well on something as easy as ARC 1 -- and can't even reliably do simple math operations. TTC and harnesses are in fact critical, and the TTC breakthrough was not obvious. 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 540 📊 1 ⚡