GPT-6 Astra评测:规划出色但执行能力差
AGI, eh? “GPT Astra is a terrible executor. It writes awful code. It makes tons of mistakes. It can'...
Gary Marcus评测了GPT-6 Astra,发现它规划能力顶尖但执行糟糕,代码错误率高,远不如Fable和Claude平衡。
GPT-6 Astra被评价为目前最佳规划模型,能处理最复杂的技术任务,也是最好的代码审计工具。该模型在创意和写作方面有所提升,但执行能力仍然很差,无法构建简单的操作序列,代码错误率高。与Fable 5.1和Claude模型相比,GPT-6 Astra在执行质量上仍有明显差距。
AGI, eh? “GPT Astra is a terrible executor. It writes awful code. It makes tons of mistakes. It can'...
AGI, eh? “GPT Astra is a terrible executor. It writes awful code. It makes tons of mistakes. It can't build even the most trivial sequence of actions — it acts first and thinks later. It doesn't consider consequences. Not once has it given me a bug-free result on the first try. Building projects with it is just painful.” Gegam @Gegam245074 First impressions of GPT-6 Astra. Part 2 After using the model for two days, my opinion is no longer so clear-cut Pros 1. Currently the best planner. It can work through even the most complex technical task so thoroughly and unconventionally that no other model can match it. Sol had exactly the same advantage before Astra. 2. The best auditor and reviewer in the world. It finds the trickiest bugs — ones that even an advanced team of QA engineers couldn't uncover. It builds extremely complex chains of logic in its analysis, comes up with hypothetical scenarios, and finds the most unexpected flaws in a project. Before Astra, Sol was the only model I trusted with reviews. 3. UX (not UI) is still the best on the market among models, just as it was with Sol. It builds the most effective and convenient interfaces and prototypes. 4. The model has become a bit more creative. Thanks to the increased parameter count, it doesn't bang its head against the wall endlessly as often. Tunnel vision is still there, but not as often as with the 5-series models. 5. Its writing is more pleasant to read than Fable's. 6. The best computer use in the world Cons 1. The most important, biggest downside still hasn't been fixed. GPT Astra is a terrible executor. It writes awful code. It makes tons of mistakes. It can't build even the most trivial sequence of actions — it acts first and thinks later. It doesn't consider consequences. Not once has it given me a bug-free result on the first try. Building projects with it is just painful 2. It introduces regressions into existing projects. Ask it to add some simple function to a file and it'll wreck your entire project. 3. It still can't complete a task and immediately verify it within a single iteration. And it tailors the tests to fit its own bugs. Never let GPT models write tests! 4. Tunnel vision. One of GPT's most irritating traits hasn't gone anywhere. It can bang its head against the wall endlessly while implementing a project. This happens far less often than with Sol, but it still comes up sometimes. 5. Overcomplication and over-engineering. This problem is still present and hasn't improved by even an inch. 6. The UI is still bad. Conclusions Fable 5.1, like the Claude models, remains unbeaten. These are the most balanced models in terms of intelligence and quality of execution. GPT-6 Astra isn't AGI at all — it's just GPT Sol bloated with more parameters 🔗 View Quoted Tweet 💬 7 🔄 7 ❤️ 54 👀 4368 📊 11 ⚡