AI模型精选

前沿模型解非线性PDE常出错,神经符号模型更可靠

even within math, frontier models are far from being general intelligences. and we still need peopl...

精选理由

Gorard团队实测了GPT-5.6、Fable 5等模型,发现它们连基础非线性PDE都解不对,而Lanyon用更少token就能搞定,想了解模型短板可以看看。

AI 摘要

Jonathan Gorard在X平台发文指出,包括GPT-5.6 Sol、Fable 5和Kimi K3在内的前沿AI模型,在实现基本非线性PDE求解器时普遍失败,常引入数值稳定性、精度阶数或物理一致性方面的显著错误,即使给出详细提示并要求用Lean形式化实现也是如此。即使成功的案例中,较可靠的模型如Fable 5也常消耗超过Lanyon这类轻量神经符号模型100倍的token。其团队认为,只有像Lanyon这样的真正神经符号模型才能为复杂科学问题生成具有端到端正确性保证的超可靠数值求解器,且目前差距明显。

原文 · Gary Marcus

even within math, frontier models are far from being general intelligences. and we still need peopl...

even within math, frontier models are far from being general intelligences. and we still need people to assess which of their output are trustworthy and which are not. Jonathan Gorard @getjonwithit All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stability, order of accuracy, or physical consistency, even when given very detailed prompting and asked to formalize their implementations in Lean. Even in cases where they succeed, the more reliable models (such as Fable 5) routinely consume >100x the tokens of a lightweight neurosymbolic model like Lanyon. Our thesis: only a truly neurosymbolic model like Lanyon is able to produce ultra-reliable numerical solvers for complex scientific problems, with end-to-end correctness guarantees. And at least right, it's not even close. Read more in our latest @lanyon_ai benchmarking post below 👇 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 8 👀 1452 📊 2 ⚡

前沿模型解非线性PDE常出错,神经符号模型更可靠 · AI 热点