行业73°

多伦多数学家称AI数学能力受验证限制

University of Toronto mathematician Daniel Litt says AI's math capabilities are bottlenecked on veri...

精选理由

多伦多数学家揭秘AI数学瓶颈:能算不会证,800页AI证明无人敢信

AI 摘要

多伦多大学数学家Daniel Litt指出,AI模型无法生成复杂长证明,因为缺乏正确性验证能力。他提到有800页AI生成的数学证明,但模型无法验证其正确性。Litt认为OpenAI和Anthropic可能已解决更多问题,但不确定结果是否正确。AI擅长计算和整合技术,但无法构建理论或保持精确哲学思考。

图片来源 · a16z
原文 · a16z

University of Toronto mathematician Daniel Litt says AI's math capabilities are bottlenecked on veri...

University of Toronto mathematician Daniel Litt says AI's math capabilities are bottlenecked on verification: "My sense is the reason [AI models are] not producing long, complicated proofs is that they cannot. The ability to check correctness is not yet there." "If you ask the models to produce a short proof, you can then ask, 'Is that correct?' And they will often say no... The problem with producing a very long thing is they might not know they're wrong." "What I wonder is, presumably internally, OpenAI and Anthropic have probably solved a lot more problems than they've released. And I imagine quite a few of them, they're just not sure if they're true." "Someone recently posted a claimed proof of resolution of singularities in positive characteristic, which was 800 AI-generated pages. I haven't read it, I haven't found an error, but there's no way it's correct. This would be a major result... Definitely no human has read it. Definitely the models are not able to check this kind of thing yet." @littmath @lishali88 Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z University of Toronto mathematician Daniel Litt and a16z's Lisha Li on AI's impact on mathematics: The models are good at a narrower slice of math than the headlines suggest. They grind long computations, pull technical ideas from more papers than any human could read, and apply every known technique better than almost anyone. What they don't do is build theory, or hold a vague philosophy long enough to make it precise, which is most of what Daniel says he actually does for a living. In this conversation, he and Lisha get into how mathematicians raided an AI proof for parts and broke several other problems with them, why a thousand AI mathematicians might all turn out to be the same mathematician, and why the proof a model handed Daniel was correct but still worth nothing. 00:00 Intro 02:10 The Erdős problem AI disproved 06:20 AI's reasoning looks recognizably human 07:55 Why English beat formal proofs 10:00 Why models can't build theory 14:50 Open problems measure your ignorance 17:45 How a graph became a Millennium Prize problem 18:58 Where AI doesn't help Daniel 21:15 Why ugly proofs are worth doing 23:42 True conjectures are harder than false ones 29:32 10 pages of calculation, zero insight 34:55 The goal of math is not to produce papers 36:25 5 conjectures, 3 bad papers, 1 hour 38:05 One mathematician duplicated 1000x 40:48 Why humans matter even if models win 46:30 When cheaper and worse beats better 49:22 Why the newest AI result isn't a big deal 57:05 How mathematicians actually check a long proof 59:38 Daniel's 3-year-old is already doing math YouTube: youtube.com/watch?v=tQI35C… @littmath @lishali88 Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 4 👀 3614 📊 2 ⚡

多伦多数学家称AI数学能力受验证限制 · AI 热点