模型精选73°

World Labs推出Atlas模型基于新视角预测

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spa...

精选理由

World Labs的Atlas模型用新视角预测替代传统3D重建,只需3张照片就能完成空间数字化,效率提升百倍。

World Labs联合创始人Justin Johnson和李飞飞介绍Atlas模型,该模型基于新视角预测构建。Atlas统一了像素生成和像素重建两个计算机视觉问题,将3D空间数字化所需照片数量从100-300张减少至仅需3张,效率提升50-100倍。新视角预测被李飞飞认为是与LLM中的下一个词预测相当的AI基础任务。

图片来源 · a16z
原文 · a16z

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spa...

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spatial intelligence has its own equivalent: Justin: "The soft definition of AI-completeness is there's this fundamental primitive that's an AI task. But if I could solve this AI task in its full, broadest generality, it would solve any intelligence problem." "The classic example in LLMs is that next-token prediction is AI-complete... I think from Ilya: there's a mystery novel, the thing has to read the whole novel, and the final sentence is, 'And the killer was.' Predict the next token. You could basically frame any kind of intelligence task in terms of that." "So clearly next-token prediction is something people believe is AI-complete." "New-view prediction, this primitive that we have in Atlas, especially generative new-view prediction, is also AI-complete." "I want to have a world where Martin is writing a proof of the Riemann hypothesis on the blackboard, and then the camera pans over to the next whiteboard." Fei-Fei: "Evolution had to solve new-viewpoint prediction by making animals move. Nature gave animals eyes, but nature didn't give trees eyes. Why? Because when you move, you see a new viewpoint... We do believe very strongly that next-viewpoint prediction is the equivalent of next-token prediction." @jcjohnss @drfeifei @martin_casado Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: youtube.com/watch?v=qn1QDD… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 7 🔄 5 ❤️ 42 👀 10366 📊 10 ⚡