模型精选73°

World Labs发布Atlas模型统一计算机视觉重建与生成

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer v...

精选理由

World Labs的Atlas模型首次统一了计算机视觉的重建与生成,3D采集效率提升50-100倍。

World Labs联合创始人Justin Johnson和李飞飞博士介绍Atlas模型,该模型首次将计算机视觉中分离的3D重建与生成功能统一在一个架构中。Atlas支持文本、图像、视频和相机姿态作为原生输入,使用3D作为原生模态工作。该模型将数字化空间3D表示的采集需求从100-300张照片减少到仅需3张,效率提升50-100倍。

图片来源 · a16z
原文 · a16z

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer v...

World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer vision has always kept apart: Justin: "Historically, reconstruction has been its own subfield in computer vision with its own specialized tasks and models. Generation is what all the text-to-video models are really good at." "Those are great for creative applications if I want to imagine something that's never been there before." "But now with Atlas, for the first time, we're putting these two different parts of visual intelligence together in one model. It can do both 3D reconstruction and generation together in one architecture." "We had to make it multimodal from the start. This thing natively works on text, images, videos, and camera poses as a native input to the model... It uses 3D as a native modality that it works on." Fei-Fei: "Computer vision has been around for more than half a century... Our field traditionally has multiple tracks. You go to a computer vision conference: you have the pixel generation track, you have some recognition track, and you have a 3D reconstruction track." "This is an elegant model that unifies the problem of reconstruction and generation by anchoring on viewpoints and viewpoint estimation. That's just incredibly powerful." @jcjohnss @drfeifei Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: youtube.com/watch?v=qn1QDD… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 4 👀 2802 📊 2 ⚡