World Labs 推出 Atlas 模型实现 3D 场景重建
Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-fo...
World Labs 的 Atlas 模型能用 3 张照片完成过去需要 300 张照片才能做到的 3D 场景重建,还能处理旧视频素材。
World Labs 发布 Atlas 模型,将传统需要 100-300 张照片的 3D 场景重建减少至仅需 3 张。该模型统一了像素生成和像素重建两个计算机视觉领域分离了半个世纪的问题。斯坦福大学演示显示,使用 3-25 张图像即可重建整个斯坦福广场。
Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-fo...
Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-founders Ben Mildenhall and Fei-Fei Li on how Atlas got there: Ben: "In a casual sense... I took three photos of this object, or six photos of this room. I look at the photos, I can understand in my mind how those piece together. I can fill in the gaps and get it." "But there's never really been any reconciliation between those data-driven priors and the brute force dense reconstruction, which is much more akin to scientific or medical imaging... When we say dense, we really mean dense." "This room, I want like 100, 200, 300 photos to capture it. And what we're trying to do is bring that down to like three. We're saying like 50, 100x reduction." "At that scale it completely flips that calculus on its head of what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and unearth a lot of footage that in the past we would never have treated as reconstructable, and bring it to life as 3D." "This is something we've been playing around with a lot with Atlas. Taking old clips. I've taken a bunch of my own old captures that never worked before and put them through the system and seen a reconstruction for the first time." Fei-Fei: "The Stanford demo is underappreciated. Anywhere between 3 to 25 images, you can reconstruct that entire Stanford quad... Everything you see is generated, but according to the laws of reconstruction. And this is really magical." @BenMildenhall @drfeifei Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: youtube.com/watch?v=qn1QDD… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 2 ❤️ 28 👀 9130 📊 5 ⚡