模型78°

Fei-Fei Li推出Atlas 3D世界模型

>be dr. fei-fei li >born in beijing, raised in chengdu >dad moves to new jersey, follow at age 16 >s...

精选理由

Fei-Fei Li创立的World Labs推出Atlas模型,只需3张照片就能创建3D世界,比传统方法快50-100倍。

Fei-Fei Li在2024年创立World Labs,2026年发布Atlas模型。该模型可将少量照片转化为完整3D世界,将3D空间数字化所需照片数量从100-300张减少至仅需3张。Atlas基于新视角预测技术,统一了像素生成和像素重建两个计算机视觉领域分离了半个世纪的问题。

原文 · a16z

>be dr. fei-fei li >born in beijing, raised in chengdu >dad moves to new jersey, follow at age 16 >s...

>be dr. fei-fei li >born in beijing, raised in chengdu >dad moves to new jersey, follow at age 16 >speak almost no english >parents open a dry cleaner >work the counter on weekends >get into princeton on a full scholarship >study physics, because why not >caltech phd >everyone in AI is obsessed with algorithms >bet on data instead >"this will never scale" they say >do it anyway >get 14 million images hand-labeled >a team from toronto trains a convnet on it in 2012 >error rate falls off a cliff >deep learning era begins >become stanford professor >become google cloud chief scientist >found stanford HAI >press starts calling you the godmother of AI >2024, most of the field is racing on language >you say the next frontier is spatial intelligence >co-found World Labs in 2024 >ship Atlas in 2026, a model that turns a few photos into a whole 3D world a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: youtube.com/watch?v=qn1QDD… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 10 🔄 6 ❤️ 146 👀 15034 📊 19 ⚡