World Labs的Atlas能从单张图片生成3D场景,视频生成长达1分钟,效果超过专业3D重建模型。
李飞飞的World Labs发布了新世界模型Atlas,能从单张图片生成完整3D场景。Atlas支持三大任务:生成最长1分钟1440p的视频,重建3D场景效果超过专业模型,以及时空仿真。模型能将多张图像整合到共享空间上下文中,理解时间演变规律。
李飞飞的 World Labs 发布了他们新的世界模型 Atlas,这个看起来比上一代厉害多了。 通过一张图片就可以创建一个非常完整的 3D 空间场景,你可以在里面随意运动,它也能理解时间的演变规律...
李飞飞的 World Labs 发布了他们新的世界模型 Atlas,这个看起来比上一代厉害多了。 通过一张图片就可以创建一个非常完整的 3D 空间场景,你可以在里面随意运动,它也能理解时间的演变规律。 它能将一张或多张输入图像整合进一个共享的空间上下文中,让每张图像都锚定在三维空间中的具体位置,再以此继续生成后续内容,既能保证当前场景的一致性,又能想象出没有看到的内容。 它主要支持三大任务: 生成图像和视频:能从一张或多张参考图出发,以像素级精确的相机控制生成不同角度的图像或完整视频,最长可输出 1 分钟 1440p 的视频 3D 场景重建:用一张或几十张图片重建真实的 3D 场景,直接生成点云、3D 高斯泼溅(3D Gaussian Splatting)等三维结果(图片越多还原越真实),效果甚至超过了专门针对 3D 高斯泼溅优化的模型 时空仿真:能把几台普通手机拍摄的视频重新取景成“子弹时间”效果,还可以为机器人训练生成测试数据 Your browser does not support the video tag. 🔗 View on Twitter Fei-Fei Li @drfeifei I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️ 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 19 👀 4420 📊 4 ⚡