论文精选73°

Nvidia 发布 Cosmos 3:统一语言、图像、视频、音频和动作的物理 AI 世界模型

Nvidia's Cosmos 3: 1 model that can understand, si…

精选理由

Nvidia 让机器人学会动作语言

AI 摘要

Nvidia 推出 Cosmos 3,一个能够理解、模拟和行动于多种物理 AI 任务的统一模型。它将动作视为世界的一等语言,把语言、图像、视频、音频和动作整合到一个共享系统中。该模型通过动作标记设计,让机器人能连接所见与可能发生的事,并决定下一步行动。论文显示,Cosmos 3 可基于视频推断动作,或与未来场景一同生成动作,从而解决机器人抓取、滑动等物理交互问题。

原文 · rohanpaul_ai

Nvidia's Cosmos 3: 1 model that can understand, si…

Nvidia's Cosmos 3: 1 model that can understand, simulate, and act across many physical AI tasks.

It treats action as a first-class language of the world.

Most AI models look at reality from the outside: images become captions, videos become descriptions, and motion becomes something to label after the fact.

Cosmos 3 tries to collapse that distance by putting language, image, video, audio, and action into one shared system, so a robot can connect what it sees with what might happen next and what it should do.

A home robot cannot simply recognize a plate, a table, and a human instruction, because the useful question is what changes when it moves, grasps, slips, bumps, or waits.

That is why the paper’s action-token design matters: it turns movement into something the model can condition on, infer from video, or generate alongside a future scene.

----

Link – arxiv. org/abs/2606.02800

Title: "Cosmos 3: Omnimodal World Models for Physical AI"