EgoLAP:用第一人称人类数据预训练机器人 VLA 模型
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
想让机器人从人类视频里学技能的可以看看这篇,EgoLAP 不硬抄人类动作,改成学语言化的动作意图,真实任务进度做到 80.1%,比其他表示高 2.3 倍。
EgoLAP 是一个 VLA 预训练框架,核心思路是不直接照搬人类第一人称视频中的底层动作,而是把动作意图表达为结构化、时间抽象化的语言动作。它通过共享的基于语言的动作思维链,让人类轨迹和机器人轨迹在同一空间中联合学习,并结合场景几何、物理和物体可供性做运动级推理。在真实世界与仿真实验中,EgoLAP 的真实任务进度达到 80.1%,相比其他动作表示有 2.3 倍的性能提升。
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.