EyeRobot 2.0 用主动注视替代腕部相机完成精细双臂操作
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
伯克利这个机器人项目挺有意思:只用一个立体相机,靠模拟人眼的主动注视,抓握遮挡场景成功率 48%,比腕部相机方案翻倍。
EyeRobot 2.0 是一个主动注视框架,用单个立体相机通过双视角对准 3D 注视点,实现细粒度双臂操作。该框架采用中央凹式处理,将更多视觉 token 分配给图像中心,并分层训练低层注视伺服策略和目标选择器,两者都用真实数据上的 RL 训练。团队在 7 个真实任务和 6 个仿真任务上收集遥操作数据,执行超过 1000 次物理和 1800 次仿真机器人试验。结果显示,仅用被动立体视觉时真实成功率从 52% 降至 27%,而 EyeRobot 2.0 在真实环境比被动立体高出 40%,在仿真中高出 20%。当被抓取物体遮挡腕部相机时,EyeRobot 2.0 成功率 48%,超过腕部相机方案的 22%。
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)