Robbyant发布LingBot-VLA 2.0:开源6B跨形态机器人操作模型

Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

精选理由

蚂蚁开源了6B的VLA模型LingBot-VLA 2.0,用6万小时数据训练,在GM-100上超过π0.5,做跨形态机器人操作很划算。

AI 摘要

蚂蚁集团Robbyant发布LingBot-VLA 2.0,一个6B参数的开源视觉-语言-动作模型(Apache-2.0许可证)。该模型在约6万小时数据上预训练,包括5万小时20种机器人配置的轨迹和1万小时人类视频。它使用55维规范动作空间统一不同形态,并采用无辅助损失的MoE动作专家扩展容量。在GM-100通用基准上,LingBot-VLA 2.0分别超过π0.5和上一代LingBot-VLA-1.0。

图片来源 · marktechpost
原文 · marktechpost

Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations and 10,000 hours of egocentric human video. It maps every embodiment into a single 55-dimensional canonical action space, covering arms, dexterous hands, waists, heads, and mobile bases. A token-level, auxiliary-loss-free Mixture-of-Experts action expert scales capacity without adding a load-balancing loss. Dual-query distillation from LingBot-Depth and DINO-Video adds geometric and temporal supervision for future-aware control. On the GM-100 generalist benchmark it outperforms π0.5 and LingBot-VLA-1.0 on both evaluated platforms. The post Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation appeared first on MarkTechPost .

Robbyant发布LingBot-VLA 2.0:开源6B跨形态机器人操作模型 · AI 热点