只用单GPU和精心数据设计,就能让机器人从实验室失败变身为靠谱的超市补货工,系统集成是关键。
DEED是一种系统级方法,旨在弥合VLA人形机器人的实验室与现实操作差距。该方法在超市芯片补货任务上评估,使用Unitree G1-Edu机器人和GR00T N1.6基础模型。它包含三个组件:数据高效后训练管道(控制频率对齐、数据整理、任务相关视觉高亮)、基于RECAP的体验驱动学习(文本优势前缀和视觉语言价值函数)以及潜在空间分析工具。结果显示,通过精心数据设计和单GPU后训练,原本在朴素微调下失败的政策可转变为可靠的现实系统。
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.