8月25日
10:40
10:40官方账号arXiv cs.AI@Sangoh Lee, Sangwoo Mo, Wook-Shin Han
We propose Intention Distillation (INDI) for Vision-Language-Action (VLA) models, improving GR00T-N1.7 from 64.3% to 84.7% on SimplerEnv-Bridge and from 64.1% to 70.3% on RoboCasa Kitchen. In real-world tasks, INDI boosts average success from 62.0% to 68.7%, with up to 12.0 pp gains on longer-horizon tasks.
推荐理由:INDI significantly enhances VLA models, outperforming GR00T-N1.7 by 20.4% on SimplerEnv-Bridge and 6.2% on RoboCasa Kitchen, making it a notable improvement over traditional behavior cloning approaches.