LiLa-WAM用一张24GB显卡就能训练,在RoboTwin 2.0上50个任务做到90%以上成功率,还不用语言标注任务,做机器人操控的值得看看。
LiLa-WAM在紧凑潜在空间中推理未来场景,可在单个24GB GPU上端到端训练,避免像素空间法的高开销。其核心是联合塑造的未来状态预测与动作生成机制,并提出了语言无关的视觉转换标记(VTT)作为任务表征。在RoboTwin 2.0的50个任务中取得90.48%成功率,且同时在LIBERO和真实机器人任务上验证了有效性。
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.