Kuafu系统将符号规划与数据驱动执行统一,无需演示或手动密集奖励,在机器人任务中表现优异。
研究人员提出SUN(Semantically UNified)程序,将几何和接触关系一次性定义并编译成MPC成本、满足谓词、RL奖励等。Kuafu系统通过大型视觉语言系统自动合成SUN程序,使用MPC筛选可行性,并在训练阶段保留语义。在九项任务中,Kuafu达到82.03%的宏观成功率,显著优于稀疏奖励(35.67%)和Stage-BC(24.75%)基线。
SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies
Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what control verified.We introduce Semantically UNified (SUN) Programs, typed executables where geometric and contact relations are defined once and compiled into aligned Model Predictive Control (MPC) costs, satisfaction predicates, RL rewards, transition guards, and diagnostics. Our system, Kuafu, driven by large vision language systems, automatically synthesizes SUN Programs from language and scene semantics, screens feasibility via MPC, and retains semantics while training stage-conditioned policies. Across nine tasks, Kuafu achieves 82.03% macro-success, outperforming sparse-reward (35.67%) and Stage-BC (24.75%) baselines. At 8192-way scale, it generates 10.57x the successful trajectory time per hour of human teleoperation. With 500 trajectories per task, Kuafu data trains DP3 policies to 46.0% simulation success (vs. 22.4% for alternatives) and 34.7% on physical Franka and Kinova robots. These results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.