论文

分层强化学习让欠驱动双足机器人动态避障行走

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

精选理由

双足机器人边走边躲障碍还容易摔,这篇论文用高低两层策略分工解决,动态场景成功率比规划器方案高了20个百分点。

一篇 arXiv 论文提出分层强化学习(HRL)框架,用于每腿四个驱动关节、无髋踝侧摆的欠驱动双足机器人。高层策略每 10 个控制步输出一次机体速度指令,底层速度条件策略通过 PD 关节目标跟踪指令,两层均用 Soft Actor-Critic 联合训练。在 PyBullet 随机环境中每方法各做 100 次评估,该方法在静态和动态避障中分别达到 98.0% 和 88.0% 的到达率,而 SAC+A*、SAC+RRT*、SAC+APF 三种规划器基线最高只有 78.0% 和 68.0%。路径长度与 A* 参考差距在 4% 以内,消融实验验证了各观测通道和奖励项的贡献。

原文 · arXiv cs.AI

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, ω_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.