分布无关鲁棒轨迹优化:基于机会约束强化学习

Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning

精选理由

航天器轨迹规划团队终于有了一个分布无关的鲁棒优化方案——无需假设不确定性分布,仅需可采样,且能跨问题复用核心结构。做深空任务或火箭着陆控制的开发者可以直接参考其强化学习鲁棒化方法。

AI 摘要

本文提出一种分布无关的鲁棒轨迹优化框架,基于机会约束强化学习。不确定性通过初始条件和过程噪声表示,仅需可采样。先离线计算确定性标称轨迹,再通过强化学习鲁棒化基线,采用结构化仿射闭环修正律(前馈调整+时变反馈增益)。概率可行性通过基于rollout的上尾分位数经验保证,终端散布通过协方差可行性惩罚调节。在地球-火星转移和大气定点火箭着陆两个案例中验证,表明该方法在保持概率可行性的同时,燃料成本竞争力强,且核心随机控制结构可跨异构航天器轨迹规划问题复用。

原文 · arXiv cs.LG

Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning

This paper presents a distribution-agnostic robust trajectory-optimization framework based on chance-constrained reinforcement learning. The uncertainty is represented here through initial conditions and process noise, with the only requirement being that it can be sampled. A deterministic nominal trajectory is first computed offline, and reinforcement learning is then used only to robustify that baseline through a structured affine closed-loop correction law comprising a feedforward control adjustment and time-varying feedback gains. Probabilistic feasibility is enforced empirically through rollout-based upper-tail quantiles, while terminal dispersion is regulated through covariance-feasibility penalties. The framework is assessed on two materially different trajectory design problems. The flagship case study is a three-dimensional multi-impulse Earth-Mars transfer, where the learned policy is benchmarked against a recent robust trajectory-optimization reference under Gaussian uncertainty and then evaluated under bounded uniform uncertainty and under process disturbances not seen during training. The second case study is a stochastic atmospheric pinpoint rocket landing problem, used to assess portability to a short-horizon continuous-thrust setting with drag, mass depletion, and glide-slope constraints. The results show that the proposed framework can remain competitive in upper-tail fuel cost while preserving probabilistic feasibility, and that the same robustification scaffold can be carried across heterogeneous spacecraft trajectory planning problems without redesign of its core stochastic-control structure.