Holonomy-Cover决策过程的稳定商最小马尔可夫化

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

精选理由

这篇论文给POMDP找了个最省记忆的压缩方法,稳定商只用3个记忆状态就能完美还原配对顺序,比传统方法省很多。

AI 摘要

该研究针对Holonomy-Cover决策过程(一类结构化的POMDP)构建稳定商,即保持一步奖励和商后继的最粗观测-wise抽象。理论证明当前观测与稳定类构成精确有限马尔可夫状态,且在可到达性和成对决策分离条件下,任何有限记忆控制器无法使用更少的记忆符号。在可重置诊断下,近原型类推理的误差呈指数衰减。实验将原始状态压缩至商状态,仅用三个决策记忆状态即实现完美配对顺序准确率,优于非商基线。

原文 · arXiv cs.LG

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In particular, we construct the stable quotient, the coarsest observation-wise abstraction preserving one-step rewards and quotient successors, and prove that the pair of the current observation and stable class forms an exact finite Markov state. When the current class is correctly initialized, exact class tracking requires exactly the minimal memory symbols, in the sense that under reachability and pairwise decision separation at a maximizing observation, no arbitrary finite-memory controller can use fewer. Under resettable diagnostics, nearest-prototype class inference has exponentially decaying error, and a calibrate-then-restart reduction transfers finite-MDP guarantees to the recovered state. The results enable \emph{Holonomy Memory Reinforcement Learning}. It represents memory by the current stable class, updates it through ordered edge transports, identifies local class coordinates when diagnostics are available, and applies a standard finite-MDP RL backbone after synchronization. Experiments recover an exact compression from raw states to quotient states and achieve perfect paired-order accuracy with three decision-time memory states, matching the quotient oracle and outperforming the non-oracle baselines.