论文

论文:解析解 broker-trader 博弈中 PPO 奖励正确仍学不准

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

精选理由

这篇用有解析解的金融博弈当考卷,发现 PPO 奖励给对了也学不准,问题出在 critic 排序动作的能力上,对做 RL 金融应用的人很有参考价值。

研究把 PPO 智能体放进一个有解析解的连续时间 broker-trader 博弈,考察奖励正确时 RL 能否学到最优策略。在无噪声流动性场景下,PPO-FFNN 能接近参考动作;但加入随机未知情订单流后,PPO-FFNN 与 PPO-LSTM 的 critic 无法可靠排序邻近动作,即使监督学习证明 actor 有表示能力,potential-based reward shaping 也无稳定改善。部分信息下,基于可观测历史的 causal certainty-equivalent 控制器误差和收益都优于 PPO。将解析策略冻结后微调,交易成本减半时 PPO 只缩小了 2.22% 的差距。

原文 · arXiv cs.AI

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.