连续时间强化学习在Hawkes跳跃扩散控制中的应用研究

Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

精选理由

这研究发布了针对Hawkes过程的连续时间强化学习新方法,用近似技术优化求解,和传统离散方法比效果不同,值得了解。

AI 摘要

该研究针对带记忆的Hawkes驱动随机微分方程,提出连续时间强化学习算法Hawkes-CT DDPG;采用有限维马尔可夫化近似方法,将问题转化为可处理的Markovian形式;对比简单指数、Erlang、幂律三类核的离散方法,验证新方法的性能优势。

原文 · arXiv cs.LG

Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation of the problem, called Hawkes-CT DDPG. We propose a model-free algorithm to solve the non-Markovian Hawkes-driven optimization by observing only the event times of the process, the realization of the solution to the SDE, and a chosen set of decay filters, while the Hawkes kernel coefficients remain unknown. We compare our continuous time reinforcement learning Hawkes-CT DDPG method with discrete time reinforcement learning techniques under three different types of kernels: simple exponential, Erlang, and power-law kernels.