EVOL:用知识追踪模拟器生成专家路径,免部署训练学习路径推荐策略
EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
一篇把机器人仿真训练搬到教育推荐上的论文,用模拟器造专家数据解决奖励稀疏问题,还对比了三种模仿学习策略,做 LPR 的可以看看。
EVOL 借鉴机器人领域的模拟器示范学习思路,用知识追踪模拟器通过进化搜索为每个学习者合成专家学习路径。模型采用非对称 actor-critic 结构,actor 在部署时盲规划,critic 在训练时利用模拟器的特权状态。在 ASSIST15、Junyi、EdNet 三个数据集(39-189 个概念)上,路径长度 L=5、10、20 时超越 8 个基线,涵盖启发式、序列模型、RL、图增强 RL 和 LLM 增强方法。论文还比较了 BC、AWR、DAPG 三种模仿策略,发现最终表现由进化专家质量决定,而非模仿目标本身。
EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.