DreamQAS用强化学习省真实VQE调用,BeH2少10.6倍,五任务里四个误差最低,做量子计算可看看。
DreamQAS提出一个基于模型的强化学习框架,保留电路构建的确定性动态,只学习昂贵的VQE后反馈。在15000集预算和冻结评估下,它在五个分子任务中的四个取得最低平均冻结策略能量误差,另一个排第二。在两种方法所有种子都能达到的精细误差目标上,它用1.6倍到2.0倍更少的真实VQE调用完成四个任务,在BeH2-8q上少用10.6倍。反事实动作排序效用五个任务平均提高0.346,95%置信区间为[0.185, 0.507]。
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. We introduce DreamQAS, a model-based RL framework that preserves these exact circuit dynamics and learns only the expensive post-VQE feedback. A recurrent randomized-prior ensemble predicts an oracle-free score relative to an empirical energy frontier and supports multi-step imagined policy learning over explicit legal circuits. Ranking-based activation, uncertainty-aware pessimism and truncation, and selective real-VQE verification form a reliability-controlled learning loop. Under a common 15,000-episode budget and frozen evaluation for the RL methods, DreamQAS has the lowest mean frozen-policy energy error on four of five molecular tasks and the second-lowest on one. At fine-error targets reached by all seeds of both methods, it uses 1.6x to 2.0x fewer real VQE calls on four tasks and 10.6x fewer on BeH2-8q. Counterfactual action-ranking utility increases across all five tasks, with a mean increase of 0.346 and a 95 percent confidence interval of [0.185, 0.507], while direct greedy and beam use of the same model does not recover the gains of imagined policy learning. Ensemble disagreement also improves risk-coverage over random rejection on all three probed tasks. These results establish a world-model design for QAS whose value lies in decision-useful feedback rather than exact energy prediction.