这篇论文提出了一种改进的源分布估计方法,通过期望最大化解决了现有方法在参数空间不准确区域的问题。
本文提出了一种通过期望最大化解决源分布估计(SDE)问题的新方法。该方法在E步骤使用当前源估计的仿真训练后验,在M步骤将源拟合到观测数据的后验平均值。研究者在三个基准任务上评估了两种参数化方法,在Lotka-Volterra任务中,现有基线方法C2ST值均高于0.96,而新方法在四种初始先验设置中有三种达到0.64-0.68。
Source Distribution Estimation by Posterior Averaging
Simulation-based science often requires a distribution over simulator parameters whose push-forward reproduces a set of real observations: this is the source distribution estimation (SDE) problem. Existing methods fit the source against a likelihood surrogate trained once from a fixed proposal prior. Their objective is therefore stated only in terms of the surrogate instead of the true simulator, which may fail for inaccurate areas in parameter space where the surrogate was never trained. We instead solve SDE by expectation maximization: an E-step trains an amortized posterior on fresh simulations from the current source estimate, and an M-step refits the source to the average of that posterior over the observed data. We give two parameterizations, (1) separate source and posterior flows and (2) a single shared conditional flow. We evaluate our method on three benchmark tasks under both broad and misspecified initial priors. Both improve on existing fixed surrogate approaches and on iterated variants of each, most clearly on Lotka--Volterra, where no baseline falls below 0.96 data-space C2ST while our methods reach 0.64-0.68 in three of four initial-prior settings.