后验匹配基准:从点估计到生成式逆求解器的后验评估
PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
这是关于如何评估生成式逆求解器后验分布准确性的新方法,比单纯看点估计更全面,对研究科学计算中的生成式模型很有用。
我们介绍 PosteriorBench 基准,用于评估生成式逆求解器的分布准确性。它通过构建高保真参考后验,评估 Darcy 流反演、泊松源恢复等四个物理逆问题。该基准使用后验均值误差、最大均值差异等五项指标,评估点估计精度、边缘不确定性、分布对齐和全局频率保真度。
PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using computationally heavy but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. We pair these references with a five-metric posterior evaluation suite: posterior-mean error, posterior-standard-deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power-spectrum error. These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching and uncertainty quantification. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness, and guidance weights and generation noise are key to posterior-variance calibration.