论文

可扩展扩散SBI处理模拟器误指定下的组合推理

Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification

精选理由

这篇论文提出了一种新方法,能处理模拟器与数据不匹配的情况,在生物信号模型中表现优异。

研究人员开发了基于扩散的推理方法,处理模拟器与观测数据不匹配的情况。该方法通过Hierarchical Blockwise Diffusion Sampling (HBDS)技术,使用单个预训练模型推断共享参数和组特定潜在状态。研究在Gaussian和Simple Likelihood, Complex Posterior基准上测试了组合采样,并在940个细胞测量数据中验证了方法的有效性。

原文 · arXiv cs.LG

Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification

Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions $ξ$. We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce Hierarchical Blockwise Diffusion Sampling (HBDS), which infers shared parameters and group-specific latent states using a single pretrained model, with the hierarchy specified only at sampling time. Together, these methods support variable observation sets and groupings without retraining. To address misspecification, we introduce path-regularized fine-tuning that adapts the learned likelihood to observations and transfers corrections to posterior inference. Using Girsanov's theorem, we quantify path divergence between pretrained and fine-tuned models across experimental designs and interpret it alongside predictive errors to distinguish candidate misspecification correction from unnecessary adaptation. We evaluate compositional sampling on exact-score Gaussian and Simple Likelihood, Complex Posterior benchmarks, HBDS with analytic and learned scores on a controlled hierarchical model, and fine-tuning and localization on a separate analytic model with known design-dependent discrepancy. Finally, we apply the framework to 940 measurements across four cell lines in a mechanistic Bone Morphogenetic Protein signaling model, where fine-tuning improves posterior-predictive accuracy relative to the pretrained model and shifts posterior marginals toward the least-squares reference while retaining spread.