这篇论文提出了一种新方法,能精确定位哪些观测值导致样本偏离参考分布,在LHC奥运会上表现出色。
研究人员开发了一种新框架,通过为每个观测值分配其在随机统计上下文中的条件或边际贡献,来解决样本偏离参考分布时的定位问题。该方法将重采样诊断和数据估值与投影理论和事件级异常检测联系起来。在对称统计情况下,固定大小替换等价于中心条件定位;对于U统计量,加法得分等于第一个Hoeffding/Hájek贡献;对于平滑分布泛函,它主要与影响函数相关;对于无偏已知背景MMD,它精确简化为MMD见证。在LHC奥运会异常检测基准上,成对估计器以预测的1/(Rm^2)缩放率收敛到直接经验MMD见证,在m=1000和R=5x10^6时达到0.9993的相关性,AUC基本相同。
Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection
Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each observation its conditional or marginal contribution across random statistical contexts. This connects resampling diagnostics and data valuation to projection theory and event-level anomaly detection. For symmetric statistics, fixed-size replacement is exactly equivalent to centered conditional localization. For U-statistics, the addition score equals the first Hoeffding/Hájek contribution; for smooth distributional functionals it is related at leading order to the influence function; and for unbiased known-background MMD it reduces exactly to the MMD witness. This viewpoint also yields more efficient estimators. Matched-context subtraction removes fluctuations unrelated to the observation, while for pairwise MMD the event-containing terms give a simple localizer. On the LHC Olympics anomaly-detection benchmark, the pair estimator converges to the direct empirical MMD witness with the predicted 1/(Rm^2) scaling, where m is batch size and R the number of batches. At m=1000 and R=5x106 it reaches correlation 0.9993 with essentially identical AUC. We also ask when context contains information beyond an event's own features. In a shared-latent toy model, the full single-event signal and background distributions are identical by construction, forcing isolated-event AUC=0.5. Discriminating information survives only in cross-event dependence induced by the shared latent parameter; the ensemble recovers this information, whereas an independent-latent control does not. This separates two roles of context: efficient localization of a global discrepancy and genuinely additional class information when the alternative contains shared structure.