论文

时空数据预测驱动推断:短标签时间序列也能给出可靠置信区间

Prediction-powered inference for time series across space

精选理由

arXiv 上的统计方法论文,解决作物产量这类短标签、长协变量的时空预测问题,把 PPI 和 HAC 结合起来修正填补偏差,做时空统计的可以看看。

这篇 arXiv 论文(编号 2610.08715)针对时空预测场景:标签(如作物产量)只在较短的近期时间段内有观测,而协变量(如气象数据)在更长时间段上可用。直接用机器学习填补缺失标签会带来明显偏差,传统预测驱动推断(PPI)的 i.i.d. 假设在时间依赖下失效。作者提出结合 HAC(异方差与自相关一致)过程的方法,在标签被部分填补的情况下仍能给出可靠的点估计和置信区间,并在实验中优于自然基线方法。

原文 · arXiv cs.LG

Prediction-powered inference for time series across space

The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.