合成增强推理的规模-权重前沿研究

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

精选理由

这篇论文教你如何用合成数据提升统计推断效果,特别适合数据稀缺场景的研究者。

AI 摘要

该研究提出了一种合成增强推理的通用框架,通过确定合成观测值的数量和权重来优化统计推断。研究人员建立了规模-权重前沿,为每个权重指定最大合成样本规模,确保所有较小规模都能达到目标任务边际覆盖率。在大型语言模型增强调查数据的实验中,该方法实现了目标覆盖率并显著缩小了置信区间。

原文 · arXiv cs.LG

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

合成增强推理的规模-权重前沿研究 · AI 热点