论文

研究:跨数据来源的群体画像能否提升模拟问卷对齐效果

Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes

精选理由

想用 LLM 模拟问卷调研的可以看看这篇:跨域数据造的画像基本没用,对上目标人群才有效。

arXiv 论文 2610.05828 研究人口群体层面的问卷模拟:用匿名的公开行为数据诱导 persona,再让智能体模拟特定人群的问卷回答。实验发现,跨域来源诱导的 persona 往往不如只用基础人口属性信息的模拟,主要原因是人群分布不匹配。但当 persona 能准确对应目标人群时,对齐度明显提升。来自目标领域问卷数据的 persona 随着问卷历史数据增多,在未见过的题目上泛化效果更好。

原文 · arXiv cs.AI

Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes

However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.