PUN协议:用于评估以人为中心的LLM的未知姓名

No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

精选理由

想了解如何评估LLM在隐私和偏见方面的表现?这篇论文提出了一个新方法,值得一读。

AI 摘要

研究人员提出PUN协议,用于构建和验证未知姓名,以评估LLM在事实性、隐私泄露、偏见和弃权方面的表现。研究发现,接受的名字更符合姓名特征,而参与者仅3%的情况下能恢复个人信息。发布300个姓名及对比控制名单。

原文 · arXiv cs.AI

No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.