长期测量:人类与AI交互的纵向研究视角

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

精选理由

一篇讲怎么测AI长期影响的论文,用社会科学的方法补NLP的短板。适合做AI安全和行为研究的人读。

AI 摘要

这篇arXiv论文认为,语言模型拟人化且融入日常使用,可能引发认知、发展和社会情感层面的长期变化。作者主张NLP应转向长期测量用户行为变化,而非局限静态文本评估。论文参考社会科学中纵向数据测量方法,并与NLP计算手段结合。该框架用于在线识别问题行为,并纳入对齐流程以缓解长期风险。

原文 · arXiv cs.AI

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.