LLM检测作为干预:策略性用户行为下的下游影响

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

精选理由

这篇论文用模型和实验告诉你,加LLM检测可能适得其反,让用户用更多LLM、产出更差内容,搞学术研究的值得看看。

AI 摘要

这篇论文研究了LLM检测工具如何在用户策略性行为下影响下游指标。作者构建了一个模型,展示不完美的LLM检测器会导致反直觉结果:用户反而可能增加LLM使用量,即使减少检测特征能提升输出质量,引入检测器也可能导致更低的输出质量。实验在arXiv摘要的词频上重现了检测属性的“先升后降”模式。论文揭示了LLM检测作为干预时的失效模式,涉及LLM使用和输出质量的扭曲。

原文 · arXiv cs.AI

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.