这篇论文用数学模型解释了为什么大家用LLM写东西越来越像,还分析了个性化模型能不能救回多样性,适合关心AI对语言影响的读者。
该论文提出了语言单一文化概念,指用户依赖相同的大语言模型(LLM)进行写作和修改,导致语言形式趋同。作者构建了数学框架,将用户和LLM表示为语言特征分布,并分析了三种交互机制:固定共享模型、递归更新共享模型和个性化模型。研究发现,共享模型会推动用户趋向共同规范,递归反馈改变共享规范但不减少个体间差异,而个性化模型能保留语言多样性。模拟显示,不同机制对长期多样性的影响不同。
Linguistic Monoculture in LLM-Assisted Language Use
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.