论文多源确认

研究:82% 的 LLM 应用在模型下线后才迁移

When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications

精选理由

有人挖了两万多个 GitHub 提交,发现八成应用都是等模型下线报错了才迁移,数据里还有迁移成本的具体行数。

arXiv 论文分析了 22,555 个 GitHub 提交,覆盖 17,703 个非 fork 仓库(2024-2026),其中 5,139 个匹配到 OpenAI、Anthropic、Google 的官方模型退役事件。估计 82%(95% CI 79-84)的迁移发生在模型停机之后,即应用已经开始报错。迁移节奏与厂商通知期强相关:Anthropic 的 60-114 天通知对应 89% 事后迁移,而 OpenAI 一年期的 Assistants API 通知仅 13%。94% 的应用硬编码模型标识符,迁移工作量中位数从纯提示应用的 6 行代码到微调应用的近 700 行,仅 8% 的迁移更换了供应商。

原文 · arXiv: Anthropic

When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications

Applications built on commercial large language model (LLM) APIs depend on model versions that providers retire on their own schedule, with notice periods ranging from one year to two weeks. We ask what actually happens to applications when a model is retired. We mine GitHub for commits that migrate away from officially deprecated models and endpoints of OpenAI, Anthropic, and Google, matching each commit to the provider's published announcement and shutdown dates. From 22,555 commits in 17,703 non-fork repositories (2024-2026), 5,139 are matched to an official event; two independent coders validated a stratified sample of 300 (kappa = 0.89-0.95), and we reweight all estimates by their labels. We find that an estimated 82% (95% CI 79-84) of migrations away from retired models were committed after the shutdown date - after the application had started failing - regardless of repository popularity, prior retirement experience, or the presence of a provider-abstraction layer. The share tracks the provider's notice policy: 89% for Anthropic's 60-114-day notices versus 13% for OpenAI's one-year Assistants API notice, and each e-fold increase in notice length reduces the odds of post-shutdown migration by about three quarters. Model identifiers are hard-coded in 94% of migrating applications, migration effort scales from a median of 6 added lines for prompt-only applications to nearly 700 for fine-tuned ones, and only 8% of migrations switch provider. We release the dataset and pipeline and discuss implications for deprecation policy, dependency-risk assessment of LLM products, and tooling.