Ben Thompson提案:训练数据应属合理使用,禁止服务条款阻止蒸馏

Who’s Afraid of Chinese Models?

精选理由

Ben Thompson提出了一个实际解决蒸馏禁令与训练数据版权矛盾的方案,同时还能帮美国开源模型对抗中国对手,值得一看。

AI 摘要

Ben Thompson提议美国立法将AI训练数据收集明确为合理使用,并禁止美国公司通过服务条款阻止模型蒸馏。他指出蒸馏仅是通过API查询模型,几乎无法禁止,因此应转向支持创新。该提案同时有助于美国开源模型与中国模型竞争。文中还提及阿里巴巴发布Qwen 3.8 Max开源权重,可能受习近平关于鼓励开源合作的讲话影响。

原文 · Simon Willison’s Weblog

Who’s Afraid of Chinese Models?

Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts: The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else. Ben also theorizes that Alibaba's decision to release Qwen 3.8 Max as open weights - a reversal from their decision not to release Qwen 3.7 Max in May - may have been influenced by a recent speech by Xi Jinping, who said: We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing. Tags: ai , generative-ai , llms , training-data , qwen , ai-ethics , ai-in-china