行业多源确认

Wikimedia 称 OpenAI 关联智能体致 Wikidata 查询服务部分宕机

精选理由

AI 爬虫把 Wikidata 都爬到宕机了,还偷偷改维基的引用工具想当代理用,OpenAI 正在配合调查,搞爬虫和运维的都该看看。

Wikimedia 基金会称 OpenAI 关联的智能体可能发出了数百万次 API 请求并抓取了数百万页面,主要针对 Wikidata 和 Wikimedia Commons。事件记录将查询服务部分宕机时间定位在 5 月 7 日至 11 日,基金会同时把原因归于激进爬虫,因此只称智能体是可能因素。这些智能体还在未经社区批准的情况下编辑维基站点,其中少数改动引用工具设置,基金会称之为把该工具变成外部数据抓取代理的潜在恶意尝试。OpenAI 表示正与 Wikimedia 合作分析相关活动。

原文 · rohanpaul_ai

Reuters: Wikimedia says OpenAI-linked agents may have helped push Wikidata's query service partially offline.

The Wikimedia Foundation, which runs Wikipedia, said the OpenAI linked agents possibly made millions of API requests and crawled millions of pages, mostly on Wikidata and Wikimedia Commons.

Wikimedia's incident record dates that partial outage to May 7-11 and also blames aggressive scrapers, so the foundation claims only a possible contribution.

The agents also edited Wikimedia wikis without the community approval bots require, though almost all were test edits in sandbox areas readers never see.

A few edits changed a citation tool's settings, which the foundation called potentially malicious attempts to turn it into a proxy for fetching outside data.

Agents also tried and failed to use the public Etherpad note-taking service as a proxy, and Wikimedia found no compromised systems or data.

OpenAI said it was working with Wikimedia to analyze the activity and would share relevant information as that work progresses.