AI训练使用合成用户数据
very relevant to today’s drama! and PSA for anyone working on anything sensitive.
Aidan Gomez爆料AI公司用用户合成数据训练模型,敏感工作者需注意数据安全。
Aidan Gomez透露,消费级AI工具的生产用户数据被用于生成合成数据并用于模型训练。大型实验室员工证实这一做法,尤其针对复杂数学、商业、软件和生物问题领域的数据会被过滤和加权。即使在ZDR和"不训练您的数据"政策下,衍生数据通常被排除在外,承诺仅不使用您输入的原始数据。
very relevant to today’s drama! and PSA for anyone working on anything sensitive.
very relevant to today’s drama! and PSA for anyone working on anything sensitive. Aidan Gomez @aidangomez Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees. In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn. Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game. 🔗 View Quoted Tweet 💬 2 🔄 3 ❤️ 25 👀 3186 📊 5 ⚡