论文精选

斯坦福研究员提出合成数据推断框架:在历史任务上校准偏差

精选理由

斯坦福Carrie的新框架,用历史任务校准合成数据,不需要真实标签,Agent Arena也在用,能解决模型和时间带来的偏差。

斯坦福博士生与Arena研究员@carrieeeeet提出一套框架,在历史任务上校准合成数据,避免依赖当前任务的真实标签。该方法覆盖社会科学调查、AI评估和Agent Arena排行榜早期信号,处理源于模型、时间和环境变化的系统偏差。核心通过任务可交换性(Task Exchangeability)从相邻任务学习,World Cup案例用于解释原理。案例还包括Bradley-Terry/AutoRater评分和Agent Leaderboard上的可操控性分数。

图片来源 · lmarena.ai
原文 · lmarena.ai

Check out this framework for inference on synthetic data through calibration on historical tasks, by Stanford PhD candidate and @arena research intern, @carrieeeeet Applied across social-science surveys, AI evaluation, and early-read signals from the Agent Arena leaderboard, this approach accounts for bias in synthetic data stemmed from models, time, and changing conditions in general — without requiring ground-truth data for the task at hand. The key idea: When the data for your task does not exist, learn from the neighboring tasks that came before it. 0:00 Intro: can you spot the real dog vs AI photo? 3:10 Synthetic data is cheap, fast, and scalable; ground-truth data is missing 4:38 The risk is non-negligible: systematic bias in synthetic data 6:52 Learn from historical tasks 9:09 Task Exchangeability: explained with a World Cup example 17:52 The method: three simple moves 23:00 Case study: social science coverage results 24:25 Case study: Bradley-Terry / AutoRater scoring 27:31 Case study: steerability score on the Agent Leaderboard 36:07 The bigger picture: from data points to tasks 39:05 Closing question: what is data, really? Your browser does not support the video tag. 🔗 View on Twitter 💬 3 🔄 1 ❤️ 10 👀 1900 📊 4 ⚡