斯坦福Carrie的新框架,用历史任务校准合成数据,不需要真实标签,Agent Arena也在用,能解决模型和时间带来的偏差。
斯坦福博士生与Arena研究员@carrieeeeet提出一套框架,在历史任务上校准合成数据,避免依赖当前任务的真实标签。该方法覆盖社会科学调查、AI评估和Agent Arena排行榜早期信号,处理源于模型、时间和环境变化的系统偏差。核心通过任务可交换性(Task Exchangeability)从相邻任务学习,World Cup案例用于解释原理。案例还包括Bradley-Terry/AutoRater评分和Agent Leaderboard上的可操控性分数。
Check out this framework for inference on synthetic data through calibration on historical tasks, by...
Check out this framework for inference on synthetic data through calibration on historical tasks, by Stanford PhD candidate and @arena research intern, @carrieeeeet Applied across social-science surveys, AI evaluation, and early-read signals from the Agent Arena leaderboard, this approach accounts for bias in synthetic data stemmed from models, time, and changing conditions in general — without requiring ground-truth data for the task at hand. The key idea: When the data for your task does not exist, learn from the neighboring tasks that came before it. 0:00 Intro: can you spot the real dog vs AI photo? 3:10 Synthetic data is cheap, fast, and scalable; ground-truth data is missing 4:38 The risk is non-negligible: systematic bias in synthetic data 6:52 Learn from historical tasks 9:09 Task Exchangeability: explained with a World Cup example 17:52 The method: three simple moves 23:00 Case study: social science coverage results 24:25 Case study: Bradley-Terry / AutoRater scoring 27:31 Case study: steerability score on the Agent Leaderboard 36:07 The bigger picture: from data points to tasks 39:05 Closing question: what is data, really? Your browser does not support the video tag. 🔗 View on Twitter 💬 3 🔄 1 ❤️ 10 👀 1900 📊 4 ⚡