One-Shot OPD技术减少训练数据量
清华团队的新研究用1个查询替代数万数据,OPD算法其实比数据更重要。
清华大学等机构提出One-Shot OPD技术,将训练集缩减至单个查询。在数学基准上,该方法从59.1提升至68.5,达到完整数据OPD 69.8分的87%。该技术在Qwen、Llama和OLMo模型上均有效,且在代码、指令跟随和智能体工具使用任务中表现一致。
Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix it visits: the student samples its own rollouts, and Qwen3, MiMo, GLM-5, DeepSeek-V4 and Kimi K3 all pair it with SFT and RL. Yet work on OPD has almost all stood on the algorithm side, treating the training set as given—so how much of OPD's gain does the data account for? Introducing One-Shot OPD, from @TsinghuaNLP (OpenBMB member) with the University of Chinese Academy of Sciences, Northeastern University, UIUC and Johns Hopkins University. It cuts the training set to one query, and the answer is that OPD is data-overfed but algorithm-starved.
1️⃣ One query, hundreds of steps. On math it goes from 59.1 to 68.5 by step 300, against 69.8 for full-data OPD—87% of its gain. It holds across code, instruction following and agentic tool use, and across Qwen, Llama and OLMo; a query the student never solves works about as well as one it always solves. 2️⃣ States, not questions. A prefix is a state where the teacher gives a target distribution, so 64 rollouts per step already yield tens of thousands of supervised positions. One query reaches 71.5% state coverage of what full-data OPD visits; 16 diverse queries reach 98.9% and match full data—extra queries buy new states, not new questions. 3️⃣Alignment slows, not the data. The student keeps improving, but each update absorbs less of what is left, which is why a run takes hundreds of steps rather than tens. This decline hardly depends on training-set size: on 1, 4, 16 and all 17k queries, alignment slowed at a similar pace. What limits a run is not how much data it gets, but how fast the student absorbs it.
📄 Paper: https://t.co/KbK6QhYA4t 💻 Code: https://t.co/dlaFpz9hNJ #AI #THUNLP #OpenBMB #LLM #PostTraining #Distillation #OpenSource