AI模型精选

递归合成框架RST:低成本生成长时程终端任务

Recursive Synthesis for Long-Horizon Terminal Tasks

精选理由

想低成本造高质量长任务数据?RST用递归验证把成本压到5分钱一个,还让模型成绩涨了10分,值得看看。

AI 摘要

RST是一种递归验证的合成框架,能从已验证种子任务出发扩展参考解、重对齐验证器与指令,并在新沙盒中验证。经过15轮递归,RST以约0.05美元/任务的成本生成了37,484个终端任务,中位参考解从67行增至374行,执行的命令数从40增至244。DeepSeek-V4-Pro的pass@4从第1轮的90%降至第15轮的2.5%,任务难度持续攀升。用Qwen3.5轨迹微调,Qwen3.5-27B在Terminal-Bench 2、Hard和Long-Horizon上最多提升10分;PPO训练后相对增益为20.0%、41.2%和21.9%。15轮后合成产率和验证率保持稳定,表明流程可继续扩展。

原文 · arXiv: DeepSeek

Recursive Synthesis for Long-Horizon Terminal Tasks

High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly \$0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90\% at $R_1$ to 2.5\% at $R_{15}$. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench~2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44\%, 32.00\%, and 22.07\% on the three benchmarks, corresponding to relative gains of 20.0\%, 41.2\%, and 21.9\% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.