想给智能体造训练任务?CalibForge 让多个求解器互相挑毛病来调任务难度,比人工标靠谱,Terminal-Bench 上能涨二十多个点。
CalibForge 是一种自动终端任务合成系统,利用验证过的求解器行为,通过对抗性求解器校准来修订候选任务。多求解器校准针对异构求解器池的分歧,对比求解器校准则针对强通过/弱失败关系。该系统构建了5,431个校准终端任务。在Terminal-Bench 2.0上,基于完整集合训练的模型达到32.58%和47.57%的成绩。相比基础模型,最大提升分别达24.71个百分点(Terminal-Bench 2.0)、27.68个百分点(SWE-bench Pro)和30.04个百分点(Doc2Repo)。
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.