斯坦福AI Lab搞了个新方法TRACE,让AI自己找短板补课。只用四分之一训练量就超过了GRPO和GEPA,还赢了大模型Codex 5.2和GLM 5。
斯坦福AI实验室提出TRACE自我改进方法,智能体通过识别自身失败背后的缺失能力并进行针对性训练来提升。TRACE训练的Qwen3.6-27B在SWE-bench Verified上达到73.2%,超过Codex 5.2和GLM 5等更大模型,同时以少于1/4的训练rollout击败GRPO和GEPA。该工作已在ICML AIWILD获得Spotlight论文。
Check out TRACE, a new self-improvement approach where the agent identifies the missing capabilities...
Check out TRACE, a new self-improvement approach where the agent identifies the missing capabilities behind its own failures and trains itself to address them. TRACE-trained Qwen3.6-27B reaches 73.2% on SWE-bench Verified, outperforming much larger models like Codex 5.2 and GLM 5, while beating GRPO and GEPA with <1/4 the training rollouts. Exciting work led by @hangoo_kang and @TarunSures41845 ! Hangoo Kang @ ICML ✈️ @hangoo_kang “TRACE: Capability-Targeted Agentic Training” got Spotlight @ ICML AIWILD 🎉 Beats direct RL, GEPA, & synthetic-agent data on SWE-Bench Verified and τ²-Bench. TRACE-Qwen3.6-27B tops GPT-5.2-Codex, GLM 5, & Claude 4.5 Sonnet on SWE-Bench. Co-led with @TarunSures41845 . Thanks to @JonSaadFalcon and our advisor @Azaliamirh . Details below 👇 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 25 👀 4899 📊 6 ⚡