Locus后训练Qwen3超越官方,登顶PostTrainBench

This is a wild result. Locus, the automated research system from @intology, post-trained Qwen3 base...

精选理由

Locus这个自动化系统后训练的模型居然干过了人类调的Qwen3 1.7B,在PostTrainBench上拿了第一,还上了Kaggle奖金赛前排。

AI 摘要

Intology的自动化研究系统Locus对Qwen3基础模型进行后训练,在PostTrainBench上达到SOTA。其训练出的模型超越官方人类调优的Qwen3 1.7B Instruct版本。PostTrainBench用10个H100小时评估代理后训练能力,扩展版PostTrainBench+将预算提升至数千H100小时。在此预算下,Locus后训练的模型集体优于官方Qwen3 1.7B。此外,Locus在Kaggle所有活跃奖金竞赛运行16天后,平均排名第4。

原文 · elvis

This is a wild result. Locus, the automated research system from @intology, post-trained Qwen3 base...

This is a wild result. Locus, the automated research system from @intology , post-trained Qwen3 base models that beat the official human-tuned Qwen3 1.7B Instruct release. SoTA on PostTrainBench! Intology @intology The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇 PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours. We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model. In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants. 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 4 👀 1258 📊 1 ⚡