AutoResearchExam基准测试发布
This is a really neat benchmark. It's super interesting that Astra starts strong and holds the lead ...
看看AI研究代理在24小时内如何解决开放性问题,以及不同模型在长期任务中的表现差异。
Alex Dimakis团队发布了AutoResearchExam基准测试,涵盖7个研究领域,包括模型训练、数据整理和AI安全。测试中,Astra模型领先19小时后,Fable 5.1反超成为最佳模型。Qwen3.8 Max、Gemini 3.8 Flash和Grok 4.6在成本效益方面表现优异,位于帕累托前沿。
This is a really neat benchmark. It's super interesting that Astra starts strong and holds the lead ...
This is a really neat benchmark. It's super interesting that Astra starts strong and holds the lead for up to 19 hours, but Fable 5.1 catches up in the final hours. Not entirely clear why. Different models solve long-horizon problems with different strategies. It's no surprise that we see different trends. And then there is also the question of cost-performance efficiency. Qwen3.8 Max, Gemini 3.8 Flash, and Grok 4.6 are more economical for research, earning a place on the cost–performance Pareto frontier. For AI research, things get complicated not because of the duration of the task but because agents (even with the best models) tend to stagnate due to low-quality exploration. Might be due to OOD. Or simply that agents are just simply not "creative" enough (well, at least not at human levels), which the top researchers are exceptional at, from years of experience building intuition and deep expertise. The other question is: where exactly are research agents spending their compute? Are they using it efficiently? Earlier work from @intology found that coding agents mostly spend compute on hyperparameter tuning, rarely attempting algorithmic research. x.com/intology/statu… Not sure where the research stands on that today, but all of these are interesting research questions. Alex Dimakis @AlexGDimakis We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability. benchmarks.bespokelabs.ai/autoresearchex… In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through experiments and feedback. We measure both speed and quality with a combined score. Our benchmark has a unique feature: testing if agents create improvements that hold up on data they never see. We find that AI research agents often overfit as they try to improve. We see an interesting head-to-head comparison at the frontier: Astra starts the strongest and holds the lead for up to 19 hours but Fable 5.1 catches up and gets the top performing spot in the final hours. Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 all sit on the cost-performance Pareto frontier, giving strong options at lower API budgets. Anthropic's Opus and Fable retain nearly all their validation performance on hidden tests, with gaps of 1.1% and 2.9%. Astra's improvement over Sol extends to generalization too, with that gap falling from 6.9% to 1.7%. (1/n) Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 3 ❤️ 12 👀 1598 📊 4 ⚡