模型精选

智能体研究模型 AREX 逐条核验答案,4B 微调模型胜过 35B

精选理由

有个叫 AREX 的研究智能体,答完题会自己逐条核对再补漏,BrowseComp 拿了 82.5%,微调过的 4B 小模型还打赢了 35B。

新发布的智能体研究模型与框架 AREX 会按需求逐条核验答案,保留已验证部分,只针对缺口继续检索。在 BrowseComp 上拿到 82.5%,WideSearch-en 达 82.0 F1。仅靠 harness 本身就能带来最多 10 个百分点的提升。微调后的 4B 模型在 6 个基准中的 5 个上超过了未微调的 35B 模型。

原文 · DeepLearning.AI

A new agentic research model and harness called AREX checks its answers requirement by requirement, keeps what's verified, and researches only the gaps. 📰✅ 📊 82.5% BrowseComp, 82.0 F1 WideSearch-en 📈 Harness alone: up to +10 pts 🧠 Fine tuned 4B beat untuned 35B on 5 of 6 benchmar hubs.la/Q04ztJKG0 aL #DeepLearningAI n #AIAgents e #LLMs LLMs 💬 3 🔄 0 ❤️ 4 👀 589 📊 3 ⚡