论文精选
Perplexity:GPT-OSS-120B 代理在 BrowseComp+ 上达 64.0%
精选理由
Perplexity 实测 GPT-OSS-120B 做搜索代理,BrowseComp+ 拿了 64.0%,比 ColBERT 高 4.9 分还搜得更少,搞 RAG 的可以看看。
Perplexity 测试了 agentic search 场景下的 GPT-OSS-120B 代理。该代理配合 9B 模型在 BrowseComp+ 上答对 64.0% 的问题,比第二名 ColBERT 高出 4.9 个百分点。同时它的搜索次数少于所有对比基线,说明检索效率也更高。
原文 · Perplexity
In agentic search, a GPT-OSS-120B agent using the 9B model answers 64.0% of BrowseComp+ questions correctly, 4.9 points ahead of the next ColBERT model, with fewer searches than any baseline. 💬 1 🔄 0 ❤️ 4 👀 498 📊 1 ⚡