Braintrust研究显示,搜索功能能让不同模型表现趋同,但搜索次数过多反而会降低准确率。
Braintrust研究评估了1,329个时事问题,对比4个模型在14种条件下的表现。搜索功能将模型间差距从47.9分缩小至5.6分。检索增益随事件时间衰减,近期事件约45分,最旧事件约24分。5次以上搜索的得分反而下降19-49%。
Incredible result.. With better search the different between models becomes much much smaller!
Incredible result.. With better search the different between models becomes much much smaller! Braintrust @braintrust Imagine doing your job without ever looking anything up on the internet. That's an agent without web search. Giving agents the web changed what they can do, especially on anything recent that isn't reflected in training data. But agents don't search like people, so optimizing their performance is a new challenge. There's a lot of great research out there on search behavior, but we wanted to answer a more operational question: when should search be on, and how should you configure it? We evaluated 1,329 current events questions across 4 models and 14 conditions, comparing @youdotcom , provider built-in search, and no search. We found that: - Search reduced the gap between models from 47.9 points to 5.6 - Retrieval gain declined with event age, from ~45 points for recent events to ~24 for the oldest - Runs with 5+ searches scored 19–49%. A fifth query was associated with lower performance Read the research → braintrustdata.link/web-search-eval 🔗 View Quoted Tweet 💬 1 🔄 2 ❤️ 7 👀 1276 📊 2 ⚡