AI搜索代理失败不在搜索,而在模糊查询时没问对问题

AI search agents don't fail at searching, they fail at asking the right questions when queries get ambiguous

精选理由

DiscoBench告诉你,AI搜索代理的问题不是搜不到,而是不敢问——模糊查询时卡住比瞎猜还差。

AI 摘要

新基准DiscoBench指出,AI搜索代理在多步研究中的真正短板不是搜索能力,而是当查询模糊时不会向用户追问澄清。测试显示,反复搜索但不追问的模型准确率仅51.9%,甚至低于直接猜测。表现最好的模型整体准确率也仅为43%。当移除查询中的歧义后,模型准确率最高可提升40个百分点。

原文 · Decoder

AI search agents don't fail at searching, they fail at asking the right questions when queries get ambiguous

AI search agents rarely fail at multi-step research because of the search itself. Their real problem is not asking the user for clarification when queries are ambiguous. A new benchmark called DiscoBench shows that models searching repeatedly instead of asking follow-up questions actually perform worse, at 51.9 percent, than those that just guess. Even the best model only hits 43 percent overall accuracy. When ambiguity is removed from the queries, accuracy jumps by up to 40 points. The article AI search agents don't fail at searching, they fail at asking the right questions when queries get ambiguous appeared first on The Decoder .