OpenAI Web Search 登上 Artificial Analysis 搜索榜:得分 74 排第 5
OpenAI 把搜索直接内置进 API 了,Artificial Analysis 实测便宜过 Perplexity,但多跳检索还打不过 Octen,做联网应用前先看看这份数据。
Artificial Analysis 搜索索引显示,OpenAI Web Search 得分 74,排在 Perplexity、Octen、Parallel、Brave 之后位列第 5。与不联网的 GPT-5.6 Luna(33 分)相比提升了 41 分。单价约 $0.05/任务,低于 Parallel($0.06)和 Perplexity($0.07),但比 Octen($0.024)贵一倍且分数低 3 分。其事实问答基准 AA-Omniscience 准确率 72%,排名第 3;多跳检索基准 BrowseComp 为 73.5%,仅排第 13。该方案在单次 Responses API 调用内完成整个搜索循环,每任务消耗约 4 万输入 tokens,远低于 Perplexity(low)的 12.5 万。
OpenAI Web Search debuts on the Artificial Analysis Search Index at 74, the 5th best provider behind Perplexity, Octen, Parallel and Brave
@OpenAI Web Search is the first integrated first-party search tool on our board. Instead of an agent calling a Search API in our Stirrup harness, GPT-5.6 Luna (medium reasoning) calls OpenAI's built-in web_search tool, and OpenAI runs the whole search loop inside one Responses API call. All setups use the same underlying model, GPT-5.6 Luna (medium reasoning).
At ~$0.05 per task, OpenAI Web Search costs less than Parallel (advanced, $0.06) and Perplexity (medium, $0.07), but about twice as much as Octen ($0.024), which scores 3 points higher.
Key benchmarking results for OpenAI Web Search (medium search context):
➤ 5th best provider on the Artificial Analysis Search Index: OpenAI Web Search scores 74, a 41-point lift over the same model with no search (33). It places 7th of 26 products, level with You . com (highlights), Nimble (standard) and Exa (auto) at 74, behind the three Perplexity variants (77 to 80), Octen (77), Parallel (advanced) and Brave (LLM context) at 75
➤ OpenAI Web Search performs best on AA-Omniscience, our proprietary factual answer benchmark, achieving 72% accuracy. This ranks 3rd of 26 search provider variants, within 1 point of the leader, Firecrawl (73%). BrowseComp, which needs multi-hop reasoning across several searches, is its weakest benchmark: at 73.5% it ranks 13th of 26, well behind the Perplexity variants and Octen (85% to 87%)
➤ OpenAI Web Search uses fewer tokens for the same tasks: OpenAI bills about 40k input tokens per task, search results included, against 125k for the leanest Search API on the board, Perplexity (low). Its model cost is among the lowest on the board, at about $0.009 per task
➤ Built-in search costs less per task than most Search APIs: At ~$0.05 per task, OpenAI Web Search is cheaper than 17 of the 25 Search API products on the board (median $0.067), including Parallel (advanced, $0.06) and Perplexity (medium, $0.07). Octen costs about half as much ($0.024) and scores 3 points higher
Key details:
➤ Setup: GPT-5.6 Luna (medium reasoning) with OpenAI's web_search tool and search_context_size set to medium, in one Responses API call
➤ Pricing: $10 per 1k web_search calls, plus the search content, which OpenAI bills as model input tokens
➤ Contamination: we block known contamination sources through the tool's domain filter