AgenticBBO-Bench:评测 LLM 智能体黑盒优化能力的跨域基准
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
想用 LLM 智能体做黑盒优化的可以看看,五个领域统一评测,还拆了哪些设计选择真有用、哪些是伪收益。
研究团队发布 AgenticBBO-Bench,覆盖合成函数、超参数优化、数据库调优、芯片设计和分子设计五个领域,统一有限预算评测协议。实验显示 agentic BBO 在全部五个领域的平均得分高于直接使用 LLM 的方法,并在其中四个领域超过最佳数值优化器。分析发现额外增加数值工具并不总是提升性能,任务语义信息普遍有用,具体先验知识反而不够可靠。在五任务前沿挑战中,七款 LLM 在 Codex agent 框架下评测,GPT-6 Astra 和 DeepSeek-V4.1-Flash 位于性能与成本的 Pareto 前沿。
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.