Research
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
Overview Research area: Large language model (LLM) agents applied to black-box optimization (BBO), and the benchmarking methodology needed to compare them fairly across scientific and engineering doma

- arXiv
- 2610.12183
- Published
- 2026-10-08
- Authors
- Ming Chen, Rong-Xi Tan, Ke Xue, Yu-Jie Zhou, Taiye Lu, Zhi-Xuan Gao, Peng Xie, Zijun Shen, Chen Lu, Haopu Shang, Chao Qian
AI summary
Overview
Research area: Large language model (LLM) agents applied to black-box optimization (BBO), and the benchmarking methodology needed to compare them fairly across scientific and engineering domains.
Technical level: Advanced. Readers should be comfortable with Bayesian optimization, evolutionary algorithms, surrogate models, hyperparameter search, and LLM agent harnesses.
Scope: The paper introduces AgenticBBO-Bench, a unified cross-domain benchmark for evaluating complete LLM-agent optimization systems under a shared finite-budget protocol, and uses it to isolate which design choices (tools, task information, LLM role) actually help.
What This Paper Is About
Black-box optimization targets problems where you can evaluate an objective but cannot see its formula or gradients, and where each evaluation is expensive. Recent work puts an LLM agent in charge of the search loop, letting it inspect the search state, run code, call tools, and decide the next candidate online — a setting the authors call Agentic BBO. The problem is that existing agentic BBO studies use different task domains and system configurations, so their results cannot be compared and individual design choices cannot be isolated. This paper builds a common benchmark across five optimization domains and uses it to test what actually drives agent performance.
Key Contributions
-
AgenticBBO-Bench, a cross-domain benchmark for agentic BBO covering five task families — synthetic numerical optimization (BBOB), hyperparameter optimization (Bayesmark), database tuning (DBTune), chip design / macro placement (BBOPlace-Bench), and molecular design (GuacaMol) — placed under a unified finite-budget evaluation protocol.
-
A systematic study of three design factors in agentic BBO: numerical optimization tools, task information and prior knowledge, and the degree of LLM involvement in the optimization loop.
-
A compact five-task frontier challenge, with one representative task per domain, used as a standardized leaderboard; seven frontier LLMs were evaluated under the same protocol and the Codex agent harness.
-
A common standardized leaderboard for comparing future general-purpose models and agent systems on BBO; code released at https://github.com/lamda-bbo/agentic-bbo.
Main Findings
-
Agentic beats Direct on all five families. The Agentic configuration outperformed the direct LLM-based method on all five benchmark families, with an average relative improvement of 30.3%. The largest gains were on BBOB (0.460 to 0.661, +43.7%) and GuacaMol (0.565 to 0.840, +48.7%).
-
Agentic beats the best numerical optimizer on four of five families. Agentic achieved the best score on BBOB, BBOPlace, HPO, and GuacaMol. TuRBO was best on DBTune, where Agentic scored 0.525 and TuRBO scored 0.660.
-
More optimizer tools do not consistently help. On an eight-task diagnostic set, base Agentic reached a mean of 0.589, compared with 0.574 when a GP candidate-proposal interface was added (+ GP Suggest) and 0.538 when surrogate prediction, acquisition scoring, diagnostics, and search-region controls were added as well (+ GP Tools). Effects varied by task, helping some and hurting others.
-
Agents keep writing their own code rather than delegating to the GP. With + GP Tools, 35.4% of rounds used only Python and another 17.4% used both Python and GP tools. In the Bigblue1 chip placement case study, the agent fit a linear proxy from observed layouts and refined a placement, improving best-so-far from 0.551 after the first new evaluation to 0.664 after four evaluations and 0.676 eventually.
-
Task semantics help broadly; specific priors are less reliable. Across six real-world tasks, restoring task and parameter semantics raised the mean score from 0.465 to 0.595. Adding domain knowledge helped Bigblue1 and Sysbench-5 but hurt PostgreSQL-5 and Adaptec1, lowering the mean to 0.569.
-
A prior only helps if it is correct and actionable. On three anonymous 12-dimensional objectives with four active variables, combining correct support information with local geometry raised the mean score from 0.685 to 0.872, while either alone gave smaller gains and replacing the support with an incorrect one dropped performance to 0.682.
-
The LLM does not need to stay online for the whole run. Restricting the LLM to controlling a GP policy lowered the score from 0.569 to 0.516. Handing the search to TuRBO after an agent warm-up achieved 0.570, essentially matching continued Agentic optimization and clearly beating TuRBO started from the initial observations (0.475).
-
Offline-developed optimizers transfer inconsistently. An optimizer developed by the LLM reached 0.501 on held-out tasks — above a fixed GP baseline at 0.451 but below online Agentic optimization — with substantial variation across independently developed programs.
-
Frontier models differ sharply across domains. On the five-task challenge, GPT-6 Astra had the highest mean score (0.591), closely followed by DeepSeek V4.1 Flash (0.584), and no model dominated every domain. Performance was only weakly aligned with token cost: GPT-6 Astra's 59.1 versus DeepSeek V4.1 Flash's 58.4 came at a much higher official token cost.
-
The challenge is far from saturated. The five-task design is compact enough for efficient evaluation while retaining diversity, and current models still vary markedly in optimization behavior.
Methodology in Plain English
The researchers assembled five existing optimization task families into one benchmark: analytic synthetic functions (BBOB), machine-learning hyperparameter tuning (Bayesmark), database knob tuning (DBTune), chip macro placement (BBOPlace-Bench, using a 32-macro setting), and goal-directed molecular design (GuacaMol, where candidates are SMILES strings). These span continuous, mixed, and structured discrete search spaces, with different levels of exposed semantics and feasibility constraints.
Every method — agentic, direct-LLM, and numerical — gets the same initial observations and the same objective-evaluation budget on each task and seed, and only objective evaluations count against the budget; model inference, code execution, tool calls, and candidate validation do not.
Agents run inside an isolated Docker container with a shared workspace and a CLI bridge (bbo_tool.py) that talks to a host-managed evaluator. The agent sees a small task card and instruction file, and pulls details on demand through shared interfaces: get_task_context, get_search_space, get_trial_history, get_incumbent, write_candidate, and submit_candidate. Invalid candidates are rejected without consuming budget and can be revised in the same round. The evaluator and hidden task state stay outside the container.
Scoring follows a normalized anytime protocol: 70% weight on the average normalized best-so-far quality over the whole trajectory, and 30% on the normalized quality of the final best solution.
The default backbone is DeepSeek V4.1 Flash with Codex as the agent harness. "Direct" methods receive the task description and history in context and generate the next candidate; "Agentic" methods work in a persistent workspace with Bash/Python and structured access to the search state. Ablations draw on a common eight-task pool with two tasks from each non-molecular family. The tool study adds a GP proposal interface (+ GP Suggest) and a fuller GP toolset (+ GP Tools). The information study compares Anonymous, Semantic, and Full Prior conditions, plus controlled anonymous objectives where support and geometry information is varied. The role study compares online Agentic control, GP-policy control, handoff to numerical optimizers after an agent warm-up, and an offline LLM-developed optimizer.
Why This Matters
Existing agentic BBO results were hard to compare because each study used its own tasks and system configuration. This paper supplies a shared protocol, a shared scoring rule, and a shared frontier challenge, which lets researchers attribute performance differences to the agent design rather than to the evaluation setup. It also delivers a counterintuitive empirical message: adding more numerical tools is not automatically better, and detailed priors can actively hurt if they are wrong.
Real-world applications the benchmark covers:
- Hyperparameter optimization — tuning machine-learning models across mixed-type search spaces based on predictive performance.
- Database tuning — configuring interacting database knobs to improve system performance, evaluated here via learned surrogates built from measurements of the underlying system.
- Chip design (macro placement) — positioning movable macros into physically valid layouts judged by a placement objective.
- Molecular design — optimizing structured discrete SMILES strings under task-specific scoring functions while maintaining chemical validity.
The broader BBO literature cited by the paper also points to chemical design, materials optimization, and complex system configuration as domains where each objective evaluation is expensive.
Industry relevance: the cost analysis matters for deployment. Models with similar aggregate optimization scores can sit at very different price points, so choosing the most expensive frontier model does not remove the need for better agent design, task adaptation, and search strategies. Teams building optimization agents for hardware, databases, or drug discovery get a reusable evaluation harness and evidence about which design choices are worth the engineering effort.
Future Directions
-
Explain why general-purpose optimizer tools fail to help consistently. The paper shows a fixed GP interface can hurt, and attributes this to agents discovering task-specific structure themselves, but the conditions under which a tool helps remain open.
-
Make prior knowledge reliable rather than merely detailed. Correct support plus geometry reached 0.872 versus 0.682 for incorrect support, so the open question is how to establish that a prior is correct and actionable before it is used.
-
Improve transfer of offline-developed optimizers. The LLM-developed optimizer beat a fixed GP baseline (0.501 versus 0.451) but fell below online Agentic optimization and varied substantially across independently developed programs.
-
Reduce cross-domain performance variance. No model dominated all five challenge tasks, and models strong on one task degraded on another, leaving room for methods that improve robustness rather than specializing in one search space; the paper also notes that future Agentic BBO directions are discussed in Appendix C.
Target Audience
Researchers and practitioners working on LLM agents, automated optimization, and AutoML; Bayesian optimization and evolutionary computation researchers who want to know where agents add value over classical solvers; and engineers in chip design, database operations, hyperparameter tuning, and molecular discovery who need a standardized way to compare agent systems before committing to a model and harness. Readers with a basic grounding in optimization will get the most from the ablation sections, while the benchmark description and results tables are accessible to a broader technical audience.
Authors’ abstract
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.