Black-box optimization (BBO) poses foundational computational hurdles across chip design, database parameter tuning, hyperparameter optimization, and molecular engineering where objective function evaluations are prohibitively expensive and gradients are unavailable. Traditional Bayesian optimization and evolutionary algorithms rely strictly on numerical samples while discarding rich task semantics, whereas direct LLM prompting struggles with numerical precision and hallucinations. Researchers from Nanjing University's National Key Laboratory for Novel Software Technology (LAMDA Group) introduce AgenticBBO-Bench (arXiv:2610.12183), the first unified cross-domain benchmark for evaluating LLM agents in black-box optimization. Spanning synthetic functions, hyperparameter tuning, database configuration, chip layout, and molecular design under a standardized finite-budget protocol, Agentic BBO achieves superior family-averaged scores over direct LLM methods across all five domains and outperforms the strongest numerical optimizers in four. Evaluating seven frontier LLMs within the Codex agent harness establishes the Pareto frontier of performance and inference cost, with the benchmark and harness fully open-sourced.
Key Takeaways
- ✓Nanjing University LAMDA releases AgenticBBO-Bench, benchmarking LLM agents across chip design, DB tuning, and molecular optimization
- ✓Agentic BBO outperforms direct LLM prompting across all 5 domains and beats leading numerical optimizers in 4 out of 5 areas
- ✓Maps Pareto frontier of frontier models within Codex harness, cutting required optimization steps by over 38% under finite budgets
Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Background and the Problem
Black-box optimization (BBO) underpins complex engineering workflows spanning VLSI chip floorplanning, distributed database tuning, hyperparameter optimization (HPO), and molecular drug discovery. In these domains, evaluation functions involve rigorous physical simulators or hardware clusters where gradients are unavailable and each trial incurs immense computational expense. Classical methods such as Bayesian optimization (GP-BO) or CMA-ES treat problems as purely numerical black boxes, ignoring code structure and domain semantics. Conversely, zero-shot prompting of foundation models suffers from hallucinations and numerical inaccuracy. Designing systematic frameworks where coding agents orchestrate numerical solvers to optimize objective functions under tight evaluation budgets is an urgent industry priority.
Architecture and How It Works
To systematically analyze agentic black-box search, Nanjing University's LAMDA Group and collaborating institutions present AgenticBBO-Bench (arXiv:2610.12183, repository: lamda-bbo/agentic-bbo):
- Unified Protocol Across Five Critical Domains: Standardizes evaluation across synthetic test functions, machine learning HPO, database tuning, chip macro placement, and molecular optimization under a uniform finite-budget protocol.
- Codex Agent Harness with Executable Tooling: Embeds foundation models in interactive environments featuring a Python interpreter, mathematical solvers (Scipy, Optuna), and real-time execution feedback. The agent synthesizes search heuristics, executes local optimization scripts, and refines search trajectories based on empirical returns.
- Dissecting Core Drivers of Agentic Success: Isolates three critical factors, demonstrating that task semantics deliver robust acceleration across domains, arbitrary tool accumulation yields diminishing returns, and classical numerical optimizers effectively capitalize on exploratory global paths mapped out by agents.
Benchmarks and Measured Results
Benchmarked against classical numerical libraries and direct LLM methods under finite search budgets:
- Outperforming Specialized Numerical Solvers in 4 of 5 Domains: Agentic BBO decisively outscores direct prompting across all five domains and surpasses state-of-the-art numerical solvers (including Optuna and CMA-ES) in hyperparameter tuning, database configuration, chip floorplanning, and molecular generation.
- Pareto Frontier Analysis: Across seven frontier LLMs evaluated under the Codex agent harness, GPT models and DeepSeek lightweight models define the Pareto frontier balancing optimization quality against inference cost.
- Over 38% Search Budget Reduction: On constrained evaluation budgets, Agentic BBO reaches high-performance target objectives in 38% fewer function evaluations compared to conventional Bayesian optimization.
Getting Started for Developers
The LAMDA Group has released the complete code, environments, and agent harness on GitHub (lamda-bbo/agentic-bbo). For platform architects and R&D teams building EDA pipelines or cloud resource optimizers, AgenticBBO-Bench demonstrates that foundation models should not act as bare parameter guessers. Instead, developers should pair coding agents with sandboxed Python runtimes and numerical libraries, leveraging LLMs to interpret domain documentation and coordinate numerical solvers for high-stakes optimization.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.