Skip to content
AI.info

Research

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Overview Research area: Natural Language Processing / LLM-driven automated scientific discovery, specifically e

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
arXiv
2609.40340
Published
2026-09-30
Authors
Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang

AI summary

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Overview

Research area: Natural Language Processing / LLM-driven automated scientific discovery, specifically evolutionary program search scaffolds augmented with web retrieval.

Technical level: Advanced. The paper assumes familiarity with evolutionary search scaffolds (FunSearch, OpenEvolve, EvoX), bi-level optimization, LLM mutation operators, and retrieval-augmented generation.

Scope: The paper proposes EvoDuet, a bi-level method that co-evolves candidate solutions (outer loop) and web search queries (inner loop) with frozen LLM parameters, and evaluates it on 21 optimization tasks against OpenEvolve.

What This Paper Is About

LLM-driven evolutionary search can stall when the knowledge needed to make progress is not present in the model's parameters or in the run's evolutionary history. Simply bolting a web search tool onto the loop does not fix this, because the same pages keep getting retrieved as the solution changes. EvoDuet addresses this by letting the LLM decide from its own knowledge gap when to search (retrieve new documents, reuse stored ones, or proceed without documents) and by evolving what to ask, so that search queries keep pace with the evolving solution.

Key Contributions

  1. A bi-level formulation of discovery. EvoDuet formalizes discovery as an outer optimization over solutions subject to an inner optimization over admissible queries, with query quality defined by the evaluator score of the candidate a retrieved document is predicted to yield. Because the inner objective is only observable after generation and evaluation, the paper approximates it with the frozen LLM as a score predictor.

  2. A knowledge-gap-based retrieval gate. At each iteration the LLM outputs a knowledge state and one of three decisions: no-op (internal knowledge suffices), look-up (reuse documents already stored in the search database), or retrieve (invoke the inner loop for new web search).

  3. Hypothetical evidence scoring. In the inner loop, a single LLM call predicts an absolute evaluator-scale score for each retrieved document, producing a surrogate signal for query optimization without generating or evaluating any candidate. Documents are ranked and the top-D are passed to the outer loop.

  4. Gated parallel candidate generation. EvoDuet generates N candidates per iteration for look-up and retrieve decisions and 1 for no-op, and passes only the best valid candidate to the selection policy.

Main Findings

  • EvoDuet improves discovery for strong backbones. At one candidate per iteration, EvoDuet raises OpenEvolve's overall normalized discovery gain (NDG) from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash.

  • Qwen3.5-9B does not benefit. Despite gaining 10.3% from oracle documents, Qwen3.5-9B shows a 14.4% decline under EvoDuet at N=1 and still loses 4.7% at N=8. Inspection of 50 randomly sampled revisions identified unused methods in 6% and incorrect implementations in 20% (26% combined).

  • New state-of-the-art results on eight tasks. EvoDuet surpasses the previously reported best scores on eight tasks across five domains and matches them on three more, with a mean cost of $45.56 across the eleven reported runs. Examples include Swap Reduction (15,186 to 14,835), Rosetta (1.552968 to 1.396424), Voyager 2 (3.430214 to 3.430206), and Erdős (0.380868 to 0.380859).

  • Web search helps, but not automatically. Oracle documents raise average NDG on 31 tasks by 10.3% with Qwen3.5-9B and 4.0% with GPT-5.6-Luna, yet lower NDG by more than 1% on 8 and 4 of the 31 tasks, respectively. The smaller GPT-5.6-Luna gain reflects limited headroom: OpenEvolve alone already reaches at least 95% NDG on 16 of the 31 tasks.

  • Exploration room matters. On Sums/Diffs, oracle documents offer no improvement at N=1 (14.8% with versus 15.4% without) but raise NDG from 25.8% to 40.7% at N=8.

  • Query evolution prevents retrieval stagnation. On Denoising, in-loop (joint-level) search used Tavily in 99 iterations producing 185 queries versus 201 for EvoDuet, but retrieved only 85 distinct URLs versus 248; 88.1% of joint-level returned URLs had appeared before versus 62.4% for EvoDuet. Joint-level search stopped improving after iteration 50 with held-out NDG of 27.7% versus 84.5% for EvoDuet.

  • Parallel generation scales monotonically. On 8 tasks with N in {1, 8, 16}, average NDG rises from 89.6% to 95.7% for GPT-5.6-Luna and from 84.5% to 89.6% for Gemini-3.8-Flash.

  • Largest gains occur on partially solved tasks. Across 126 task/model/candidate-count combinations, median gains were -5.4% (OpenEvolve NDG 5–20%), +17.5% (20–40%), +9.2% (40–60%), and +2.4% (60–80%). When OpenEvolve NDG was at least 80%, the median gain was 0.0% across 75 comparisons. Pearson r = -0.41 between OpenEvolve NDG and EvoDuet gain.

  • Mathematics sees the largest average gain. EvoDuet improves NDG on mathematics in five of six model/budget settings, with an average gain of 7.1% across all six; at N=8, Qwen3.5-9B gains 5.5% and Gemini-3.8-Flash gains 6.7%.

  • Knowledge-gap gating beats alternatives. On Denoising and Erdős respectively: random gating (p=0.5) 58.5% / 93.9%, stagnation heuristic 60.6% / 99.5%, knowledge-gap gate 84.5% / 100.0% — a 23.9% gain over the stagnation heuristic on Denoising.

  • Bi-level coupling beats sequential and in-loop coupling. On Molecule, Burgers, and CP (n=26), EvoDuet scores 0.8524 / 0.7846 / 2.635983, versus OpenEvolve 0.8496 / 0.6937 / 2.635983, joint-level 0.8474 / 0.6919 / 2.635980, and DeepEvolve 0.8149 / 0.6666 / 2.581971.

  • The loop packs with existing scaffolds. ΔNDG on Sums/Diffs and Denoising respectively: OpenEvolve +23.5% / +84.5%, Top-K +46.7% / +63.6%, EvoX +55.1% / +37.8%.

  • Predicted scores track real scores. Across 4,324 retrievals, hypothetical evidence scores correlate with selected candidates' evaluated scores at Spearman ρ = +0.86, positive on all 21 tasks with a median of +0.74.

  • Promising documents raise improvement rates. When retrieve retains at least one document predicted to beat the parent, the selected candidate improves on its parent in 67.1% of iterations and sets a new run best in 14.1%, versus 52.1% and 9.6% under no-op; under look-up the rates are 61.2% and 10.1%. On 17 of 21 tasks, medians were 69% versus 54%. However, 36.5% of retrievals yield no promising document, and there the parent improvement rate is only 38.3%.

  • Documents are used mostly for method transfer. Across 8,200 iterations from 82 GPT-5.6-Luna runs, applying principles and methods from retrieved documents occurred in 55 of 82 runs, using published best scores as reference targets in 40 runs, unsuccessful revisions inspired by retrieved ideas in 29 runs, and reuse of published solutions in 6 runs. Public artifact reuse occurred in only 6 of 82 GPT-5.6-Luna runs (7.3%), and only two of the 11 SOTA results reuse public artifacts in their best programs.

  • Cost efficiency on Denoising. EvoDuet achieves 100.27% NDG at $38.45, a 6.9× cost reduction versus SimpleTES (100% NDG at an estimated API-equivalent cost of $265.39).

Methodology in Plain English

The researchers start from a standard evolutionary search scaffold: a frozen LLM proposes candidate solutions, a deterministic evaluator scores them, and good candidates seed the next round. They first run a diagnostic study on 31 tasks, comparing plain OpenEvolve against OpenEvolve given oracle documents at every iteration, and find that knowledge helps unevenly — sometimes a lot, sometimes negatively.

They then build EvoDuet around two nested loops. The outer loop does the usual solution evolution, but the prompt can now include retrieved web documents. The inner loop runs only when the gate says to search. It works in R rounds: the model writes a batch of J queries targeting its remaining knowledge gaps, the queries are executed (via Tavily), the new documents join the retained pool, and a single LLM call assigns each document a predicted absolute evaluator score — the score the model thinks a candidate built on that document would get. The top-D documents by predicted score are kept, and the model's knowledge state is updated so the next round of queries can target what is still missing. No candidate is generated or evaluated during the inner loop.

A knowledge-gap gate ties the loops together: at each outer iteration the LLM reports what it knows, what it learned from prior searches, and what questions remain about the current parent, then chooses no-op, look-up, or retrieve. When documents are used, the model generates N candidates in parallel from the same prompt and only the best valid one is passed to the selection policy; the documents are stored alongside the candidate's actual evaluator score so later iterations can decide to reuse them.

Evaluation uses Normalized Discovery Gain (NDG), which measures the percentage of the gap between the initial solution's score and the previously reported SOTA score that a run closes. Each baseline is run with four random seeds, and results are reported as the best NDG across four runs.

Why This Matters

The paper shows that "add a search tool" is not the same as "open the search loop." Making the LLM judge its own knowledge gap, and evolving the queries as the solution evolves, changes both how many distinct pages are found and whether progress continues — on Denoising, joint-level search retrieved 85 distinct URLs and stalled at 27.7% held-out NDG, while EvoDuet retrieved 248 and reached 84.5%. This is evidence that retrieval design, not just retrieval availability, is the bottleneck for LLM-driven discovery.

Real-world applications named or implied by the task set:

  • Quantum compilation: the Swap Reduction task on Q20, where EvoDuet found a router using 14,835 SWAPs, 2.31% fewer than the released SimpleTES program (15,186) under the same evaluator.
  • Astrodynamics: the Voyager 2 trajectory task, improved from 3.430214 to 3.430206.
  • Scientific ML and biology: single-cell RNA-seq denoising, where EvoDuet reached 100.27% NDG at $38.45 (6.9× cheaper than SimpleTES), and protein-related optimization via the Rosetta task (1.552968 to 1.396424).
  • Mathematics and algorithm engineering: Erdős minimum-overlap, Sums/Diffs, Hadamard, and circle packing at n=26 and n=32; the intro also cites GPU kernel design as a target domain for such scaffolds.

Industry relevance: the cost results matter for anyone paying for LLM inference at scale — a 6.9× cost reduction on one task while exceeding the prior SOTA score suggests that gated, query-evolved retrieval can be cheaper than brute-force search. The cross-scaffold results (gains of +23.5%/+84.5% on OpenEvolve, +46.7%/+63.6% on Top-K, +55.1%/+37.8% on EvoX for Sums/Diffs and Denoising) mean teams already running an evolutionary scaffold can adopt the method without replacing their infrastructure. The negative Qwen3.5-9B results also function as a practical warning: co-evolution appears to require a backbone capable of both finding and applying evidence, with 26% of sampled revisions either leaving retrieved methods unused or misimplementing them.

Future Directions

  • Closing the weak-backbone gap. Qwen3.5-9B gained 10.3% from oracle documents but lost 14.4% at N=1 and 4.7% at N=8 under EvoDuet, with 26% of sampled revisions failing to use or correctly implement retrieved methods. Whether stronger gating, verification of retrieved methods before use, or better prompt scaffolding can make weaker models benefit remains open.

  • Validating predicted gains before spending budget. Documents predicted to help still produced much worse candidates in 29 of 82 runs, and 36.5% of retrievals yield no promising document. The paper notes predicted gains must still be validated by the evaluator, raising the question of whether cheap intermediate validation could be inserted.

  • Better handling of already-solved tasks. When OpenEvolve's NDG was at least 80%, EvoDuet's median gain was 0.0% across 75 comparisons, and of 52 comparisons with NDG at least 95%, six declined by more than 5%. Knowing when not to search — or how to allocate a search budget under high headroom — is unresolved.

  • Tuning the inner-loop hyperparameters and extending beyond the tested scaffolds. The inner loop has rounds R, query batch size J, and retained-document count D, and the paper reports applicability across OpenEvolve, Top-K, and EvoX on Sums/Diffs and Denoising; testing broader scaffold families and task domains is a natural next step.

Target Audience

Researchers and engineers working on LLM-driven automated discovery, evolutionary program search, and agentic retrieval systems will get the most from this paper, since it assumes fluency in scaffold terminology and bi-level optimization. Practitioners who already run scaffolds such as OpenEvolve, Top-K, or EvoX and want a drop-in retrieval layer will find the cross-scaffold results and the $38.45 versus $265.39 cost comparison directly actionable. Readers interested in retrieval-augmented generation more broadly will find the negative Qwen3.5-9B result and the document-behavior taxonomy useful as a reality check on when retrieval actually pays off.

Authors’ abstract

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.

Read the original paper