Skip to content
AI.info

Research

AIM: Agentic Idea Management for Automated Research

AIM: Agentic Idea Management for Automated Research Overview Research area: Automated scientific research with LLM agents, specifically the design of search procedures that decide which research idea

AIM: Agentic Idea Management for Automated Research
arXiv
2609.38445
Published
2026-09-29
Authors
Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen, Mihir Parmar, Rui Meng, Chun-Liang Li, Xiangru Tang, Sharon Li, Jinsung Yoon, Tomas Pfister

AI summary

AIM: Agentic Idea Management for Automated Research

Overview

Research area: Automated scientific research with LLM agents, specifically the design of search procedures that decide which research idea to try next under a limited experimental budget.

Technical level: Intermediate. The method is described as a pipeline of agent modules, which is readable without deep background, but the paper also includes a formal analysis section with assumptions, definitions, and propositions, and it assumes familiarity with baselines such as Bayesian optimization, MCTS, and beam search.

Scope: The paper proposes the Agentic Idea Manager (AIM), a fully autonomous framework that organizes research ideas into clusters, estimates their relative promise, dispatches selected ideas to solver agents, audits the resulting solutions for validity, and adaptively allocates the remaining experimental budget, evaluated on 10 AutoLab tasks.

What This Paper Is About

Automated research agents can propose approaches, write code, run experiments, and read verifier feedback, but implementation and evaluation are expensive, so an agent must decide how to spend a fixed budget across many possible directions. The authors separate existing methods into solution-driven search (optimizing code artifacts directly) and idea-driven search (reasoning over research ideas and delegating implementation to a solver), and argue that idea-driven search raises three management problems that current methods handle with fixed rules. AIM is the authors' answer: an idea-driven framework whose components mirror the surrogate-and-acquisition structure of Bayesian optimization, evaluated for score and for wall-clock efficiency.

Key Contributions

  1. A categorization of automated research. The paper introduces what it describes as the first categorization of automated research by primary unit of search: solution-driven versus idea-driven. It identifies three central challenges of idea-driven approaches: idea management, idea selection, and idea–solution integrity.

  2. The AIM framework. A fully agentic idea-driven framework that manages ideas as semantic clusters via an Agentic Surrogate, selects ideas through an Agentic Acquisition mechanism with explicit explore/exploit actions, and audits solutions with a Solution Auditor. A Resource Planner dynamically allocates the remaining budget across parallel search branches.

  3. Empirical results on 10 AutoLab tasks. AIM reports the highest average score in both task groups: 67.0% on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline (ScientistOne) by 1.6 and 4.9 percentage points respectively, and reaching the best baseline performance up to 3.1× faster in wall-clock time.

  4. A theoretical analysis of when idea-level search helps. The paper formalizes semantic coverage and derives a bound relating idea-driven search to any procedure evaluating the same number of solutions, plus a probabilistic condition under which broader coverage is valuable.

Main Findings

  • Best averages in both task groups: AIM scores 67.0% on System Optimization and 55.8% on Model Development & CUDA, versus 65.4% and 50.9% for the strongest baseline, ScientistOne (gains of 1.6 and 4.9 points).

  • System Optimization detail: AIM leads the strongest solution-driven baseline, AdaEvolve (63.3%), by 3.7 points. The largest gain is on Flash Attention, where AIM exceeds ScientistOne by 4.3 points and AdaEvolve by 5.2 points.

  • Per-task System Optimization scores for AIM: Flash Attention 90.5, Radix Sort 69.5, FFT Rust 56.5, AES128 Ctr 66.4, Z-order Range Scan 52.1.

  • Long-horizon tasks: AIM has the highest mean score on three of five tasks — Moving MNIST World Model (61.5), Huffman Canonical Decode (43.4), and NTT Butterfly (59.0) — improving over ScientistOne by 4.2, 0.4, and 2.1 points respectively. AIM also scores 61.8 on Data Selection IFEval and 53.5 on ICP Correspondence Step.

  • One clear loss: On Data Selection IFEval, AIM (61.8) outperforms the evaluated idea-driven baselines but falls below AdaEvolve (82.7). The authors speculate this is a smooth, local-search-friendly landscape where score is dominated by fine-tuning a single heuristic recipe, which suits AdaEvolve's mutation loop against a persistent best program.

  • Missing baseline results: Several baselines could not handle the Moving MNIST World Model task because it requires multiple file outputs, so their averages are omitted or shown as "–".

  • Wall-clock efficiency: AIM reaches the final score levels of all competing baselines within roughly the first 1–2 hours, whereas baselines need between 2.2 and 5.5 hours to reach those scores. AIM reaches its best score of 90.5% after 3.3 hours, compared with AdaEvolve's 85.3% at 5.5 hours and ScientistOne's 86.2% at 3.5 hours.

  • Ablations (Flash Attention, full model 90.5%): Removing the Agentic Surrogate drops to 85.9% (−4.6); removing the Solution Auditor drops to 87.6% (−2.9); replacing dynamic allocation with a fixed 5×5 schedule drops to 88.1% (−2.4); removing explicit acquisition actions drops to 89.2% (−1.3).

  • Qualitative case (Data Select IFEval, iteration 3): The Organizer partitions a 33-idea pool into four clusters and ranks "metadata stratification" first, grounded on its best member idea at 0.377. In that iteration, an exploit/exploit branch (Source-Balanced IO-Length Stratification, rank 1/6) regressed to 0.119, while an exploit-cluster/explore-idea branch (Unsupervised TF-IDF + KMeans Stratification, rank 5/6) matched the trailing best at 0.377; a third branch scored 0.230.

  • Theory, part 1: Under the stated assumptions, the semantic coverage of an extreme idea-driven procedure that sends N executions to N distinct ideas is an upper bound on the coverage of any procedure evaluating at most N executable solutions. The paper notes this does not mean solution-driven search must have lower coverage — rather that idea-driven search makes semantic breadth explicitly controllable.

  • Theory, part 2: Under a competitive-direction exchangeability assumption, a search with coverage C finds at least one ε-optimal direction with a stated probability, and a sufficient condition on C guarantees success probability 1 − δ (the rendered text of this final condition is truncated).

  • Not reported: Specific dates, dataset sizes beyond task counts, hardware details, and code-release status are not given in the provided content; a project page URL is listed.

Methodology in Plain English

AIM treats each research idea as a natural-language object — a short title, a hypothesis or abstract, and experiment plans — and keeps a search state containing the task context, the idea pool, accumulated implementation lessons, a cluster organization, promisingness rankings, and past idea–score observations. The task context and ideator structure follow ScientistOne; the research brief is generated from the task description using Claude Code.

The pipeline has four parts:

  1. Agentic Surrogate. An Organize operator rebuilds a cluster map from the full idea pool and evaluation history at every iteration, so ideas from different generation lineages can be grouped together and the map can shift as evidence arrives. An Estimate operator then produces ordinal (ranked, not calibrated numeric) promisingness estimates at both cluster and idea level, considering observed performance, evidence scarcity, semantic novelty, and implementation lessons.

  2. Agentic Acquisition. A Dispatch operator assigns each parallel solver branch a two-tier action — explore or exploit at the cluster level, and explore or exploit at the idea level — through two sequential LLM calls, then picks a target cluster and idea per branch. Each selected idea goes to an independent Solver that returns executable code, a verifier score, and an execution record. An Expand operator then extracts lessons into memory and generates new candidates using one of four modes: score-guided refinement, cross-pollination, error-guided repair, or novel idea generation.

  3. Solution Auditor. It flags results as trivial, task mismatch, idea mismatch, or reward hacking. Invalid results are discarded and excluded from the history and lesson extraction (though the execution still consumes budget). If only idea mismatch is detected, the auditor reconstructs the idea and re-associates the score and lessons with it, closing the loop from idea to implementation to audit to alignment.

  4. Resource Planner. Given a total execution budget N and B_tot total branches, each branch receives N_branch = floor(N / B_tot) executions. The planner chooses how many branches run per iteration, trading parallel breadth against how often the search can update from feedback, while preserving total branch and execution budgets.

Evaluation uses the Gemini-3.1-Pro-Preview backbone for all methods and the Gemini Deep Solver used in ScientistOne (with a Claude Code substitution analysis placed in the appendix). Budgets were at most 300 executions and 6 wall-clock hours for System Optimization, at most 60 executions and 24 hours for Model Development, and at most 60 executions and 12 hours for CUDA. Comparisons span solution-driven baselines (EvoX, AdaEvolve, AIRA in evolutionary and MCTS variants) and idea-driven baselines (DeepScientist, AI-Scientist-v2, Arbor at depth 2 and 3, and ScientistOne).

Why This Matters

Impact on research. The paper reframes automated research as a budget-allocation problem over abstract directions rather than only over code. It argues that making semantic breadth explicit makes it directly controllable — and it shows that when many plausible directions exist but few are competitive, breadth should matter more. That gives builders of research agents a concrete design vocabulary (idea management, idea selection, idea–solution integrity) and a warning that a solver can silently produce code that does not implement the selected idea, corrupting later lessons.

Real-world applications (as studied in the benchmark tasks):

  • Kernel and systems optimization, such as Flash Attention, Radix Sort, AES128 Ctr, FFT Rust, and Z-order range scan.
  • GPU kernel development, such as Huffman canonical decode, NTT butterfly, and ICP correspondence step.
  • Data curation for LLM fine-tuning, such as selecting a training subset to hit a target benchmark (IFEval).
  • Model development workflows, such as a Moving MNIST world model.

Industry relevance. Resource efficiency is the paper's practical hook: AIM reaches competing baselines' final scores within about the first 1–2 hours, and reaches the best baseline's performance up to 3.1× faster in wall-clock time, at up to 3.3 hours for its own best score of 90.5% on Flash Attention. For teams paying for compute and inference, being able to control how a fixed experiment budget is spread across directions — rather than only tuning code — is a direct cost lever.

Future Directions

  • Hybrid designs. The paper's own discussion suggests that a hybrid of solution-driven and idea-driven approaches may be desirable to ensure faithful implementation of ideas, given that Assumption 2 (faithful realization) is described as somewhat idealistic and is argued to be supported by the Solution Auditor.

  • Solver portability. The framework is demonstrated mainly with the Gemini Deep Solver; a solver substitution analysis using Claude Code is placed in Appendix D.1, leaving open how well AIM transfers across different solver backends.

  • Extending the theory. The analysis establishes an upper bound on semantic coverage and a probabilistic success condition, but the practical guidance for choosing how much breadth a given task warrants is left at the level of a sufficient condition.

  • Closing the gap on locally friendly tasks. AIM underperforms AdaEvolve on Data Selection IFEval (61.8 versus 82.7), which the authors attribute to that task rewarding fine-grained mutation against a persistent best program rather than exploration of qualitatively different strategies — an open question for idea-driven designs.

  • Component-level depth. The paper defers further analysis to appendices, including an in-depth ablation of the Agentic Surrogate, the preciseness of the Estimate operator's ordinal predictions, and the distributions of Dispatch explore/exploit actions and Expand generation modes.

Target Audience

Researchers and engineers building autonomous research or AutoML-style agents, particularly those working on LLM-based scientific discovery pipelines; practitioners who need to allocate expensive experiment budgets across many candidate directions; and readers interested in the formal question of when searching over abstractions beats searching over artifacts. The paper will be most useful to those already familiar with baselines such as AIDE, AlphaEvolve, AI Scientist-v2, and ScientistOne, since much of its argument is framed as a contrast with them.

Authors’ abstract

Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/

Read the original paper