Skip to content
AI.info

Research

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

Overview Research area: Natural language processing, specifically retrieval-augmented generation (RAG) pipeline configuration and hyperparameter optimization (HPO) using large language model agents. T

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents
arXiv
2610.08452
Published
2026-10-06
Authors
Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft

AI summary

Overview

Research area: Natural language processing, specifically retrieval-augmented generation (RAG) pipeline configuration and hyperparameter optimization (HPO) using large language model agents.

Technical level: Advanced. The paper assumes familiarity with RAG components (chunking, embeddings, reranking, generators), multi-objective optimization, Pareto frontiers, Bayesian optimization (TPE), and LLM-as-judge evaluation.

Scope: The paper introduces an LLM-agent optimizer that configures RAG pipelines by diagnosing whether each failed exam question failed at retrieval or at generation, and by proposing the next configuration using a knowledge base of public model rankings and per-token pricing.

What This Paper Is About

Configuring a RAG pipeline is an expensive search over many coupled choices — chunking, embedding model, index type, retrieval depth, reranker, query expansion, passage compression, and generator — and the best configuration changes from corpus to corpus. Existing optimizers (greedy search, multi-armed bandits, Bayesian optimization, genetic search) reduce each trial to a single aggregate score, so they cannot tell whether a low score came from retrieval failing to surface the right passages or from the generator answering incorrectly despite having them. The paper's goal is an optimizer that uses the per-question failure evidence every RAG pipeline already records — whether the gold answer spans appeared in the retrieved chunks — to drive its next proposal, and that also weighs accuracy against API cost.

Key Contributions

  1. A diagnose-then-propose loop. After each trial, a Diagnoser reads the per-question failure evidence, attributes each failure to retrieval or generation, and compresses it into a short diagnosis. A Proposer then reads that diagnosis and selects the next configuration.

  2. A knowledge-base-grounded, cost-aware agent. The Proposer prices and ranks candidate configurations against public model leaderboards and per-token pricing before spending a trial on them, so it can navigate the accuracy-cost trade-off without first learning from trial outcomes which models are strong or cheap.

  3. A corpus-derived exam. Open-ended, multi-hop questions are drafted with verbatim gold-answer spans, filtered through a sufficiency oracle, and kept only when they separate configurations spanning the search space. The exam is frozen once and reused as the optimization target for every trial. Exam generation is optional — a user can supply a question-answer set instead.

  4. Empirical comparison against random search, MO-TPE with and without a warm start, a knowledge-base-and-diagnosis ablation of the authors' own agent, and a no-search "KB Greedy" reference, across three multi-hop QA benchmarks plus a real-world healthcare corpus.

Main Findings

  • Higher held-out LLM-judge accuracy than every baseline on all three benchmarks. The agent reaches 0.825 on HotpotQA, 0.330 on MuSiQue, and 0.790 on MultiHop-RAG, above random search (0.774, 0.270, 0.706), MO-TPE (0.784, 0.301, 0.724), and warm-started MO-TPE (0.786, 0.301, 0.731).

  • Sample efficiency. Within its first 10 trials, the agent matches or beats the statistical baselines' full 30-trial judge accuracy: 0.817 against 0.786 on HotpotQA, 0.773 against 0.731 on MultiHop-RAG, and a tie on MuSiQue at 0.303 against 0.301. At 20 trials the numbers are 0.835, 0.326, and 0.783 respectively.

  • The knowledge base and diagnosis add only modest lift over a plain LLM optimizer in the trace. The no-KB/diag ablation tracks closest to the full agent; on MultiHop-RAG it matches the full agent's exact match (0.729) and F1 (0.748) and trails slightly on the judge (0.780 versus 0.790).

  • MuSiQue accuracy is low for every method. The paper attributes this to the search space: MuSiQue's up-to-four-hop questions need multi-step retrieval that conditions each hop on the previous hop's results, a pipeline shape the single-round search space does not support. The low ceiling reflects the search space rather than the optimizer, and the agent still ranks best on MuSiQue.

  • Cost-aware results on the UniDoc healthcare corpus. The agent's median curve reaches 77% exam accuracy at $0.000741 per query, while the strongest baseline, warm-started MO-TPE, peaks at 71.5% and pays $0.001284 — so the agent's best point is more accurate at 57.7% of the cost. It matches that 71.5% at $0.000288, 22.4% of what the baseline pays (reported in the abstract as about 22% of the cost). The agent leads every baseline above $0.000236 per query; warm MO-TPE leads only in a band from $0.000097 to $0.000236, where neither curve reaches 60% accuracy.

  • The cost-accuracy gaps survive per-seed testing. Using two-sided Mann-Whitney U tests with Holm correction per family of three baseline comparisons: the agent's per-seed peak exam accuracies span 74% to 82% and rank systematically above each baseline's (Holm-corrected p = 0.017 against warm MO-TPE, 0.025 against cold MO-TPE, 0.001 against random search). All ten agent seeds reach warm MO-TPE's median peak of 71.5% within their 30 trials, against five of ten warm MO-TPE seeds, four of ten cold MO-TPE seeds, and two of ten random-search seeds; ranked by the lowest per-query cost at which they reach that accuracy, the agent leads every baseline (Holm-corrected p ≤ 0.001).

  • The generated UniDoc exam skews toward multi-span reasoning. Of the 100 questions that survive validation and probe-based selection: Extraction 16, Definitional 0, Numeric (single-span) 1, Inference 25, Bridge 19, Comparison 25, Numeric (multi-span) 14. 98 questions keep all spans within one document; 2 cross into a second.

  • Corpus sizes and search space scale. HotpotQA has 19,276 documents / 1.81M words / 94 words per document; MuSiQue 17,629 / 1.32M / 75; MultiHop-RAG 609 / 1.07M / 1,760; UniDoc 250 / 1.30M / 5,214. The shared search space spans 2 chunking strategies over 4 chunk sizes and 4 overlaps, 13 embedding models, 2 index types with continuous retrieval depth and fusion weight, 4 query-expansion strategies, 5 reranker choices with continuous rerank depth, and 35 generator LLMs with a per-trial reasoning toggle.

  • KB Greedy is far below the search methods. Taking each dimension's top-rated option directly from the knowledge base without any trial gives 0.702 judge accuracy on HotpotQA, 0.254 on MuSiQue, and 0.544 on MultiHop-RAG.

Methodology in Plain English

The system runs a fixed budget of trials (30 in both experiments, repeated ten times with independent seeds). Each trial builds or loads a cached index for one candidate configuration, runs the whole pipeline over a frozen exam of question-answer pairs, scores the answers, and records which gold spans appeared in the retrieved chunks. That span record is the key: it lets the loop say a failure was a retrieval failure when the gold spans were missed, or a generation failure when they were retrieved but the answer was still wrong.

Two reasoning agents then interpret the trial. The Diagnoser reads a compact summary — overall accuracy, accuracy on questions where retrieval succeeded, the retrieval-versus-generation split of failures, a sample of failed questions spanning different failure types and question types, and how accuracy moved across earlier trials (but not their configurations, cost, the Pareto frontier, or its own earlier diagnoses). It emits a short narrative, findings tied to the trial's numbers, and illustrative questions. It does not prescribe configuration changes.

The Proposer then selects the next configuration, reading the diagnosis, the current configuration, the legal search space ranges, and a knowledge base of model rankings and pricing. It sees every configuration evaluated so far, which generators, embedding models, and rerankers remain untried, and how much budget is left. In cost-aware mode it reads the current accuracy-cost frontier and decides where along the cost axis to probe next.

Exam generation happens once, before optimization. The Composer reads neighborhoods of related chunks (an anchor chunk plus same-document siblings plus cross-document chunks ranked by TF-IDF cosine similarity), drafts questions tagged with one of seven types, and over-generates. Deterministic filters then reject questions that cite one span twice, refer to their source by name, have empty spans, or fail arithmetic verification. Multi-hop questions pass a sufficiency oracle that must return the smallest supporting subset of gold spans; a question is rejected if that subset is smaller than its hop count. Finally, four probe configurations spanning the search-space extremes (LLM, embedding model, reranker, chunk size, and retrieved top-k) evaluate every candidate: questions all probes answer are discarded because they no longer separate configurations, questions only one or a few probes solve are favored, and a small share that no probe solves is kept.

For the Pareto experiment the corpus is the healthcare subset of UniDoc-Bench — 230 PDFs and 20 page images, mostly biomedical research articles and medical studies — parsed with Docling using table-structure extraction and OCR. The authors use the documents as a corpus and generate their own exam from them rather than using UniDoc-Bench's labeled question-answer pairs. Cost is computed by summing the price of generation and query-expansion calls from LiteLLM's model catalog per query and averaging; embedding and reranking models run locally at no API cost.

In the Accuracy experiment, methods are scored each trial on a fixed 100-question validation exam built from the benchmark's labeled questions, and returned configurations are re-evaluated on a separate 300-question held-out gold test set disjoint from the search questions. MO-TPE is the multivariate TPE sampler from Optuna, following syftr; the authors restrict its space to dense and hybrid pipelines (leaving out agentic-RAG variants), disable its Pareto-frontier pruning, and omit the cross-dataset transfer seeding. The warm-started variant is seeded with the 30 completed trials of a paired random-search run on the same corpus, so it draws on information from 60 evaluations while every other method sees 30.

Why This Matters

Impact on research. The paper reframes RAG pipeline tuning from blind score optimization to evidence-driven search. Prior optimizers treat a trial as a number; this work shows that the retrieved chunks already contain a free signal about where the pipeline broke, and that feeding that signal back to an LLM proposer improves both final accuracy and sample efficiency. It also connects RAG optimization to the LLM-agent-as-optimizer literature and to corpus-derived automated evaluation, and it is one of the few RAG-tuning studies to run the full pipeline end to end on a real document corpus with a cost objective rather than an isolated benchmark.

Real-world applications.

  • Deploying question-answering over an organization's private document collection, where a practitioner wants the cheapest pipeline that clears an accuracy bar because inference cost is paid on every live query.
  • Healthcare and biomedical document search, the setting of the Pareto experiment, where corpora are multi-page PDFs with tables, figures, and mixed layouts.
  • Cost-controlled enterprise deployments where the search space includes dozens of generator models with very different per-token prices.
  • Any setting where the target corpus changes frequently enough that a pipeline tuned elsewhere underperforms and the search must be repeated.

Industry relevance. The cost-aware mode directly addresses the operational question teams face: not "what is the most accurate pipeline at any price" but "what is the cheapest pipeline that is accurate enough." The knowledge base also lets the optimizer skip configurations that public benchmarks predict will underperform, which matters when each trial costs real API spend. The paper reports per-trial search cost and embedding compute in Appendix D, though the appendix contents are not included in the truncated text available here.

Future Directions

  1. Richer cross-document bridges. The exam pipeline ranks cross-document bridges by lexical TF-IDF similarity, which connects only chunks sharing rare surface terms. An entity-relation graph over the corpus, as in graph-based RAG, would also surface bridges that share an entity without sharing its wording.

  2. Expanding the search space beyond one pipeline family. The diagnose-then-propose loop is agnostic to what the space contains, so it could tune whole architectures such as agentic RAG or graph-based RAG, searching across retrieval paradigms rather than within one.

  3. More objectives. The optimizer weighs only accuracy and cost today, but its knowledge base already records throughput and memory, making latency or memory footprint natural additions.

  4. Warm starts across runs. Each run begins cold on a new corpus; carrying diagnoses and configuration priors across runs would let the optimizer accumulate a retrievable agent memory and generalize configuration knowledge from one corpus to the next.

The paper also names open limitations: parsing is fixed outside the search space (Docling is run once and cached, since reparsing every trial would dominate wall-clock time); the knowledge base draws on Artificial Analysis rankings and LiteLLM pricing, neither peer-reviewed, and is a point-in-time snapshot that degrades as vendors re-rank, deprecate models, and update prices; agent nondeterminism remains because several reasoning models require a temperature of 1.0 and hosted APIs do not guarantee bit-exact outputs; and accuracy in both experiments is scored by an LLM judge with no human-agreement study, so any systematic judge bias carries into the reported numbers.

Target Audience

This paper is most useful to researchers and engineers working on RAG systems, LLM-agent-based optimization, and automated machine learning for NLP. It will appeal to practitioners who must choose RAG configurations under a real cost budget, to HPO researchers interested in replacing aggregate scores with structured diagnostic feedback, and to readers following the emerging line of work on LLM agents as optimizers. Some parts — the Pareto analysis, the MO-TPE comparisons, and the statistical testing — assume familiarity with multi-objective optimization.

Authors’ abstract

Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.

Read the original paper