Skip to content
AI.info

Research

Is Agentic RAG worth it? An experimental comparison of RAG approaches

Overview Research area: Retrieval-Augmented Generation (RAG) system design for large language model applications, specifically the comparison of "Enhanced" (fixed-pipeline) RAG against "Agentic" RAG,

Is Agentic RAG worth it? An experimental comparison of RAG approaches
arXiv
2601.07711
Published
2026-01-12
Authors
Pietro Ferrazzi, Milica Cvjeticanin, Alessio Piraccini, Davide Giannuzzi

AI summary

Overview

  • Research area: Retrieval-Augmented Generation (RAG) system design for large language model applications, specifically the comparison of "Enhanced" (fixed-pipeline) RAG against "Agentic" RAG, with attention to enterprise deployment.
  • Technical level: Intermediate. The paper assumes familiarity with RAG terminology (retrieval, reranking, query rewriting, embedding models), but the comparisons and conclusions are presented in accessible terms.
  • Scope in one sentence: An experiment-driven evaluation of Enhanced versus Agentic RAG across four dimensions of Naïve RAG weaknesses—user intent handling, query rewriting, document list refinement, and underlying LLM choice—plus a cost and latency analysis, across four datasets (NQ, FiQA, FEVER, CQADupStack-English).

What This Paper Is About

Basic ("Naïve") RAG systems retrieve documents for every query, which causes problems such as unnecessary retrieval for out-of-scope questions, weak matching between short queries and long documents, and noisy retrieved context. Two families of solutions have emerged: "Enhanced" RAG, which adds fixed modules (a router, a rewriter, a reranker) in a predefined sequence, and "Agentic" RAG, in which the LLM itself orchestrates the workflow and decides which actions to take and whether to iterate. This paper asks which paradigm practitioners should actually choose, and at what cost, by running both under matched conditions.

Key Contributions

  1. An experiment-driven comparison of Enhanced and Agentic RAG across four evaluation dimensions that map to specific Naïve RAG shortcomings: accuracy of RAG usage for in-scope versus out-of-scope queries, query rewriting impact, retrieved-document refinement, and the impact of generator (LLM) strength.
  2. A detailed cost and computational time analysis of the two paradigms, expressed in fixed hardware costs, processed input/output tokens, and end-to-end latency.
  3. A practical summary of findings intended to support architectural choices in real-world RAG deployments, including a recommendation to combine the two approaches.
  4. A released dataset of valid and invalid queries for user intent handling, hosted on Hugging Face (anonymousubmission/user-intent-handling).

Main Findings

  • User intent handling: Enhanced routing (using the semantic-router framework with OpenAI text-embedding-3-small) scored F1 95.7 on FIQA, 87.9 on FEVER, and 96.6 on CQADupStack-EN. Agentic scored higher on FIQA (F1 98.8, recall 97.7) and CQADupStack-EN (F1 99.8, recall 100), but much lower on FEVER (F1 64.6, recall 49.3) — a gap of −28.8 F1 points. The authors attribute this to FEVER's less restrictive domain definition, which makes it harder for the agent to recognise what counts as "valid". A Naïve RAG baseline scored recall 100 and F1 66.7 on all three datasets.
  • Query rewriting: Agentic RAG outperformed Enhanced on average by +2.8 NDCG@10 points. The largest gains appeared on NQ (+7.8 points, 51.7 versus 43.9 for Enhanced) and on the two IR/E tasks (FEVER +2 points, CQADupStack-EN +1.5 points). On FIQA the two performed essentially equally (43.2 Agentic versus 43.5 Enhanced). The authors attribute the Agentic advantage to its ability to decide dynamically whether and how to rewrite.
  • Document list refinement: Reranking helped Enhanced substantially, while the Agent gained no benefit from repeating retrieval. On NDCG@10 averages across FIQA and CQADupStack-EN, Enhanced without rewriting scored 48.0, Enhanced with rewriting 49.5, the Agent 43.9, and Naïve RAG 45.5. The agent chose to redo retrieval in 10% of cases, and 53% of the retrieved documents in those repeated attempts remained the same as in the previous step.
  • Underlying LLM strength: Swapping between Qwen3-0.6B (without thinking), Qwen3-4B, Qwen3-8B, and Qwen3-32B produced no significant difference in the performance patterns of the two systems — as the model gets larger, both Enhanced and Agentic improve at comparable rates.
  • Cost and time: Agentic settings required an average of 3.3 times more input tokens, 1.9 times more output tokens, and 1.5 times more time across datasets. The conclusion states Agentic RAG is systematically more expensive, up to 3.6 times more.
  • Latency breakdown for Enhanced RAG on valid queries: roughly 45–50% of total time goes to answer generation, a similar proportion to query rewriting, 0–5% to retrieval, and 0–2% to document re-ranking, making LLM calls the dominant factor.
  • Scale limitation: For NQ and FEVER, Agentic RAG would have taken more than 7 days on each of their knowledge bases, so those datasets were excluded from the document-refinement and underlying-LLM evaluations.

Methodology in Plain English

The authors start from a published list of Naïve RAG weaknesses and turn each one into an evaluation dimension, then build one Enhanced and one Agentic implementation to address each dimension and pit them against each other.

  • Enhanced RAG is a fixed pipeline: a router (semantic-router, text-embedding-3-small embeddings, top-20 most similar labelled examples, cosine similarity, with a similarity threshold that can return "None") decides whether to retrieve; HyDE-style query rewriting is forced before retrieval by prompting gpt-4o; a 300M-parameter ELECTRA-based cross-encoder reranker orders the 20 most similar documents; and the generator answers from the resulting context. The routing classifier uses example-based comparison against labelled "valid" and "invalid" query sets.
  • Agentic RAG is built on PocketFlow and has three nodes: an orchestrator, an answer node, and a retrieval node. The agent decides whether to call retrieval or answer, can rewrite the query itself, and can call retrieval again. The authors deliberately restrict this to a single-tool agent so its functional scope matches the Enhanced pipeline; multi-tool agents with planning or external APIs were excluded to avoid confounding the comparison.
  • Datasets: NQ (3,452 queries, 2,681,468 documents, average 1.2 relevant documents per query), FiQA (648 queries, 57,638 documents, 2.6), CQADupStack-English (1,570 queries, 40,221 documents, 1.4), and FEVER (6,666 queries, 5,416,568 documents, 1.2). Ground-truth query–document links come with each dataset.
  • Intent-handling test set: 500 valid and 500 invalid queries per dataset; valid queries came from the train splits and invalid queries were generated by prompting gpt-4o with a 5-shot prompt. NQ was excluded because it handles any type of query by design.
  • Metrics: F1 and recall for intent handling; NDCG@10 for retrieval quality in rewriting and refinement; and LLM-as-a-judge, using Selene-70B (a fine-tuned version of Llama-3.3-70B-Instruct), for final answer quality. For CQADupStack-EN the judge used a pairwise metric, validated by manual annotation of 5% of the test data (312 answer pairs): inter-annotator agreement was 71.9% and human-model agreement 65.4%, with 1.5 minutes per pair and 15.5 hours of annotation in total. For FIQA a 0–1 classification metric was used.
  • Cost setup: a t3.large EC2 instance at 0.09 $/h for the pgvector database, a t2.medium at 0.05 $/h for the backend, gpt-4o and OpenAI text-embedding-3-small with cosine similarity, and Qwen3 models on an 8×A40 (46GB) cluster — 0.6B, 4B, and 8B on one GPU, 32B on four — with an equivalent AWS g4ad.8xlarge costing 1.9 $/h. Retrieval chunks were set to 5 for the LLM comparison.

Why This Matters

Impact on research: The paper provides empirical numbers where prior work, such as Neha and Bhati (2025), proposed definitions and evaluation dimensions but stopped short of a full study. It shows that neither paradigm dominates, and it isolates which pipeline stages benefit from agency (routing, rewriting) and which do not (document refinement), giving future architecture work a finer-grained target than "agentic versus not."

Real-world applications:

  • Enterprise question answering over proprietary data and internal knowledge bots, which the paper notes are offered by providers including IBM, AWS, and Azure.
  • Financial question answering, represented by the FIQA dataset, where grounding answers in expert knowledge matters.
  • Claim verification against a knowledge base, represented by FEVER, where the system must retrieve evidence for or against a user statement and summarise it.
  • Forum and blog-post retrieval, represented by CQADupStack-English, where the system finds previously resolved posts related to a user's question and summarises them rather than answering directly.

Industry relevance: The cost analysis is the practical core. Agentic RAG's 3.3× input-token, 1.9× output-token, and 1.5× time overheads (up to 3.6× more expensive by the conclusion's framing) mean a well-optimised Enhanced pipeline can match or exceed agentic performance while being cheaper. The authors' concrete recommendation is a hybrid: use agentic components for user-intent routing (which needs no manually crafted examples) and query rewriting, but add an explicit reranking step to agentic pipelines because the agent does not reliably improve its own retrieved document set.

Future Directions

  1. Testing multi-tool agents with planning and external API access, which this study deliberately excluded to keep the comparison fair — the authors leave open how added capabilities would shift the trade-off.
  2. Adding explicit reranking into agentic pipelines, which the results suggest "could provide substantial gains" but which was not implemented here.
  3. Scaling the reported experiments to large knowledge bases such as NQ and FEVER, where Agentic RAG would have taken more than 7 days per knowledge base in this setup.
  4. Designing latency optimisations focused on LLM calls, since 45–50% of Enhanced RAG time on valid queries goes to answer generation with a similar share going to query rewriting.

Target Audience

Practitioners and architects deciding how to build production RAG systems, engineering teams weighing agentic orchestration frameworks against traditional pipelines, and applied NLP researchers who need empirical evidence (rather than definitions alone) for when Agentic RAG is worth its cost. The paper is also useful to researchers studying evaluation methodology for RAG sub-components, given its dimension-by-dimension metrics and its manual-annotation validation of an LLM-as-a-judge metric.

Authors’ abstract

Retrieval-Augmented Generation (RAG) systems are usually defined by the combination of a generator and a retrieval component that extracts textual context from a knowledge base to answer user queries. However, such basic implementations exhibit several limitations, including noisy or suboptimal retrieval, misuse of retrieval for out-of-scope queries, weak query-document matching, and variability or cost associated with the generator. These shortcomings have motivated the development of "Enhanced" RAG, where dedicated modules are introduced to address specific weaknesses in the workflow. More recently, the growing self-reflective capabilities of Large Language Models (LLMs) have enabled a new paradigm, often referred to as "Agentic" RAG. In this approach, an LLM orchestrates the entire process, deciding which actions to perform, when to perform them, and whether to iterate. Despite the rapid adoption of both paradigms, it remains unclear which approach is preferable under which conditions. In this work, we conduct an empirically driven evaluation of "Enhanced" and "Agentic" RAG across multiple scenarios and dimensions. Our results provide practical insights into the trade-offs between the two paradigms, offering guidance on selecting the most effective RAG design for real-world applications, considering both performance and costs.

Read the original paper