Skip to content
AI.info

Research

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Overview Research area: Information retrieval (cs.IR), specifically the intersection of dense retrieval, large language model reasoning, and agentic systems. Technical level: Intermediate. Readers wil

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
arXiv
2610.05750
Published
2026-10-05
Authors
Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira, Oliver Holworthy, Jie He, Ronay Ak, Jiarui Cai, Ryan Chesler, Bo Liu, Even Oldridge

AI summary

Overview

Research area: Information retrieval (cs.IR), specifically the intersection of dense retrieval, large language model reasoning, and agentic systems.

Technical level: Intermediate. Readers will benefit from familiarity with retrieval metrics such as nDCG@10, embedding-based dense retrieval, and the ReAct-style tool-calling agent pattern, but the paper's argument is accessible without deep mathematical background.

Scope: The paper evaluates a ReAct-based retrieval agent built on top of standard dense retrievers across two complex retrieval benchmarks, measuring both accuracy gains and the compute costs of those gains.

What This Paper Is About

Standard dense retrieval ranks documents by the cosine similarity between query and document embeddings, which captures surface-level semantic overlap but not the reasoning, world knowledge, or iterative exploration that complex queries demand. The authors build a retrieval agent in which an LLM decides when, how often, and with what rewritten queries to call a dense retriever inside a ReAct loop, then test whether that agent outperforms standard retrieval on complex tasks, how well it transfers across domains, and what it costs to run.

Key Contributions

  1. An agentic retrieval pipeline (NeMo Retriever Agent) that pairs LLM reasoning with dense retriever corpus exploration in a ReAct loop using three tools: Retrieve (dense top-k search returning document IDs), Think (internal reasoning tokens that do not alter the environment), and Final_Results (returns an ordered list of document IDs plus a rationale, and errors if the count is wrong).
  2. A reliability mechanism: when the agent fails (for example by exceeding the context window or triggering content violation errors), the pipeline collects the ranked lists from each search attempt before the error and merges them with Reciprocal Rank Fusion (RRF).
  3. An infrastructure redesign: replacing a Model Context Protocol (MCP) server with a thread-safe singleton retriever living in the same process as the agent, loading model and corpus embeddings once, protecting access with a reentrant lock, and exposing the same retrieve() interface to concurrent agents. The authors report this eliminates a class of deployment errors and substantially improves GPU utilization and experiment throughput.
  4. A joint evaluation of effectiveness and cost on ViDoRe v3 and BRIGHT, contrasting agents built on a proprietary model (Opus 4.5) and an open-source model (gpt-oss-120b) against standard dense retrieval baselines.

Main Findings

  • Agentic retrieval beats standard retrieval with the same embedding model. The paper reports an average improvement of 8.7 points in nDCG@10. On ViDoRe v3, Opus 4.5 with colembed-vl-8b-v2 scores 69.22 versus 64.36 for standard retrieval with that embedding model; gpt-oss-120b with colembed-vl-8b-v2 scores 66.38. On BRIGHT, Opus 4.5 with llama-nv-reasoning-3b scores 50.79 versus 38.28 for standard retrieval; gpt-oss-120b with llama-nv-reasoning-3b scores 41.27.
  • The LLM's reasoning strength matters. Using the best embedding model for each dataset, the agent with Opus 4.5 achieves on average a 6.2 point higher nDCG@10 than the same pipeline with gpt-oss-120b. The gap is larger on BRIGHT, at 9.5 nDCG@10 points.
  • Stronger models explore more. On ViDoRe v3, Opus 4.5 makes 9.2 search calls per query on average, while gpt-oss-120b makes 2.4, which the authors suggest may help explain the performance difference.
  • The agent partially compensates for weaker embedding models. The performance gap between weaker and stronger embedding models is much smaller in the agentic setup than in standard retrieval. Agentic retrieval reduces that gap by almost half, 57 percent, on average.
  • Agentic retrieval generalizes across domains; specialized methods do not. As of March 13, 2026, the pipeline held the #1 spot on the ViDoRe v3 leaderboard and the #2 spot on the BRIGHT leaderboard. Evaluated on ViDoRe v3, the #1 BRIGHT method, INF-X-Retriever, scored 51.01, far below the agent's 69.22; when INF-X was re-run with the agent's embedding model (nemotron-colembed-vl-8b-v2), it scored 62.31, still below the 64.36 dense-only baseline. On BRIGHT the agent is reported at 50.90 in Table 2 (compared with 63.40 for INF-X-Retriever), while Table 1 reports 50.79 for the Opus 4.5 configuration.
  • The accuracy comes with substantial cost. The abstract reports agentic retrieval taking 107.4 seconds per query on average versus 0.67 seconds for standard retrieval, and consuming 764.1K input and 5.8K output tokens per query. Section 3.3 states the average agentic time as 107.45 seconds per query. Per-model breakdowns in Table 3: Opus 4.5 averages 136.3 seconds, 759.4k input tokens, and 6.2k output tokens; gpt-oss-120b averages 78.6 seconds, 768.9k input tokens, and 5.4k output tokens.
  • Three recurring successful search patterns are identified: query refinement (adjusting later queries based on newly discovered information), persistent rephrasing (reformulating until useful information is found), and complexity decomposition (breaking multi-part queries into simpler ones).
  • Retrieval cost shifts away from index size. In standard retrieval, costs correlate mainly with search index and corpus size; in agentic retrieval, costs are a function of the number of reasoning steps and tokens consumed, which depend on query complexity and dynamic agent behavior.

Methodology in Plain English

The authors start from the observation that LLMs can reason but cannot scan millions of documents, while retrievers can scan millions of documents but cannot reason. Their agent gives the LLM a system prompt describing the goal of finding all relevant documents, hands it the original query plus the dense retriever's initial results in the first user message, and asks it to find any documents that initial pass missed. Showing the initial results serves two purposes: it avoids redundant searches and lets the agent infer what kind of documents live in the corpus so it can adapt its strategy.

The agent then loops, calling Retrieve with a query and a value of k, optionally calling Think to reason or plan, and finally calling Final_Results with an ordered list of document IDs and a justification.

Evaluation covers two collections. ViDoRe v3 is an enterprise document retrieval benchmark spanning six languages and 10 domains, including finance, pharmaceutical, and government energy reports, with seven query types such as extractive, multi-hop, and numerically reasoning queries. BRIGHT is a reasoning-intensive text retrieval benchmark measuring logical deduction, code understanding, and mathematical reasoning across domains. Agent configurations pair Opus 4.5 or gpt-oss-120b with colembed-vl-8b-v2 on ViDoRe v3 and llama-nv-reasoning-3b on BRIGHT, with llama-nemotron-1b also reported on both datasets. Standard retrieval using the same embedding models serves as the baseline.

Why This Matters

The paper argues, with measurements rather than assertions, that semantic similarity alone is insufficient for emerging retrieval workloads, and it quantifies both the benefit and the price of moving to an agentic alternative. It also shows that a single unmodified pipeline can be competitive on two very different leaderboards, which is a meaningful result for anyone who has watched a specialized retriever degrade the moment it leaves its training domain.

Real-world applications the paper points to:

  • Retrieval-augmented generation (RAG) and DeepResearch-style workflows that need to explore large unstructured corpora.
  • Enterprise document search across finance, pharmaceutical, and government energy reports, including multi-hop and numerical-reasoning queries.
  • Math and science assistance, where an agent given a problem should retrieve useful theorems even when the problem statement shares little surface vocabulary with them.
  • Tool and API specification lookup, where the agent must reason about real-world system dynamics rather than match keywords.

Industry relevance: the cost findings are directly operational. A jump from 0.67 seconds and no LLM inference to roughly 107 seconds and 764.1K input tokens per query changes the economics of serving search at scale, and the paper's MCP-to-in-process migration is a concrete engineering lesson for teams building agent tooling. The observation that open-source models lag noticeably behind proprietary ones on this task points to a clear commercial gap.

Future Directions

  • Cost-efficient retrieval agents. The authors state that further work is needed before retrieval agents can be deployed at scale, and identify LLM inference overhead as the main limitation.
  • Cost-control mechanisms analogous to approximate nearest neighbor search. Just as ANN variants let practitioners trade index accuracy for speed, the paper calls for ways for users to tune token and reasoning costs to their application.
  • Better open-source models for agentic retrieval. The task is described as clear, well defined, and requiring specific skills such as iterative exploration, which the authors take as motivation for models that are more effective than current open-source options and more efficient than proprietary ones like Opus 4.5.
  • Task-specific infrastructure and tool-calling patterns. The paper suggests parallel tool calls, meaning multiple search attempts in a single turn, as a better fit for retrieval than the sequential one-at-a-time pattern common in general agents, since it can reduce steps while giving more context for the next search.

Target Audience

Information retrieval researchers studying reasoning-intensive and enterprise retrieval benchmarks; engineers building RAG, DeepResearch, or agentic search systems who need to weigh accuracy against latency and token budgets; and teams deciding whether to adopt agent frameworks such as ReAct or MCP for high-volume retrieval workloads.

Authors’ abstract

Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.

Read the original paper