Research
Follow the Entities: A Corpus Map for Agentic Search
Overview Research area: Natural Language Processing, specifically retrieval-augmented generation (RAG), LLM agents, and agentic search over large document collections. Technical level: Advanced. The p

- arXiv
- 2609.37226
- Published
- 2026-09-29
- Authors
- Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam
AI summary
Overview
Research area: Natural Language Processing, specifically retrieval-augmented generation (RAG), LLM agents, and agentic search over large document collections.
Technical level: Advanced. The paper assumes familiarity with RAG, agentic multi-step retrieval, entity linking and cross-document coreference resolution, and retrieval evaluation metrics.
Scope: The paper introduces CorpusMap, an offline-built, entity-centric navigation layer that represents recurring entities as Entity Pages linked to every document mentioning them, and shows across 7 models and 3 benchmarks that it improves answer quality and evidence discovery for LLM agents while using fewer input tokens than raw-corpus agentic search.
What This Paper Is About
Answering questions over large document collections often requires connecting evidence spread across several documents, such as an approval in one, requirements in another, and a status update in a third. Current LLM agents handle this by iteratively searching the full corpus, but because the corpus is presented only as a flat collection of files, a relevant document gives no hint of how it relates to others, so the agent must rediscover those relationships on every query, frequently missing complementary evidence while spending many extra tokens. The paper's goal is to build a persistent, reusable navigation layer that makes cross-document relationships explicit by anchoring them on the recurring entities mentioned in the documents themselves.
Key Contributions
-
CorpusMap, an entity-centric navigation layer. CorpusMap represents each recurring entity (such as a person, project, or incident) as an Entity Page that consolidates what the documents state about it, tags each fact with its source document, lists the names under which the entity appears, and links to every document that refers to it. Formally, the map is a bipartite graph of entity nodes and document nodes, and only entities linked to at least two documents are retained, since only these provide reusable navigational paths. Documents linked to no retained entity remain in the corpus as isolated nodes, reachable through raw-corpus search.
-
An offline construction protocol (Algorithm 1). The map is built in four stages: cataloging, which induces from the corpus itself a catalog of entity types (each with a name, definition, identity criteria, and observed examples); extraction, which identifies, types, and groups entity mentions per document; resolution, which grounds each name in its source occurrence and processes each document in turn against a shared registry, deciding to LINK an observation to an existing entry, ADD a new entity, or leave it UNRESOLVED; and rendering, which keeps only cross-document entities and renders each neighborhood as an Entity Page. Because construction depends only on the corpus and not on any query, the links are shared across queries rather than rediscovered at inference time.
-
A navigation procedure that reuses the same tools as the raw corpus. Each Entity Page is stored as a file alongside the raw documents and lists the file paths of its linked documents. The agent receives, along with the question, the file paths of documents linked to the Entity Pages relevant to the question, and can move in either direction — from a page to a listed document, or from a document by searching for the pages that list it.
-
Extensive empirical validation and practical analyses. The paper evaluates CorpusMap on 7 models, 3 benchmark datasets, and 5 comparison methods, and additionally examines amortized construction cost, reuse of a map built by one LLM with a different answering LLM, incremental map updates, construction with off-the-shelf deterministic components, scaling to larger corpora, and an open-weight model.
Main Findings
-
Best answer and retrieval quality across the board. CorpusMap consistently achieves the best answer and retrieval quality across all benchmarks and LLMs, with significant overall gains over every baseline, at a lower average cost per query than raw-corpus agentic search.
-
Size of the gains. Compared with raw-corpus agentic search, CorpusMap improves overall quality by 6.4 to 11.7 points while using 34% to 57% fewer input tokens on average.
-
Overall Quality and relative token ratios (Table 1). On the dataset-balanced Overall metric, CorpusMap scores 72.55 with GPT-5.5 (versus 66.11 for Raw Corpus, relative tokens 0.43×), 65.58 with GPT-5.6 Luna (versus 53.85, 0.65×), 72.51 with GPT-5.6 Terra (versus 65.75, 0.66×), and 74.68 with GPT-5.6 Sol (versus 67.69, 0.43×).
-
Organizing the corpus is not automatically beneficial. The four alternative navigation layers — Document Page, Group Page, LLM Wiki, and Corpus2Skill — do not consistently improve over Raw Corpus, indicating that simply adding a navigation layer does not guarantee improvement. CorpusMap, by contrast, outperforms all four.
-
Largest token savings on the most expensive models. CorpusMap improves quality while reducing tokens, with the largest savings for GPT-5.5 and GPT-5.6 Sol, the two most expensive LLMs.
-
Narrows the gap to an idealized oracle. CorpusMap substantially narrows the gap to the non-comparable Gold Documents (Oracle) condition, which supplies only the gold documents as reference; Oracle Overall scores are 79.62 (GPT-5.5), 80.00 (Luna), 79.90 (Terra), and 79.69 (Sol).
-
Beneficial even without cross-document evidence. CorpusMap also outperforms Raw Corpus with fewer tokens on the remaining EnterpriseRAG-Bench questions, most of which are grounded in a single document (Table 5).
-
Generalizes to other model families (Table 2, EnterpriseRAG-Bench). With DeepSeek-V4-Pro, CorpusMap reaches 71.11 quality (1,043.3k tokens) versus 68.02 for Raw Corpus (1,349.4k tokens). With MAI-Thinking-1, CorpusMap reaches 45.07 (158.8k tokens) versus 34.06 for Raw Corpus (129.6k tokens). The other navigation layers again do not consistently beat Raw Corpus.
-
Beats retrieval-based approaches (Table 3, EnterpriseRAG-Bench). Under a retrieve-then-generate paradigm, CorpusMap scores 76.60 with GPT-5.5 and 73.20 with Luna, compared with BM25 (63.66 / 60.78), dense retrieval (65.05 / 58.84), HippoRAG (64.66 / 61.74), GraphRAG (47.95 / 43.46), and Raw Corpus (65.60 / 63.31).
-
A map built by one LLM transfers to others (Table 4). Even the map built by the least expensive LLM, at a one-time construction cost of no more than $74.65, improves overall quality over Raw Corpus for every answering LLM (GPT-5.5, Sol, Luna, and DeepSeek), including across model families.
-
Incremental updates work. Ordering the EnterpriseRAG-Bench corpus chronologically and extending the map over its growing prefix — reusing the existing registry and re-rendering only Entity Pages whose linked documents change — each incremental update saves a substantial fraction of the tokens of a full rebuild, quality improves overall with each update, the final map performs comparably to a full rebuild over the same documents, and questions whose evidence is newly incorporated improve after the update even though they were already searchable in the raw corpus.
-
Construction cost is amortized. Adding the one-time construction cost divided by the number of queries served to the per-query cost, CorpusMap is initially more expensive than Raw Corpus but becomes cheaper beyond a certain number of queries for every LLM.
-
Off-the-shelf construction is comparably effective. Instantiating extraction, resolution, and rendering with GLinker (open GLiNER models, with Entity Pages rendered without any LLM calls) yields a map comparably effective to the LLM-constructed map that outperforms Raw Corpus on every metric at a lower cost.
-
Advantage persists as the corpus grows and with open-weight models. Expanding the EnterpriseRAG-Bench corpus with distractor documents while keeping all gold documents, CorpusMap outperforms Raw Corpus at every corpus size with fewer tokens; the same holds for the open-weight Qwen3.8-27B over a GLinker-built map.
-
Case study (Table 10). For a question about how signing is represented in the v1 specification of a manifest, the answer requires an earlier draft and the v1 specification stored in different sources. Corpus2Skill's agent explored several branches of its topical tree but failed to locate them and concluded no such document existed. CorpusMap's agent navigated between Entity Pages and documents — opening the page of a work item that links the draft, reading the draft, searching Entity Pages with its terms, then opening the Serving Runtime page that links the v1 specification — and answered with the fields defined in v1.
-
Not reported. The provided content is truncated mid-sentence in Appendix A (at "artifacts such as Slac"), so the remaining appendices (evaluation metric details, prompt listings in Appendix C, and further analyses in Appendix B) are not summarized here.
Methodology in Plain English
The researchers start from the observation that documents are naturally written around recurring subjects — people, projects, products, incidents — and that these subjects can be identified from the documents alone, without knowing the questions in advance. They therefore build a map of the corpus before any query arrives.
Construction proceeds offline in four stages. First, an LLM proposes a catalog of the kinds of entities worth extracting for this particular corpus (for instance, "Project" or "configuration flag"), drawing candidates from several small samples of documents, synthesizing them into one catalog, then verifying and revising it. Second, each document is processed to find and type entity mentions and group mentions of the same subject within that document. Third, documents are processed one at a time against a shared registry that starts empty: each local entity is compared with plausible existing registry entries, together with evidence from the current document, and is either linked to an existing entity, added as a new entity, or left unresolved. Every link or add records a connection between a grounded entity and a document. Fourth, only entities appearing in at least two documents are kept, and each one is rendered as an Entity Page containing a brief overview, key facts each tagged with its source document, the names the entity appears under, and links to all its documents.
At inference, the pages are simply stored as files next to the raw documents, so the agent uses the same shell-style tools it already uses to read and search the corpus. Along with the question, the agent receives the file paths of documents linked to the Entity Pages relevant to that question, and decides which to read. It can then move along the links in either direction, letting one document lead to related documents in other folders instead of relying on searching for the query text again.
Evaluation follows three axes: answer quality (correctness, completeness, factuality, content), retrieval quality (document recall, context recall), and efficiency (input tokens accumulated over the full agent trajectory per question). Three benchmarks with multi-document questions are used — EnterpriseRAG-Bench (80 questions from the Project Related, Conflicting Info, and Completeness categories, over a fixed set of 2,819 documents), WixQA (79 multi-document questions from the ExpertWritten and Simulated splits, over the full 6,221-article knowledge base), and HERB (238 content-based questions, over a fixed set of 6,365 documents) — with the same LLM both constructing each method's artifacts and answering the questions.
Why This Matters
The paper argues for a shift in how agentic search is approached: instead of only improving the search policy an agent runs, improve how the corpus itself is organized. Making relationships between documents into a persistent, reusable artifact means the same cross-document structure does not have to be rediscovered for every question, which the results show improves quality while reducing token cost. The paper also reports that CorpusMap can be built without LLMs and updated incrementally as the corpus grows, making it relevant to deployed systems rather than only to controlled experiments.
Real-world applications:
- Enterprise question answering across heterogeneous sources such as project trackers, shared drives, email, and chat channels, where a single answer requires combining a status, an approval, requirements, and a decision recorded in meeting notes.
- Customer-support knowledge bases built from many separate help articles, where a question spans more than one article (the setting of WixQA).
- Software-company workspaces containing artifacts such as specifications, drafts, and issue records, where a revision and the document it supersedes live in different places (the setting of HERB, and the paper's own case study).
- Any large, growing document collection where agentic search is already used, since the map's construction cost is amortized across queries and it can be maintained incrementally.
Industry relevance: The reported token savings are largest for the most expensive LLMs tested, which bears directly on the operating cost of agentic search deployments. The finding that a map built by a cheap LLM, at a construction cost of at most $74.65, improves results for every answering LLM tested means organizations can build the map once with an inexpensive model and reuse it with stronger ones. The demonstration that the map can be built with off-the-shelf, deterministic components without any LLM calls lowers the barrier to deployment further.
Future Directions
- Permission-aware and privacy-preserving maps. The ethics statement notes that consolidating information about an entity from many documents into a single Entity Page may aggregate private or sensitive information, or expose documents to users not permitted to access them. It suggests constructing and serving the map according to the corpus's access permissions (for example, separately per permission level) and incorporating safeguards such as privacy and content filters.
- Broader evaluation across corpora, languages, and domains. The paper's evaluation covers 3 benchmarks and 7 models; whether the entity-centric map is equally effective for other corpus types, languages, or entity-type catalogs is not reported.
- Longer-horizon maintenance of the map. Incremental updates are demonstrated on one chronologically ordered benchmark corpus; how the map behaves under sustained growth, deletion, and editing of documents over long periods is left open.
- Extending the shift from searching isolated sources to navigating connected knowledge. The conclusion frames CorpusMap as a foundation for a broader change in agentic search over large document collections, leaving room for further work on what other reusable structures a navigation layer should expose.
Target Audience
Researchers and engineers working on LLM agents, retrieval-augmented generation, and search over large enterprise document collections will benefit most. It is also relevant to practitioners who deploy agentic search systems and care about both answer quality and token cost, and to researchers in entity linking, cross-document coreference, and entity resolution, since the method builds directly on that line of work. Because the paper assumes familiarity with RAG pipelines, agentic multi-step retrieval, and retrieval evaluation metrics, readers without that background will find the method section demanding, though the high-level idea of linking documents through shared entities is accessible.
Authors’ abstract
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.