Research
Self-Evolving Search Index
Self-Evolving Search Index Overview Research area: Information retrieval (cs.IR), specifically index optimization and the application of self-evolving agent frameworks to retrieval indices. Technical

- arXiv
- 2609.19656
- Published
- 2026-09-17
- Authors
- Sangam Lee, Wonjae Lee, Sunghwan Kim, Deogyong Kim, Jaehoon Kim, Daye Nam, SeongKu Kang, Dongha Lee
AI summary
Self-Evolving Search IndexOverview
Research area: Information retrieval (cs.IR), specifically index optimization and the application of self-evolving agent frameworks to retrieval indices.
Technical level: Advanced. The paper assumes familiarity with retrieval evaluation metrics (nDCG@10), sparse and dense retrievers, LLM-based optimization pipelines, and multi-stage agentic loops.
Scope: This paper proposes Self-Index, a framework in which a search index autonomously diagnoses its own retrieval failures, selectively revises its index keys, validates those revisions, and proactively generates new queries to guide further evolution, evaluated across natural language, code, math, and table corpora with three retrievers and two downstream agent applications.
What This Paper Is About
Retrieval quality depends on the index keys — the representations used to match and rank documents for a query — but no single index optimization strategy works well across all corpora and retrievers, so an index must be evolved to fit its own retrieval environment. Today that evolution is human-driven: people must diagnose why retrieval failed, redesign the optimization strategy, and reprocess the entire index, which is slow and expensive to repeat. Self-Index aims to remove the human from this loop by letting the index diagnose, revise, and validate its own keys, and by having a query simulator surface retrieval demands the index has not yet been optimized for.
Key Contributions
-
The Self-Index framework. A framework that enables an index to self-evolve by autonomously diagnosing and refining its representations across diverse retrieval demands, with an Optimizer that runs a three-stage loop (Self-Diagnosis, Self-Revision, Self-Validation) and only updates the index with revisions that pass validation.
-
Proactive self-evolution via a Query Simulator. Beyond reacting to queries it receives, Self-Index adds a Query Simulator that performs Self-Exploration by generating plausible new retrieval demands grounded in sampled documents and filtering them by Jaccard-similarity-based Dissimilarity, so the index evolves beyond the queries already available for optimization.
-
Consistent retrieval gains across diverse retrievers and corpora. Self-Index achieves the highest average nDCG@10 for every retriever on BRIGHT and on the table retrieval benchmarks, outperforming existing index optimization methods, which sometimes only marginally improve or even degrade performance.
-
Benefits extended to downstream applications. The paper shows gains for search agents on BrowseComp-Plus (higher answer accuracy and evidence recall, fewer search calls, generally lower calibration error) and for agent memory systems on LongMemEval-V2.
Main Findings
-
Highest average nDCG@10 for every retriever on BRIGHT and table retrieval. On BRIGHT, Self-Index reaches a final average of 20.4 with BM25 (versus 14.5 for the base index, +40.4%), 21.8 with BGE (+57.0%), and 26.1 with Qwen3-Emb-8B (+38.8%). On the table retrieval datasets, Self-Index reaches 49.5 with BM25 (+49.1%), 53.1 with BGE (+17.9%), and 57.0 with Qwen3-Emb-8B (+16.2%).
-
Competing methods are inconsistent. Doc2Query improves table retrieval with BM25 (+37.2%) but degrades performance on the code corpora; on dense BGE it falls below the base index on BRIGHT (−6.2%) and on tables (−3.2% with BGE, −7.1% with Qwen3-Emb-8B). SPIKE and RL-Index yield only marginal improvements in some settings.
-
Self-Validation is the most important component. On BRIGHT (nDCG@10 averaged across retrievers), removing validation drops NL to 16.1 (−10.5), Code to 13.0 (−9.1), and Math to 11.9 (−5.3) — below the base index in every corpus type. Removing any single criterion also hurts: without Faithfulness 21.0/18.1/15.6; without Specificity 20.6/16.1/13.0; without Separation 23.9/16.8/12.1 (NL/Code/Math). The full Self-Index scores 26.6/22.1/17.2.
-
Co-retrieval profiles matter for diagnosis. Removing the co-retrieval profile C_k from Self-Diagnosis lowers NL to 19.7 (−6.9), Code to 16.7 (−5.4), and Math to 15.9 (−1.3).
-
The Dissimilarity filter in Self-Exploration helps. Removing it lowers NL to 22.2 (−4.4), Code to 18.5 (−3.6), and Math to 14.8 (−2.4) with the number of optimization queries held fixed.
-
The Optimizer also works with externally sourced queries, but simulator queries are better. Replacing the Query Simulator with ReasonIR HQ queries still improves average nDCG@10 over the base index across all corpus types, but queries from the Query Simulator yield larger gains in every corpus type.
-
Search agents improve in effectiveness and efficiency. On BrowseComp-Plus, GPT-OSS-120B with BM25 improves accuracy from 31.08 to 58.92 (+89.53%), recall from 37.54 to 65.84 (+75.39%), search calls from 21.16 to 16.62 (−21.46%), and calibration error from 42.05 to 26.20 (−37.69%). The same pattern holds for GPT-5.4-nano, Gemini-3.7-Flash, and Kimi-K2.5 under both BM25 and Qwen3-Emb-8B: Self-Index achieves the highest answer accuracy and evidence recall in every setting, while SPIKE decreases accuracy or recall in some cases and increases search calls in some cases.
-
Robustness to corpus scale. When the BrowseComp-Plus corpus expands from 100K to 200K and 400K documents, Self-Index maintains stable answer accuracy while slightly reducing online cost per query relative to its 100K-document baseline, whereas SPIKE's accuracy stays below its own baseline and its cost rises; DCI declines sharply in accuracy and rises substantially in cost.
-
Agent memory utilization improves. On LongMemEval-V2, Self-Index raises overall accuracy for every evaluated memory system: Query→Slice from 0.415 to 0.472 (+13.9%), Query→Slice+Notes from 0.448 to 0.503 (+12.4%), and AgentRunbook-R from 0.532 to 0.581 (+9.2%). Gains appear consistently in the static, dynamic, and workflow abilities; gotchas is the exception (e.g., 0.276 unchanged for Query→Slice and AgentRunbook-R, 0.310 unchanged for Query→Slice+Notes).
-
Progressive rather than one-shot evolution. Across BRIGHT corpus types, Self-Index surpasses SPIKE in early iterations and continues to gain in later iterations; an example from the Biology domain shows the gold document starting with the full document text as its only key, then being organized into aspect-specific keys, and finally receiving a key that explicitly represents the query-relevant knowledge, moving the document near the top of the ranking.
Methodology in Plain English
Task setup. A corpus D = {d₁, …, d_N} is indexed by assigning each document d a set of index keys K(d); the whole index is K = ⋃ K(d). A document's score for query q is the maximum relevance over its keys, s(q, d) = max_{k ∈ K(d)} rel(q, k). Self-Index revises keys only and leaves the underlying corpus unchanged. Each document's original text is kept as a fixed key that is never revised or validated; initially K₀(d) = {d}. The maximum-over-keys aggregation removes the need to tune weights between representations. The key budget is m_max = 10.
The Optimizer loop. Each iteration takes a query set Q and runs three stages:
- Self-Diagnosis — for each query, invoke the retriever on the current index and collect results. For each retrieved key k, build a co-retrieval profile C_k recording which keys from other documents were retrieved alongside k and how often. In the spirit of pseudo-relevance feedback, this uses retrieval outcomes as feedback instead of human relevance annotations, and the Optimizer diagnoses what the current keys fail to expose.
- Self-Revision — targeted document key sets are revised once per iteration as a whole (rather than key-by-key) so that new keys do not duplicate information already covered by other keys of the same document, producing a proposed set K′(d).
- Self-Validation — every generated key in K′(d) \ {d}, including retained keys, is checked against three criteria: Faithfulness (supported by d without distortion), Specificity (emphasizes knowledge specific to d rather than broadly shared content), and Separation (lower maximum relevance to observed competing keys than the current key set). Failing keys are dropped. If a newly proposed key passes all three, K(d) becomes the fixed original-text key plus all passing generated keys; otherwise it stays unchanged.
Query Simulator. It samples documents from D and generates queries reflecting plausible retrieval demands grounded in those documents, then applies a Jaccard-similarity Dissimilarity filter to limit lexical overlap with queries already used for optimization and with queries already accepted in the current simulation step. Retained queries go to the Optimizer.
Experimental setup. Evaluation uses BRIGHT (natural language: Biology, Earth Science, Economics, Psychology, Sustainability; code: Robotics, Stack Overflow, LeetCode, Pony; math: AoPS, TheoremQA, Theorems) plus three table datasets (Spider 2.0, FIBEN, BEAVER), with nDCG@10 as the retrieval metric. Downstream evaluation uses BrowseComp-Plus (search agents) and LongMemEval-V2 (agent memory), following the official protocols. Baselines are Doc2Query, SPIKE, and RL-Index, plus EnrichIndex on table retrieval. Retrievers are BM25, BGE-Large, and Qwen3-Embedding-8B. Agent backbones on BrowseComp-Plus are GPT-OSS-120B, GPT-5.4-nano, Gemini-3.7-Flash, and Kimi-K2.5; LongMemEval-V2 uses Qwen3.5-9B for both the memory controller and the downstream reader. The Optimizer and Query Simulator are built on Qwen3.6-35B-A3B, and all baselines were reproduced with the same backbone LLM for fairness. In the main experiments, every index is evolved solely with Query Simulator queries, so evaluation queries remain unobserved during optimization. The paper notes a released code link.
Why This Matters
Impact on research. The paper extends the self-evolving paradigm — previously applied to reasoning models and agentic systems — to index optimization, showing that the index itself can be the object of self-improvement. It also argues that no fixed optimization strategy generalizes across retrieval environments, which reframes index construction as a continuous, environment-specific adaptation problem rather than a one-time preprocessing step. Its ablations give concrete evidence about which design choices (validation criteria, co-retrieval context, exploration diversity) actually drive those gains.
Real-world applications:
- Search agents and deep research assistants that must answer complex questions through many retrieval calls — fewer search calls with higher accuracy lowers the API cost of running them.
- Agent memory systems that store and reuse past interactions — improving the retrieval keys lets agents surface useful earlier interactions without changing how the memory is constructed, which the paper shows works across raw trajectory slices, slice-plus-notes designs, and a dedicated system such as AgentRunbook-R.
- Enterprise and domain-specific retrieval over heterogeneous corpora — natural language documents, code repositories, and tables — where a single hand-designed indexing strategy cannot be tuned per environment.
- Large-scale web or document search — the corpus-expansion experiment from 100K to 400K documents suggests the approach preserves answer quality and cost efficiency as collections grow, where the index-free DCI alternative degrades.
Industry relevance. Because Self-Index modifies only keys and is described as largely orthogonal to how memory or corpora are constructed and organized, it can be layered on top of existing retrieval stacks. The cost argument is central for commercial deployments: index-based search has typically been cheaper online than index-free direct corpus interaction but weaker in task performance, and the paper positions Self-Index as closing that gap while further reducing cost. The work was partly supported by Samsung Research, and code is released.
Future Directions
- Reducing optimization-time cost. Self-Index automates the human effort of index evolution, but the paper does not report the computational cost of running the Optimizer and Query Simulator with its Qwen3.6-35B-A3B backbone; a natural next question is how the number of iterations and LLM calls trade off against retrieval gains.
- Scaling beyond the tested corpus sizes and key budget. The robustness study covers 100K, 200K, and 400K documents and the algorithm sets a key budget of m_max = 10; whether the same behavior holds at much larger scales or with different key budgets is not reported.
- Handling abilities that resist retrieval-side improvement. On LongMemEval-V2, gotchas did not improve under Self-Index, and the paper attributes this to its dependence on how the memory system processes past interactions into memory contents rather than on retrieval; how to serve such abilities remains open.
- Fully proactive evolution. The paper notes that the Optimizer is still fundamentally reactive and that the Query Simulator is what extends it toward proactive behavior; better or more targeted demand exploration is an open direction. Beyond the stated hope that Self-Index "will contribute to future works that effectively support users," the paper does not enumerate specific planned next steps.
Target Audience
This paper is most useful to information retrieval researchers working on index representation and retrieval-augmented generation; engineers building search agents, deep research systems, or agent memory; practitioners tuning retrieval over mixed corpora (text, code, tables) who cannot hand-design a per-environment indexing strategy; and readers interested in self-evolving or self-improving system designs more broadly. It is written for an audience comfortable with retrieval benchmarks, nDCG@10, LLM-based pipeline design, and ablation-style evaluation.
Authors’ abstract
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.