Research
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Overview Research area: Information retrieval and AI-for-science — specifically, benchmarks for scientific literature search and scientific agents. Technical level: Intermediate. The retrieval concept

- arXiv
- 2610.02202
- Published
- 2026-10-01
- Authors
- Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn
AI summary
Overview
Research area: Information retrieval and AI-for-science — specifically, benchmarks for scientific literature search and scientific agents.
Technical level: Intermediate. The retrieval concepts (Recall@N, nDCG, BM25, dense retrievers, agentic search) are standard IR material, and the paper explains its own task definition clearly, but familiarity with embedding retrieval and LLM agent architectures helps.
Scope: ScholarCatalyst is a benchmark of 894 author-written research queries paired with author-judged "catalyst papers," built to test whether retrieval systems can surface prior work that actually inspired a research project rather than work that merely shares a topic.
What This Paper Is About
Great researchers can look at a half-formed question and sense which earlier paper holds the idea that unlocks it. Current retrieval systems pick papers by topical or semantic similarity, but the paper argues that an inspiring prior paper often looks dissimilar to the question and may never have been cited by the final publication. The authors build a benchmark, grounded in 184 researchers' firsthand accounts of their own projects, to measure whether AI systems can make that same judgment.
Key Contributions
-
A new retrieval task with author-provided ground truth. Given a query written before a project's key findings, a system must rank "catalyst papers" — prior work whose ideas did or could have advanced the project — from a corpus restricted to papers published before the source paper was completed.
-
ScholarCatalyst, a benchmark of 894 queries from 207 papers. 184 researchers supplied 207 core research queries (CoreQ) and 687 subfield-specific queries (SubQ), each with positives, topically related hard negatives the authors reviewed and rejected, and a rationale for every judgment.
-
A scalable automated data-collection pipeline. The pipeline needs only a source paper's arXiv ID to produce author-reviewable draft data (queries, candidate papers, rationales), which lets the benchmark refresh with newly published papers and stay ahead of model training cutoffs.
-
A systematic evaluation and failure analysis. The paper benchmarks sparse, multi-vector, dense, and agentic retrieval across CoreQ and SubQ and analyzes why systems fail, including similarity analyses, citation-intent checks, reranking experiments, and agent context ablations.
Main Findings
- Every system misses most catalyst papers. The strongest system recovers only 48% of gold papers in its top 20. General-purpose dense retrievers place 39% of author-identified inspiration papers in the top 20 for CoreQ and 51% for SubQ.
- Scientific-document retrievers do not help. SPECTER2 and OpenScholar trail the general-purpose models by at least 16 points on CoreQ and 27 points on SubQ at Recall@20, suggesting domain-specific training on citation signals does not by itself solve the task. BM25 and LateOn trail by 15–16 points on CoreQ and 18–25 points on SubQ.
- Agentic search does no better than embedding retrieval. Agentic search reaches 0.42 versus 0.48 Recall@20 for embedding retrieval, despite the agents calling that same retriever as a tool.
- Grep-style search fails badly. A grep-based agent finds only 8% of gold query-paper pairs during search, versus 46% for the tool-calling agent. The authors attribute this to candidate coverage: an agent can only interact with the corpus through its search queries.
- Even a post-cutoff model struggles. Claude Fable 5.1, whose training data may include the source papers, reaches only 0.51 Recall@20 across agent designs and still misses nearly half of the gold papers in its top 20.
- Similarity is a weak signal. Hard negatives are at least as similar to the query as positives on both lexical similarity (fraction of unique query words in the candidate's title and abstract) and semantic similarity (cosine similarity under Qwen3-Embedding-8B).
- Inspiration types are diverse. Across 663 key-inspiration rationales classified under a nine-type taxonomy, the most common are a technique the authors adapted (method), empirical findings that supported a direction (evidence), an idea extended to a new setting (generalization), and a limitation that motivated a new approach (limitation); more than one type applies to 55.8% of rationales.
- The published record hides inspiration. 43.6% of SubQ positives are not cited by the source paper, and 67.8% of those share no references with it. Among cited papers, only 46.2% of CoreQ and 34.0% of SubQ positives carry a "uses" or "extends" citation-intent label. Given the entire finished source paper, Gemini 3.1 Pro covers only 32.2% of CoreQ positives on average and finds none of the key inspirations in 28.0% of instances.
- Coverage, not recognition, is the bottleneck. In a 50 CoreQ / 50 SubQ reranking experiment, stronger rerankers gain more from gold injection (swapping missing positives into a fixed-size pool). Letting an agent read beyond abstracts changes Recall@20 by at most 0.04, and the best SubQ reading policy (0.50) still trails the retriever alone (0.55).
- Source context helps by guiding search, not by memory. For a Gemini 3.7 Flash tool-calling agent on a 203-query split, source title and abstract raise Recall@20 by 0.11 (CoreQ) and 0.02 (SubQ); web search changes it by at most 0.01; the source bibliography (containing 95% of CoreQ gold papers) raises it by 0.35. Without search tools, the model with title and abstract reaches only 0.31 CoreQ Recall@20, below corpus search alone.
- Query rewriting does not help. On 100 queries with Gemini-Embedding-2 and GPT-4.1 rewrites, single-query expansion moved Recall@20 by +0.04 (CoreQ) and −0.05 (SubQ), multi-query generation by −0.03 and −0.01, and both HyDE variants reduced recall substantially (down to 0.18 and 0.20).
- A human ceiling estimate suggests room to improve. A preliminary comparison on one paper with three author-annotators found that coauthors marking about seven positives per thread recovered 43–60% of the first author's labels at 84–88% precision, well above the best system's overall R@5 of 0.24.
- Cross-area inspiration is common. 49.5% of papers the authors credit come from outside their project's subject area, and 43.6% of subfield-specific positives were never cited in the source paper.
Methodology in Plain English
The authors asked researchers to reconstruct the thinking behind a paper they had already published. Using an automated pipeline, they parsed each source paper and its bibliography, used Gemini 3.1 Pro to draft a core research question and several subfield-specific questions along with candidate papers, retrieved uncited candidates with BM25 and Qwen3-Embedding-8B, and reranked the pooled candidates with Gemini 3.6 Flash, keeping 10 candidates per query. Authors then reviewed everything through a web interface: they could rewrite or approve the queries, mark each candidate as offering an insightful idea or not, and explain their reasoning. The pipeline was applied to 207 source papers from 184 researchers, mostly first authors of oral, spotlight, or award papers at major AI conferences in 2025–2026 (more than 100 of the 207 papers received such a distinction).
To keep the task honest, the search corpus only contains papers published before the source paper was completed. The corpus of 190,896 papers combines references resolved from the source papers with 181K additional arXiv papers from 2020–2024 in major computer-science categories. Web access is disabled. Queries are scored with Recall@N and nDCG@N. Evaluating systems are drawn from four families: BM25 (lexical), LateOn (multi-vector), seven dense retrievers (Qwen3-Embedding-4B and 8B, Gemini Embedding 2, SPECTER2, OpenScholar, ReasonEmbed, Inf-Retriever-v1-Pro), and three agents (grep, tool-calling, and deep research) whose backbones, GPT-4.1 and o3, have a June 2024 knowledge cutoff predating all source papers. Claude Fable 5.1 (June 2026 cutoff) is evaluated separately as a comparison point outside the cutoff rule.
Of the 894 finalized queries, 764 (85.5%) passed manual review unchanged; the rest were reworded to remove phrasing that gave away the eventual solution.
Why This Matters
Impact on research. The paper separates a previously unmeasured skill — sensing which prior idea a new problem needs — from the retrieval capabilities existing benchmarks test. It shows that similarity-based search, citation-graph methods, and agentic search all leave a large gap, and it gives a measurable target for training retrieval models with what the authors call an expert-level sense of which ideas matter. It also reframes the corpus-size problem: the benchmark's 191K-paper corpus is far smaller than the three million papers on arXiv, and Semantic Scholar alone indexes over 225M papers, so real-world difficulty is likely greater than the benchmark shows.
Real-world applications
- Literature review and related-work assistants that surface the non-obvious paper a researcher has not yet found, including the 49.5% of credited papers from outside the project's subject area.
- Scientific agents that take a half-formed idea and point to the prior research it needs, as the authors envisage.
- Grant, peer-review, or patent workflows where the question is which prior work genuinely informs a new proposal rather than which work shares terminology.
- Cross-disciplinary discovery tools aimed at "undiscovered public knowledge," where a method from one field applies to a problem in another.
Industry relevance. The benchmark has direct implications for teams building research copilots, literature search products, and deep research systems. Its finding that agent recall tracks backbone capability but stays near the embedding retriever's suggests that scaling agent reasoning alone will not close the gap, and that retrieval front-ends remain the limiting component. The pipeline's ability to generate new instances from an arXiv ID also makes it a maintenance-friendly evaluation target for commercial systems whose models keep advancing past the benchmark's data.
Future Directions
- Training retrievers with expert intuition. The authors explicitly call for new training recipes, since current agents only reach the corpus through similarity-based retrievers whose candidate pools are the bottleneck.
- Establishing the task's performance ceiling. The paper names the "skyline" — the performance ceiling of the task — as a fundamental open question, with only a small three-annotator comparison on a single paper as a preliminary reference point.
- Reaching beyond any individual expert. The authors set two milestones: expert-level retrieval within a field, and retrieval at that level across a breadth of fields no single researcher can follow.
- Broadening coverage and addressing hindsight. The benchmark currently covers computer-science papers from 2025–2026 with uneven representation across research areas, and its labels are retrospective author judgments subject to hindsight; extending to other fields and testing label robustness are open problems. Because the pipeline can ingest new papers, keeping the benchmark ahead of model training cutoffs is also an ongoing requirement.
Target Audience
Researchers and engineers working on information retrieval, scientific literature search, and AI-for-science agents; benchmark designers interested in author-elicited ground truth and automated annotation pipelines; and product teams building research assistants, literature discovery tools, or deep research systems who want to understand where current retrieval and agent designs fall short on questions with latent, rather than explicit, relevance.
Authors’ abstract
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.