Skip to content
AI.info

Research

BioMedSearch: A Multi-Source Biomedical Retrieval Framework Based on LLMs

Overview Research area: Natural Language Processing / biomedical information retrieval and retrieval-augmented generation (RAG) with large language models. Technical level: Intermediate. The paper ass

BioMedSearch: A Multi-Source Biomedical Retrieval Framework Based on LLMs
arXiv
2510.13926
Published
2025-10-15
Authors
Congying Liu, Xingyuan Wei, Peipei Liu, Yiqing Shen, Yanxu Mao, Tiehan Cui

AI summary

Overview

Research area: Natural Language Processing / biomedical information retrieval and retrieval-augmented generation (RAG) with large language models.

Technical level: Intermediate. The paper assumes familiarity with LLMs, retrieval-augmented generation, search agents, and embedding-based semantic similarity, but the framework itself is a modular system built from off-the-shelf components rather than a new model architecture.

One-sentence scope: The paper introduces BioMedSearch, a training-free framework that decomposes biomedical queries into sub-queries and keywords, retrieves from literature databases, protein databases and web search engines, and generates a structured report, evaluated on a newly constructed 3,000-question multi-level benchmark called BioMedMCQs.

What This Paper Is About

Biomedical questions often require synthesizing knowledge across gene regulation, protein function, disease mechanisms and clinical outcomes, and general-purpose LLMs frequently hallucinate in this domain — fabricating protein functions, interactions and structural details, or misattributing physiological processes to disease mechanisms.

The stated goal is to build a retrieval framework that connects LLMs to authoritative biomedical sources (literature databases, UniProt protein data, real-time web search) so that answers are grounded in retrieved evidence, and to build an evaluation benchmark that measures retrieval quality and multi-level reasoning rather than only factual memory.

Key Contributions

  1. BioMedSearch framework: A multi-source biomedical search agent that performs real-time online retrieval by integrating protein databases, biomedical knowledge bases and general web search, without any additional training.

  2. Biomedical search planner and retrieval executor: A planner that decomposes sub-queries into fine-grained keywords and organizes retrieval paths, plus an executor that performs retrieval across general literature and web sources, protein information databases, and specialized biomedical literature repositories according to those planned paths.

  3. BioMedMCQs benchmark: Described as the first benchmark for biomedical search queries, built by randomly generating biomedical research topics to simulate real-world user intents. It contains 3,000 multiple-choice questions across three difficulty levels (1,000 per level) and is designed to evaluate retrieval effectiveness, biomedical reasoning, and evidence–answer alignment.

  4. Empirical evaluation across five LLMs and multiple baselines: Experiments with ChatGPT-4.1, DeepSeek-R1, Llama-4, Gemini-2.5 and Qwen3 as backends, compared against general search agents (MindSearch, DeepSearcher, MMSearch, PaSa) and biomedical RAG methods (Self-BioRAG-7B, MedRAG).

Main Findings

  • Consistent gains across all levels and models: BioMedSearch outperforms all baselines at every reasoning level and across all model architectures tested. Reported average accuracy (abstract) rises from 59.1% to 91.9% at Level 1, from 47.0% to 81.0% at Level 2, and from 36.3% to 73.4% at Level 3.

  • Level 1 (mechanistic identification): BioMedSearch achieves a highest accuracy of 94.2% with Gemini-2.5, 93.8% with ChatGPT-4.1, and 92.1% with Llama-4. MedRAG reaches a maximum of 87.2% (ChatGPT-4.1) and PaSa 85.9% (Llama-4), which the paper describes as a 6–8% improvement across models.

  • Level 2 (non-adjacent semantic integration): Llama-4 reaches 84.9%, followed by ChatGPT-4.1 at 83.3% and Gemini-2.5 at 82.6%. MedRAG reaches up to 80.5% and MMSearch 71.3%, which the paper frames as a 2–13% advantage across models.

  • Level 3 (temporal causal reasoning, the most challenging level): BioMedSearch reaches 78.0% with ChatGPT-4.1, 75.5% with Gemini-2.5 and 74.1% with Llama-4. MedRAG reaches a maximum of 72.8%, DeepSearcher 71.2%, and PaSa 70.9%.

  • Average comparison (Figure 4): At Level 1 BioMedSearch averages 91.8%, versus MedRAG at 81.2% and PaSa at 79.3%. At Level 2 BioMedSearch reaches 81.02%, versus MedRAG 74.80% and PaSa 70.74%. At Level 3 BioMedSearch reaches 73.42%, versus MedRAG 66.38% and PaSa 61.2%. (Note: the abstract states the Level 1 average as 91.9%, while Figure 4 reports 91.8%; both figures are as given in the paper.)

  • Baseline (no retrieval) performance is much lower: ChatGPT-4.1 scored 63.5 / 53.8 / 49.5, Llama-4 65.0 / 57.4 / 50.6, Gemini-2.5 61.6 / 56.1 / 42.7, Qwen3 54.8 / 41.8 / 32.8, and DeepSeek-R1 50.7 / 39.7 / 25.9 across Levels 1–3. Self-BioRAG-7B scored 58.1 / 47.8 / 35.4.

  • Ablation — removing literature retrieval hurts most: Under the "w/o literature" setting, DeepSeek-R1's Level 3 accuracy falls from 67.7% to 56.6% and Gemini-2.5's falls from 75.5% to 71.0%, indicating literature retrieval is critical for complex biomedical questions.

  • Ablation — removing keyword extraction also degrades performance: Under "w/o keywords", Qwen3's Level 3 accuracy drops from 71.8% to 68.2%, suggesting keyword decomposition improves retrieval precision and reduces distracting information.

  • Full model beats ablated variants for every LLM: The paper states the full setting consistently outperforms both ablated variants across all LLMs and difficulty levels.

Methodology in Plain English

The framework has three stages and is used without any fine-tuning of the underlying models.

Stage 1 — Planning. A user query in natural language is broken into bounded sub-queries along biomedical semantic dimensions such as developmental effects, endocrine regulation, clinical phenotypes and molecular mechanisms. Each sub-query is then further broken into biomedical keywords along finer dimensions. All extracted keywords are combined into a directed acyclic graph, which defines retrieval paths and assigns an appropriate retrieval tool to each sub-query.

Stage 2 — Retrieval. Three modules run in parallel depending on what the plan calls for:

  • Literature: queries PubMed, PMC and ScienceDirect simultaneously, retrieving up to 100 results from each. If a database returns nothing, the query is rebuilt using the three most salient keywords. Two filters follow: at least 80% of sub-query keywords must appear cumulatively in title and abstract, and the remaining papers are ranked by cosine similarity against the sub-query using PubMedBert embeddings, keeping the top 10.
  • Protein information: when a gene name is detected, the system performs a UniProt ID lookup, then extracts functional annotations, interactions and sequence data. For queries about 3D structure or spatial conformation, it calls AlphaFold using the UniProt ID to generate and visualize a predicted structure in PDB format.
  • Web search: queries DuckDuckGo, Google and Brave independently per sub-query, taking the top 100 results from each (titles, URLs, summaries), then filters with an LLM plus PubMedBert embeddings to keep only high-relevance pages above a relevance threshold.

Stage 3 — Report generation. The LLM evaluates each sub-query's retrieval results, extracts content that answers it to form sub-answers, links interrelated sub-answers supporting the original intent, and assembles them into a structured research report with background, key findings and source references.

Evaluation setup. For answering, the system extracts coherent paragraphs from the markdown reports as context, vectorizes each question and option with PubMedBert, and selects the top-k most relevant paragraphs. All LLMs receive the same prompt template and are explicitly forbidden from using external knowledge, so answers must come solely from the provided context. The implementation uses Python 3.10.16, and all LLMs are accessed through official APIs with default configurations and no fine-tuning.

Benchmark construction. BioMedMCQs was generated by having an LLM randomly generate a biomedical research topic, decompose it into sub-topics, extract keywords, retrieve approximately 600 articles from PMC, PubMed and ScienceDirect, and filter to the top 300 per topic using keyword matching and PubMedBert semantic validation. Questions span Level 1 (fundamental causal mechanisms involving a single factor), Level 2 (non-adjacent semantic integration across sentence boundaries), and Level 3 (hierarchical reasoning under temporal dependency and feedback loops, e.g., drug metabolism pathways and inflammation regulatory networks). Each question was manually reviewed by biomedical domain experts for scientific validity and clarity.

Why This Matters

The paper targets a practical failure mode of LLMs in science: confident but fabricated biomedical content. By grounding answers in UniProt, AlphaFold, PubMed/PMC/ScienceDirect and live web results, it offers a way to make biomedical question answering more traceable and verifiable, and it argues that existing biomedical RAG methods lack real-time web access and are mostly evaluated on general datasets like MedQA, MedMCQA and MMLU, which limits their coverage of open-ended, real-world scenarios.

Real-world applications:

  • Drug target identification and design: combining functional annotations, sequence data and predicted 3D structures from UniProt and AlphaFold to support research decisions.
  • Clinical diagnosis support: retrieval-augmented reasoning over disease mechanisms and phenotypes, particularly at the paper's higher reasoning levels.
  • Gene regulatory mechanism research: answering questions about gene knockout effects, developmental phenotypes and hormonal regulation, as in the paper's own example query about cyp17a1 knockout in zebrafish.
  • Capturing patient-reported and emerging information: real-time web search surfaces expert opinions and patient experience with treatment, recovery and medication responses that static literature databases would miss.

Industry relevance: the approach is training-free and built on public APIs and openly available databases, which lowers the barrier to deployment for research tools, pharmaceutical informatics pipelines, and clinical decision-support products that need auditable sources rather than ungrounded generation.

Future Directions

  • Refining the task graph structures and retrieval paths used to schedule multi-source retrieval.
  • Incorporating domain-specific knowledge graphs to supplement the current literature, protein and web sources.
  • Adding multimodal data such as biomedical images and pathways to support reasoning that current text-only retrieval cannot address.
  • Open questions the work raises: how to generalize beyond the three reasoning levels defined in BioMedMCQs, how the framework behaves outside the multiple-choice format (e.g., on free-form research report generation, which the paper does not quantitatively evaluate), and how retrieval quality would scale to query domains where no UniProt or AlphaFold analogue exists.

Target Audience

This paper is most useful to researchers and engineers working on retrieval-augmented generation, LLM search agents, and biomedical or scientific question answering, as well as teams building clinical or pharmaceutical knowledge tools. It also suits readers interested in evaluation methodology, since the BioMedMCQs benchmark and its three-level reasoning taxonomy are a substantial part of the contribution. Readers with no background in LLM retrieval systems will need some familiarity with RAG concepts to follow the architecture, but the method description itself is largely procedural and accessible.

Authors’ abstract

Biomedical queries often rely on a deep understanding of specialized knowledge such as gene regulatory mechanisms and pathological processes of diseases. They require detailed analysis of complex physiological processes and effective integration of information from multiple data sources to support accurate retrieval and reasoning. Although large language models (LLMs) perform well in general reasoning tasks, their generated biomedical content often lacks scientific rigor due to the inability to access authoritative biomedical databases and frequently fabricates protein functions, interactions, and structural details that deviate from authentic information. Therefore, we present BioMedSearch, a multi-source biomedical information retrieval framework based on LLMs. The method integrates literature retrieval, protein database and web search access to support accurate and efficient handling of complex biomedical queries. Through sub-queries decomposition, keywords extraction, task graph construction, and multi-source information filtering, BioMedSearch generates high-quality question-answering results. To evaluate the accuracy of question answering, we constructed a multi-level dataset, BioMedMCQs, consisting of 3,000 questions. The dataset covers three levels of reasoning: mechanistic identification, non-adjacent semantic integration, and temporal causal reasoning, and is used to assess the performance of BioMedSearch and other methods on complex QA tasks. Experimental results demonstrate that BioMedSearch consistently improves accuracy over all baseline models across all levels. Specifically, at Level 1, the average accuracy increases from 59.1% to 91.9%; at Level 2, it rises from 47.0% to 81.0%; and at the most challenging Level 3, the average accuracy improves from 36.3% to 73.4%. The code and BioMedMCQs are available at: https://github.com/CyL-ucas/BioMed_Search

Read the original paper