Research
UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval
Overview Research area: Biomedical question answering (QA) within natural language processing, combining retrieval-augmented generation (RAG), multi-hop reasoning, and in-context learning. Technical l

- arXiv
- 2601.06974
- Published
- 2026-01-11
- Authors
- Quoc-An Nguyen, Thi-Minh-Thu Vu, Bich-Dat Nguyen, Dinh-Quang-Minh Tran, Hoang-Quynh Le
AI summary
Overview
Research area: Biomedical question answering (QA) within natural language processing, combining retrieval-augmented generation (RAG), multi-hop reasoning, and in-context learning.
Technical level: Advanced. The work assumes familiarity with large language models, retrieval pipelines, ensemble classifiers, TF-IDF/cosine similarity, and QA evaluation metrics.
Scope (one sentence): The paper describes a biomedical QA system that routes questions into "direct" or "sequential" tracks, decomposes only the sequential ones into sub-questions, retrieves context from search-engine snippets plus Wikipedia, and reports an Exact Match of 0.84 and second place on the BioCreative IX MedHopQA shared task leaderboard.
What This Paper Is About
Biomedical questions often require pulling together facts from several sources in a step-by-step chain, and LLMs struggle with this both because of domain complexity and because decomposition itself can introduce hallucinated or irrelevant sub-questions. The authors build a system for the BioCreative IX MedHopQA shared task that applies multi-hop decomposition only when a question actually needs it, leaving simpler "direct" questions untouched to avoid unnecessary errors and overhead. Their goal is to improve answer accuracy on both short-answer Exact Match and concept-level evaluation.
Key Contributions
- A QA framework that dynamically distinguishes between direct and sequential questions and applies a different reasoning strategy to each, so that only sequential questions are decomposed.
- A lightweight machine learning classifier (a stacking ensemble of Random Forest and XGBoost with a Logistic Regression meta-classifier) that identifies sequential questions before any decomposition, reducing unnecessary complexity and shielding direct questions from decomposition errors.
- A retrieval and grounding strategy that combines multiple sources — search-engine snippets and crawled Wikipedia content ranked by TF-IDF and cosine similarity — with in-context learning to supply rich context to the answering LLM.
- A Wikipedia-based answer normalization step that uses the Wikipedia API to map a generated short answer to a page heading title, improving string-level Exact Match scores.
Main Findings
-
Best submission score: The system's best run (Run 5, ID 302816) reached an Exact Match score of 0.840 and a Concept Level Score of 0.863, ranking second on the leaderboard. The abstract reports this as an Exact Match of 0.84.
-
Progressive gains across five submissions: Run 1 (baseline, snippets only, no normalization) scored 0.755 Exact Match and 0.812 Concept Level. Run 2 (adding Wikipedia retrieval for yes/no questions) scored 0.783 / 0.832. Run 3 (adding Wikipedia-based answer normalization) scored 0.811 / 0.836. Run 4 (adding Wikipedia retrieval for WH-questions) scored 0.829 / 0.851. Run 5 (reprocessing unanswered questions and experimenting with a stronger LLM on a subset) scored 0.840 / 0.863.
-
Wikipedia retrieval helps when targeted by question type: Adding Wikipedia-based retrieval first for yes/no questions (Run 2) and later for WH-questions (Run 4) both improved scores; the paper attributes this to the LLM receiving more accurate supporting information.
-
Normalization improves Exact Match: The Wikipedia-based normalization applied in Run 3 raised the Exact Match score, with the authors stating it successfully identifies appropriate biomedical normalization terms.
-
Classifier performance: The question-type classifier, trained on 453 human-annotated samples with labels from two annotators and conflicts resolved by a third, achieves an accuracy of 0.89 and an F1 score of 0.83. The Logistic Regression meta-classifier assigns weights of 0.72 to Random Forest and 0.28 to XGBoost outputs.
-
System configuration: The pipeline uses the Google Custom Search API with a maximum of 10 returned results, crawls Wikipedia articles and selects the highest TF-IDF sentences up to a 300-token limit, and uses gpt-4o-mini for decomposition/sub-query generation and gpt-o3-mini for answer generation, with temperature set to 0. Run 5 additionally experimented with gpt-4o-search-preview on a subset of samples.
-
Dataset scale: Evaluation used the MedHopQA Track 1 dataset from BioCreative IX, with 45 development questions and 1000 test questions. The organizers released roughly 10,000 questions, 1,000 of which were concealed as the hidden test set, and performance is measured only on those hidden questions.
-
Not reported: The paper does not report statistical significance tests, inference latency, token cost, or separate ablation figures beyond the five-run comparison, and it does not state a numeric F1 or accuracy for the question simplification step.
Methodology in Plain English
The system works in a pipeline:
-
Simplify. If a question is attached to a preceding explanatory sentence, an LLM rewrites it into a cleaner question of fewer than 50 words, so the core inquiry is not buried under distracting description.
-
Classify. Before any decomposition, a stacking ensemble (Random Forest plus XGBoost, combined by a Logistic Regression meta-classifier) uses linguistic, structural, and embedding-based features to decide whether the question is "direct" or "sequential." This step exists because LLM decomposition can hallucinate irrelevant sub-questions.
-
Decompose — selectively. Sequential questions are broken into a chain of self-contained sub-questions, each paired with a distilled sub-query (keywords) and an initial anchor (extracted entities). Direct questions skip decomposition; only a sub-query and anchor are extracted. This is done through in-context learning with examples embedded in the prompt, and the prompt instructs the model to use the minimum number of steps needed.
-
Retrieve per hop. At each hop, the system concatenates the current sub-query with the previous hop's answer (the anchor) and sends it to a search engine. It also crawls Wikipedia, converts sentences to TF-IDF vectors, and uses cosine similarity to pick the most relevant sentences as extra context, capped at 300 tokens.
-
Answer per hop. The LLM receives the sub-question, sub-query, anchor, and retrieved context, and produces two outputs: a long answer that lets the model reason over the context, and a short answer that follows strict formatting rules.
-
Normalize. The short answer is run through the Wikipedia API; if a matching page exists, its heading title becomes the normalized answer. That normalized answer becomes the anchor for the next hop, or the final prediction if it was the last hop.
Why This Matters
Impact on research: The paper argues that selective decomposition — applying multi-hop reasoning only where needed — is a practical alternative to decomposing every question, and it shows that a lightweight classifier can serve as a gatekeeper. It also demonstrates that Wikipedia's heading titles double as useful biomedical normalization targets, and that mixing search-engine snippets with crawled Wikipedia sentences improves grounding, offering a template for other domain QA systems that cannot afford fine-tuning.
Real-world applications:
- Clinical decision support tools that must answer chained questions such as identifying a condition, then a gene, then the chromosome that gene sits on.
- Biomedical literature search assistants that need to combine evidence scattered across multiple documents.
- Patient-facing or clinician-facing question answering over trusted encyclopedic and medical reference content.
- Question routing in medical information systems, where the same classifier that separates direct from sequential questions could triage queries to cheaper or more expensive pipelines.
Industry relevance: The approach requires no model fine-tuning and runs on commercially available LLMs, which makes it attractive to health-tech and pharma-informatics teams that want retrieval-grounded accuracy within budget. The layered design — cheap classifier first, expensive reasoning only when necessary — is directly relevant to cost and latency management in production QA systems.
Future Directions
- Determining whether the question-type classifier generalizes beyond the 453 annotated samples used for training, and whether it would hold up on other biomedical QA datasets.
- Isolating the individual contribution of each component (simplification, classifier, Wikipedia retrieval per question type, normalization, stronger LLM) with a controlled ablation, since the five submitted runs bundle multiple changes at once.
- Measuring cost, latency, and token usage for the multi-hop pipeline, which the paper does not report, especially given that Run 5 involved a stronger LLM.
- Replacing the search-engine-plus-TF-IDF retrieval with other sources or re-rankers, and testing whether the Wikipedia-heading normalization strategy transfers to other knowledge bases or languages.
Target Audience
This paper is most useful to NLP researchers and graduate students working on biomedical or multi-hop question answering, RAG pipelines, and prompt engineering. It also serves applied machine learning engineers and health-informatics practitioners who need a reproducible, no-fine-tuning recipe for retrieval-grounded QA, as well as shared-task participants interested in how leaderboard submissions were structured and sequenced.
Authors’ abstract
Biomedical Question Answering systems play a critical role in processing complex medical queries, yet they often struggle with the intricate nature of medical data and the demand for multi-hop reasoning. In this paper, we propose a model designed to effectively address both direct and sequential questions. While sequential questions are decomposed into a chain of sub-questions to perform reasoning across a chain of steps, direct questions are processed directly to ensure efficiency and minimise processing overhead. Additionally, we leverage multi-source information retrieval and in-context learning to provide rich, relevant context for generating answers. We evaluated our model on the BioCreative IX - MedHopQA Shared Task datasets. Our approach achieves an Exact Match score of 0.84, ranking second on the current leaderboard. These results highlight the model's capability to meet the challenges of Biomedical Question Answering, offering a versatile solution for advancing medical research and practice.