Skip to content
AI.info

Research

MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking

Overview Research area: Biomedical natural language processing — specifically Named Entity Recognition (NER) and Entity Linking (EL), ontology/knowledge-graph harmonization, and hierarchy-aware evalua

MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
arXiv
2511.10887
Published
2025-11-14
Authors
Nishant Mishra, Wilker Aziz, Iacer Calixto

AI summary

Overview

Research area: Biomedical natural language processing — specifically Named Entity Recognition (NER) and Entity Linking (EL), ontology/knowledge-graph harmonization, and hierarchy-aware evaluation.

Technical level: Intermediate. The paper assumes familiarity with concepts such as UMLS CUIs, controlled vocabularies, BioKGs, and standard retrieval/reranking pipelines, but its central ideas are explained clearly enough for readers with general NLP background.

Scope: MedPath is a large-scale, multi-domain biomedical entity linking resource that normalizes nine expert-annotated corpora to a single UMLS version, maps concepts across up to 62 vocabularies, and adds full hierarchical (root-to-leaf) paths for concepts in 11 vocabularies.

What This Paper Is About

Biomedical entity linking datasets are fragmented: each one is anchored to a single knowledge graph or a single text domain, so models trained on one do not transfer well to another. Existing datasets also offer little ground truth for explainable models, and they are evaluated with "flat" metrics such as precision, recall and F1 that treat a near-miss link (e.g., confusing two heart diseases) the same as a completely wrong one. MedPath addresses all three problems by harmonizing nine existing expert-curated datasets into a single resource with canonical UMLS identifiers, cross-vocabulary mappings, and hierarchical path annotations.

Key Contributions

  1. Integration of nine expert-annotated datasets. MedPath unifies corpora covering clinical notes (ShARe/CLEF 2013, SNOMED-CT EL Challenge), biomedical literature (BC5CDR, NCBI Disease, MedMentions), drug-label prose (TAC 2017 ADR), and social media (CADEC, COMETA), plus Mantra GSC, totalling over 500,000 mentions and 45,000 unique concepts as stated in the abstract (513,218 mentions and 44,259 unique concepts in the detailed counts).

  2. Canonical UMLS normalization plus cross-vocabulary mapping. Every entity is normalized to a canonical UMLS Concept Unique Identifier using the UMLS 2025AA release, and each CUI is mapped to corresponding codes in up to 62 other biomedical vocabularies, resolving the semantic fragmentation of single-vocabulary datasets.

  3. Hierarchical multi-vocabulary path annotations. Concepts are annotated with full ontological paths (general to specific) for 11 biomedical vocabularies that expose a usable API or downloadable hierarchy, yielding 573,786 distinct hierarchical paths for the 44,259 unique concepts.

  4. A hierarchy-aware evaluation framework and preliminary EL benchmarks. The paper introduces exact, ancestor-based, descendant-based and hierarchy-based metrics and reports vocabulary-agnostic entity linking experiments with TF-IDF and SapBERT retrieval plus cross-encoder reranking.

Main Findings

  • Scale and breadth: The corpus comprises over 5 million tokens, 513,218 expert-annotated mentions and 44,259 unique CUIs (500,384 mentions and 43,396 unique concepts if fallback-matched examples are excluded). Annotated mentions span 126 of the 127 possible high-level UMLS Semantic Types.

  • Low cross-dataset concept overlap: The combined mapped UMLS concept count across datasets is roughly 54,000, while unique mapped CUIs number 44,259, indicating only about 20% concept overlap between source datasets.

  • Semantic group distribution: The most prominent semantic groups are Disorders and Diseases (25.5%), Drugs and Chemicals (21.6%), and concepts (21.2%), as reported in the paper.

  • Genre imbalance: Social media posts account for 20% of documents but only 6% of mentions, reflecting their short-form nature.

  • SNOMED-CT centrality: Seven of the nine datasets map over 80% of their mentions to SNOMED-CT concepts, even datasets with different native knowledge bases. SNOMED-CT is also the most structurally complex vocabulary covered: 84% of its mapped concepts have multiple inheritance paths, averaging over 20 distinct paths per concept.

  • Path depth varies by vocabulary: SNOMED-CT shows a wide spread of path lengths, NCBI contains the deepest hierarchies on average, while ICD-9, ICD-10 and MedDRA have more concentrated path lengths of 3–5 levels.

  • Retrieval results (candidate generation only): TF-IDF leads on Acc@1 (51.46% vs. 48.12% for SapBERT) and MRR@32 (0.5756 vs. 0.5594), while SapBERT leads at larger k (Acc@5 65.44% vs. 64.85%; Acc@32 73.68% vs. 72.01%).

  • Reranking substantially improves all results: TF-IDF plus reranker reaches Acc@1 80.84%, Acc@5 91.22%, Acc@32 96.36% and MRR@32 0.857; SapBERT plus reranker reaches Acc@1 79.02%, Acc@5 92.60%, Acc@32 98.76% and MRR@32 0.861.

  • Hierarchy-aware metrics reveal near-misses hidden by exact accuracy: TF-IDF's Hierarchy@1 is 68.60% despite Acc@1 of 51.46%, about a 17-point gain, and SapBERT shows a comparable 13-point gain (48.12% to 61.39%). Roughly 20% of errors are over-general (ancestor) and 20% over-specific (descendant), pointing to granularity ambiguity rather than synonym mismatch.

  • Training on the full MedPath helps: In Figure 5, a reranker trained on all MedPath datasets (macro-averaged Acc@16 per semantic class) consistently outperforms models trained in-dataset or in-domain, sometimes by a large margin.

  • Annotation fidelity trade-off: Fallback exact-match and semantic containment mapping costs 2.5% of unique mentions (2.13% exact match, 0.37% semantic containment) and 1.15% of unique concepts (0.81% exact match, 0.35% semantic containment). These examples are explicitly labelled so users can filter them out.

Methodology in Plain English

The authors built MedPath with a four-stage automated pipeline.

  1. Unification. Dataset-specific parsers ingest the nine source corpora from their native formats (BRAT, PubTator, XML, TSV) into a common JSON schema, with light text cleaning that preserves the original annotations.

  2. Canonicalization. Every concept ID is mapped from its native vocabulary to a UMLS CUI using a dictionary derived from the UMLS database. When a direct mapping is unavailable, the pipeline falls back first to exact string matching against UMLS concept names and synonyms, then to bidirectional substring containment, choosing the closest match by token overlap and length similarity.

  3. Semantic enrichment. Each CUI receives its UMLS Semantic Type (TUI) and a set of parallel concept identifiers in other biomedical vocabularies, derived from atom-level information inside UMLS.

  4. Hierarchical path extraction. From the top 25 most frequent vocabularies in the data, the authors identified 11 that have both a formal hierarchy and an accessible taxonomy (public API or downloadable files). They wrote bespoke extractors for each — respecting different relationship types such as is-a relations in SNOMED-CT or tree numbers in MeSH — that iteratively walk from each concept to its parents until a root node is reached, capturing all paths where multiple inheritance exists. Caching, filtering of inactive codes and API callback handling were used for scalability.

For evaluation, a two-stage entity linking model adapted from the X-MEN library was benchmarked. Retrieval uses either a TF-IDF vectorizer over character 3-grams or SapBERT embeddings, both indexing a unified dictionary of UMLS CUIs with their names, synonyms and lexical variants. Retrieved candidates are reranked by a cross-encoder initialized from cambridgeltl/SapBERT-from-PubMedBERT-fulltext, trained on the top-32 candidates plus the gold entity with a categorical cross-entropy loss, and selected by best validation top-1 accuracy. Hierarchical metrics check whether any top-k prediction is an ancestor, a descendant, or otherwise hierarchically related to the gold CUI in the same knowledge graph, skipping the top three levels from the vocabulary root. These hierarchical metrics could be applied to approximately 98.7% of mentions.

Why This Matters

Impact on research. MedPath shifts biomedical entity linking from flat, single-vocabulary benchmarking toward interoperable, hierarchy-aware evaluation. Because every entity shares a canonical UMLS identifier, supervision can be pooled across datasets and domains, and per-semantic-type comparisons become meaningful. The hierarchy annotations give the community a testbed for designing partial-credit metrics and inherently explainable models, which the authors argue is a prerequisite for trustworthy clinical deployment.

Real-world applications.

  • Clinical decision support systems that need to map mentions to the right granularity of concept — distinguishing a plausible near-miss from a dangerous mislink matters more than whether the surface strings look similar.
  • Cross-vocabulary interoperability, letting a single system report a mention's equivalent concepts in SNOMED-CT, MeSH and ICD-10 at once, which is useful for billing, coding and record exchange across institutions.
  • Pharmacovigilance and adverse-event monitoring across drug labels, forum posts and clinical notes, where the source text domains and vocabularies differ sharply.
  • Literature curation and knowledge-graph construction, where long-tail concepts and deep hierarchies are common.

Industry relevance. The paper explicitly notes the regulatory push-back against "black-box" models in safety-critical domains and the absence of ground truth for interpretable clinical NLP. Cross-vocabulary mappings and hierarchical paths offer a concrete route to models whose predictions can be inspected and whose errors can be graded by clinical severity rather than treated as binary.

Future Directions

  1. Inherently explainable models. Rather than post-hoc explanations, train generative models that predict a mention's entire hierarchical path, not just an entity ID, so the prediction process itself is visible.

  2. Multi-vocabulary fluency and multi-task learning. Use the cross-BioKG mappings to train one model that predicts equivalent concepts across all 11 vocabularies simultaneously.

  3. Richer hierarchical evaluation metrics. The path annotations provide a testbed for metrics that capture domain-specific error semantics within and across BioKGs, going beyond the basic ancestor/descendant/hierarchy accuracy used in the preliminary experiments.

  4. Fine-grained error analysis and grounding. Open questions include whether models confuse sibling concepts more often than distant ones and at what hierarchy depth models begin to fail; the paper also suggests knowledge graph generation and using MedPath for pretraining or fine-tuning LLMs to be more factually grounded in medical knowledge.

The authors also flag open maintenance questions: annotations are tied to specific ontology versions (they report using UMLS 2025 AA, SNOMED CT May 2025 and MedDRA 27.1), knowledge bases evolve constantly, adding novel BioKGs requires community contributions, and no expert clinical validation was performed on the newly added mapping and path layers.

Target Audience

Biomedical and clinical NLP researchers, especially those working on entity linking, NER, ontology alignment or knowledge-graph grounding. It is also relevant to clinical informatics practitioners and industry ML engineers building vocabulary-agnostic or explainability-oriented clinical systems, and to evaluation-focused researchers looking for hierarchy-aware metrics beyond F1. Readers interested in dataset construction methodology, licensing constraints around clinical corpora, and reproducible multi-source harmonization pipelines will find the construction and availability sections particularly useful.

Authors’ abstract

Progress in biomedical Named Entity Recognition (NER) and Entity Linking (EL) is currently hindered by a fragmented data landscape, a lack of resources for building explainable models, and the limitations of semantically-blind evaluation metrics. To address these challenges, we present MedPath, a large-scale and multi-domain biomedical EL dataset that builds upon nine existing expert-annotated EL datasets. In MedPath, all entities are 1) normalized using the latest version of the Unified Medical Language System (UMLS), 2) augmented with mappings to 62 other biomedical vocabularies and, crucially, 3) enriched with full ontological paths -- i.e., from general to specific -- in up to 11 biomedical vocabularies. MedPath directly enables new research frontiers in biomedical NLP, facilitating training and evaluation of semantic-rich and interpretable EL systems, and the development of the next generation of interoperable and explainable clinical NLP models.

Read the original paper