Research
A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective Analysis
Overview Research area: Natural Language Processing, specifically named entity recognition (NER) using large language models (LLMs), in-context learning (ICL), multi-agent LLM systems, and retrieval-a

- arXiv
- 2511.19083
- Published
- 2025-11-24
- Authors
- Wenxuan Mu, Jinzhong Ning, Di Zhao, Yijia Zhang
AI summary
Overview
Research area: Natural Language Processing, specifically named entity recognition (NER) using large language models (LLMs), in-context learning (ICL), multi-agent LLM systems, and retrieval-augmented knowledge injection.
Technical level: Intermediate. Readers should be comfortable with standard NER terminology (spans, entity types, F1), the idea of in-context learning, and the general notion of agent pipelines and prompt construction. No training or architectural derivations are required.
Scope in one sentence: The paper proposes KDR-Agent, a multi-agent LLM framework that combines Wikipedia knowledge retrieval, entity disambiguation, and reflective self-correction with type definitions and a small static set of contrastive demonstrations, and evaluates it on ten NER datasets across five domains and three LLM backbones. The paper states it has been accepted by AAAI 2026.
What This Paper Is About
Standard in-context NER relies either on retrieving annotated examples from a labeled support set (few-shot) or on auto-labeling unlabeled text to build a provisional support set (zero-shot). The authors argue both approaches break down when annotated data is scarce, when the target domain is unfamiliar to the LLM, and when an entity mention is ambiguous without external context. The goal is to build an NER system that works across many specialized domains with only a small, static amount of labeled data, by actively fetching external knowledge and explicitly resolving ambiguity rather than just selecting better demonstrations.
Key Contributions
-
A multi-agent framework (KDR-Agent) for multi-domain low-resource in-context NER. It integrates external knowledge retrieval, entity mention disambiguation, and reflective analysis into the prompt construction and inference process, coordinated by a central LLM planner.
-
Entity-type definitions plus an entity-level positive–negative contrastive demonstration strategy. Instead of retrieving demonstrations from a large labeled corpus, the method uses concise natural-language type definitions and a small static support set in which each gold mention–type pair is paired with a perturbed negative. Negatives are built by randomly applying one of four error types: mention boundary alteration, incorrect type label, spurious mention not present in the input, or omission of a valid mention. The authors state this significantly reduces reliance on retrieving few-shot demonstrations from extensive labeled datasets.
-
Explicit handling of the two failure modes the authors identify as unaddressed by prior work. The central planner formulates targeted Wikipedia queries for domain-specific concepts and flags potentially ambiguous mentions; a Knowledge Retrieval Agent executes the queries and a Disambiguation Agent produces concept–context interpretation statements for the flagged mentions.
-
A two-stage reflective correction loop. A Reflective Analysis Agent compares the initial prediction against the input under four named error categories (span error, type error, spurious detection, omission) and produces a structured diagnostic report that drives a second prediction round.
Main Findings
-
KDR-Agent outperformed all zero-shot and few-shot baselines across all five domains and all three backbones. On GPT-4o, KDR-Agent reached F1 of 82.47 (BC5CDR), 79.41 (NCBI), 76.16 (MIT Movie), 69.98 (MIT Restaurant), 83.34 (CoNLL-2003), 71.85 (OntoNotes 5.0), 74.90 (Twitter Broad), 60.87 (Twitter NER-7), 74.37 (WikiANN), and 80.78 (WNUT-17).
-
Results held across backbones. With Qwen-2.5-72B, KDR-Agent scored 81.45 (BC5CDR), 79.22 (NCBI), 76.16 (MIT Movie), 67.70 (MIT Restaurant), 73.73 (CoNLL-2003), 67.37 (OntoNotes 5.0), 73.55 (Twitter Broad), 59.99 (Twitter NER-7), 71.06 (WikiANN), and 79.48 (WNUT-17). With DeepSeek-V3 it scored 81.38, 78.66, 76.31, 65.99, 75.67, 66.07, 74.95, 60.73, 70.44, and 78.52 on the same ten datasets in the same order.
-
Few-shot methods generally beat zero-shot methods. The paper reports that the overall performance of KDR-Agent and all few-shot baselines surpasses the zero-shot baselines, which the authors read as evidence that demonstrations help the LLM learn labeling guidelines such as entity type definitions and output formatting.
-
Gains were largest in Biomedical and Social Media domains. The authors attribute this to the value of external domain knowledge and context-aware disambiguation in these settings.
-
Ablation on GPT-4o shows every component contributes. On NCBI, removing Reflection dropped F1 from 79.41 to 75.91, removing the Knowledge Retrieval Agent to 76.21, removing the Disambiguation Agent to 75.49, removing both to 74.16, and removing the negative contrastive examples to 78.36. On Twitter NER-7 the corresponding numbers were 60.87 (full), 57.81, 59.34, 55.81, 55.07, and 58.99. On OntoNotes 5.0 they were 71.85 (full), 70.17, 71.70, 70.73, 69.94, and 70.69.
-
Different domains depend on different components. The authors report that biomedical NER relies heavily on both domain knowledge and disambiguation, social media NER depends more on ambiguity resolution, and the news domain changes relatively little, indicating lower reliance on domain-specific knowledge or disambiguation.
-
Smaller LLM backbones perform worse. Figure 2 shows a clear downward trend in F1 as Qwen model size decreases across three datasets, with the degradation more pronounced on NCBI (biomedical) and Twitter NER-7 (social media) than on OntoNotes 5.0 (news). The specific parameter sizes of the four Qwen variants are not listed in the text available here.
-
Reflection reduces all four measured error types on GPT-4o. Span error rate on NCBI fell from 22.03% to 9.18%, on OntoNotes 5.0 from 17.33% to 12.55%, and on Twitter NER-7 from 9.86% to 7.09%. Spurious detection rate fell from 16.44% to 5.57% (NCBI), 24.19% to 17.67% (OntoNotes 5.0), and 24.27% to 12.57% (Twitter NER-7). Omission rate fell from 49.62% to 17.78% (NCBI), 24.94% to 20.64% (OntoNotes 5.0), and 48.97% to 30.38% (Twitter NER-7). Type error rate fell from 6.14% to 4.49% (OntoNotes 5.0) and from 17.22% to 7.78% (Twitter NER-7); it is not applicable to NCBI because that dataset contains only a single entity type.
Methodology in Plain English
The system is organized as a pipeline with a coordinator and three specialist agents.
Stage 1 — building a richer prompt. The authors first write short natural-language definitions for each entity type, including what should and should not count as that type; these definitions come from the annotation guidelines that typically accompany NER datasets. They then prepare a small, fixed set of example sentences with labels. Critically, each correct mention–type pair is accompanied by a deliberately wrong counterpart produced by one of four perturbations (truncating or extending the span, changing the type, inventing a mention, or dropping a real one), so the model sees what boundary and type mistakes look like rather than only correct output. Because this example set is static, nothing has to be retrieved from a labeled corpus at inference time.
A central planner LLM then reads the input sentence and does two jobs: it spots concepts that probably need outside facts (rare or domain-specific terms) and writes Wikipedia search queries for them, and it flags mentions that could plausibly take more than one type. A Knowledge Retrieval Agent runs those queries through the MediaWiki Action API and returns the lead paragraph of each matched article as a short, source-attributable snippet; queries are restricted to information available before May 1, 2025, and only the summary section of each article is kept. A Disambiguation Agent writes a short natural-language statement interpreting what each flagged mention means in its local context. All of these pieces — task instruction, type definitions, contrastive demonstrations, retrieved knowledge, and disambiguation notes — are assembled into one prompt and the LLM produces an initial prediction.
Stage 2 — reflecting and correcting. The initial prediction is handed to a Reflective Analysis Agent, which is instructed to look for four specific error categories (span error, type error, spurious detection, omission), justify each suspected problem with evidence from the input, and propose revisions. That diagnostic report, plus a correction instruction, is fed back into the LLM for a second prediction round. The number of demonstrations is set to 10 for the datasets with more entity types (MIT Movie, MIT Restaurant, OntoNotes 5.0) and 5 for the rest, and the planner is capped at 5 Wikipedia query terms and 5 mentions to disambiguate. F1 is the reported metric, evaluation is on official test sets, and WikiANN — being large — is evaluated on a random sample of 5,000 instances from its development and test sets.
Why This Matters
Impact on research. The paper argues that the dominant paradigm in in-context NER — retrieving demonstrations from a support set — has a hidden assumption of available annotated data, and that improvement in specialized domains requires explicit external knowledge and ambiguity handling rather than better demonstration selection. It provides an empirical case, across ten datasets and three backbones, that a static contrastive demonstration set can replace retrieval, and that agent-decomposed knowledge retrieval and disambiguation can close gaps in domains where LLM internal knowledge is thin.
Real-world applications:
- Biomedical and clinical text mining, where the paper reports the largest absolute F1 scores (82.47 on BC5CDR and 79.41 on NCBI with GPT-4o) and where domain terminology and single-type label spaces make external knowledge and disambiguation especially valuable.
- Social media monitoring, where the paper reports the largest gains over baselines and the most improvement in type error rate after reflection (17.22% to 7.78% on Twitter NER-7).
- Task-oriented dialogue processing, for extracting movie and restaurant attributes from user utterances, a setting where the paper uses 10 demonstrations because the type sets are larger.
- News and open-domain knowledge graph construction pipelines, where entity mentions must be linked to consistent types across many documents.
Industry relevance. The framework's claim of practical value rests on two operational points: it removes the need to maintain and query a large annotated retrieval corpus at inference time, and it lets practitioners improve accuracy by adding domain knowledge and disambiguation rather than by collecting more labels. The paper asserts this reduces latency, though no latency measurements are reported in the content available here.
Future Directions
- Quantifying the cost side of the trade-off. The paper argues the static demonstration design reduces latency relative to retrieval-based methods, but reports no latency or token-cost measurements. Measuring the overhead introduced by the extra agent calls (planner, retrieval, disambiguation, reflection) against the retrieval savings remains open.
- Understanding the model-scale boundary. Figure 2 shows performance declining as Qwen backbone size decreases, and the authors note smaller LLMs "may be inadequate for complex domains." Where the cutoff lies, and whether agent decomposition can compensate for smaller backbones, is unresolved.
- Improving domains that benefited least. News-domain improvements were modest in both the ablation and the error analysis, which the authors attribute to syntactic regularity and semantic clarity. Whether that domain needs different agents, or none, is an open question.
- Reducing reliance on correct reflection criteria. The reflection stage is driven by predefined error categories and a manually constructed guideline; the paper does not report how sensitive results are to those definitions or whether they can be learned.
Target Audience
Researchers and practitioners working on named entity recognition, information extraction, and low-resource NLP who are already familiar with in-context learning and prompt-based extraction. It is also relevant to engineers building domain-specific extraction pipelines — especially in biomedical, social media, or dialogue settings — who need to operate with small labeled budgets and want to know which of knowledge retrieval, disambiguation, or self-reflection is worth adding, and for which domains. Readers primarily interested in LLM agent architectures will find the planner/specialist decomposition and the structured reflection loop the most transferable parts.
Authors’ abstract
In-context learning (ICL) with large language models (LLMs) has emerged as a promising paradigm for named entity recognition (NER) in low-resource scenarios. However, existing ICL-based NER methods suffer from three key limitations: (1) reliance on dynamic retrieval of annotated examples, which is problematic when annotated data is scarce; (2) limited generalization to unseen domains due to the LLM's insufficient internal domain knowledge; and (3) failure to incorporate external knowledge or resolve entity ambiguities. To address these challenges, we propose KDR-Agent, a novel multi-agent framework for multi-domain low-resource in-context NER that integrates Knowledge retrieval, Disambiguation, and Reflective analysis. KDR-Agent leverages natural-language type definitions and a static set of entity-level contrastive demonstrations to reduce dependency on large annotated corpora. A central planner coordinates specialized agents to (i) retrieve factual knowledge from Wikipedia for domain-specific mentions, (ii) resolve ambiguous entities via contextualized reasoning, and (iii) reflect on and correct model predictions through structured self-assessment. Experiments across ten datasets from five domains demonstrate that KDR-Agent significantly outperforms existing zero-shot and few-shot ICL baselines across multiple LLM backbones. The code and data can be found at https://github.com/MWXGOD/KDR-Agent.