Research
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
Overview Research area: Information retrieval (IR) evaluation and benchmark construction, with connections to retrieval-augmented generation (RAG) and multi-agent LLM debate. Technical level: Intermed

- arXiv
- 2602.06526
- Published
- 2026-02-06
- Authors
- Minjeong Ban, Jeonghwan Choi, Hyangsuk Min, Nicole Hee-Yeon Kim, Minseok Kim, Jae-Gil Lee, Hwanjun Song
AI summary
Overview
- Research area: Information retrieval (IR) evaluation and benchmark construction, with connections to retrieval-augmented generation (RAG) and multi-agent LLM debate.
- Technical level: Intermediate.
- One-sentence scope: The paper introduces DREAM, a multi-agent debate framework for labeling query-chunk relevance that produces the refined BRIDGE benchmark by filling missing relevant chunks ("holes") in BEIR and RobustQA.
What This Paper Is About
IR benchmarks are built by labeling only a small subset of text chunks, so many chunks that really are relevant stay unlabeled and are effectively treated as irrelevant; these gaps are called "holes." Existing LLM-based labeling either trusts a single model (risking overconfidence) or escalates uncertain cases to humans based on poorly calibrated confidence scores. The paper proposes DREAM, which uses two LLM agents with opposing stances who debate over multiple rounds, accepts labels when they agree, and sends only genuine disagreements to humans.
Key Contributions
- DREAM framework: A debate-based relevance labeling framework with two LLM agents initialized with opposing stances (one arguing "relevance," the other "irrelevance") that critique each other's arguments over multiple rounds, using inter-agent agreement as the reliability signal instead of confidence scores or threshold tuning.
- BRIDGE benchmark: A refined IR benchmark built by re-labeling subsets of BEIR (MS MARCO, NQ) and RobustQA (Lifestyle, Recreation, Science, Technology, Writing), uncovering 29,824 missing relevant chunks — 428% of the originally annotated 6,976 gold chunks — for a combined total of 36,800.
- Bias analysis: A demonstration that holes distort retrieval performance estimates and rankings, and that DREAM's hole filling reduces this bias, based on growth-rate and marginal-contribution metrics that converge toward zero.
- Retrieval–generation insight: A new RAGAlign metric and analysis showing that unaddressed holes cause retrieval–generation misalignment in RAG evaluation, which BRIDGE improves by 0.14 on average.
Main Findings
- Labeling accuracy and human effort: DREAM achieves 95.2% balanced accuracy with only 3.5% escalation to humans, measured on non-escalated cases. This surpasses the Human-Only reference of 93.8% balanced accuracy, which uses three MTurk workers per case with majority voting at 100.0% human involvement.
- Single-agent labeling is weak on irrelevance: LLMJudge reaches only 73.9% balanced accuracy with 0.0% escalation, driven by irrelevant-case recall of 50.2% (relevance recall 97.5%).
- Confidence-based escalation is inefficient: LARA reaches 82.1% balanced accuracy at 3.5% escalation, 83.9% at 12.5%, 87.8% at 25.0%, and 96.3% at 50.0%. To match DREAM's quality it needs escalation near 50.0% of cases.
- DREAM's class-wise detail: 91.9% irrelevance recall and 98.4% relevance recall, compared with LARA at 3.5% escalation (74.5% / 89.6%) and Human-Only (89.9% / 97.8%).
- Cost and latency: A cost analysis reports DREAM is 200× cheaper and 3.5× to 7.0× faster than Human-Only.
- Two debate rounds suffice: With LLM adjudication, DREAM scores 90.0% balanced accuracy at R=1, 93.3% at R=2, and 93.2% at R=3, so gains saturate after two rounds and R=2 is the default. Human adjudication at R=2 gives 95.1% balanced accuracy, higher than LLM adjudication at the same round count.
- Debate history helps human annotators: Providing debate history raises Fleiss' Kappa inter-annotator agreement from 0.50 to 0.62 and balanced accuracy from 87.3% to 92.0% on escalated cases.
- More debate diversity does not help: Increasing the number of agents, mixing heterogeneous LLM families, and raising temperature all failed to improve annotation performance; more agents reduced agreement on relevant cases, heterogeneous families suffered from performance disparities, and higher temperatures produced longer rationales that degraded quality.
- Holes are pervasive: The average Hole@10 ratio across 25 retrieval systems is 17.1%. Advanced fusion and newer systems show higher ratios — for example Arctic at 20.8%, MuGI at 22.3%, Aggretriever at 17.5%, BM25 + Rewrite at 21.2%, while Contriever is lowest at 6.5% and BM25 at 14.6%.
- Retrieval scores rise after hole filling: All 25 systems gain Hit@10 on BRIDGE, with widely varying magnitude. BM25 + Rewrite, which had the highest Hole@10 ratios, shows the largest improvement; BM25 goes from 0.65 to 0.65 with an improvement of 0.23, SPLADE gains 0.11, and Contriever gains 0.19.
- Rankings shift: 20 out of 25 systems change their retrieval rankings after hole filling — for example SBERT and ANCE drop while BM25+MuGI and SPLADE rise.
- RAG alignment improves: Average RAGAlign@10 rises from 0.70 on the original benchmark to 0.84 on BRIDGE, a gain of 0.14, with consistent trends across subsets (for example MS MARCO 0.59 to 0.88, Science 0.64 to 0.85).
- Bias converges to near zero: Growth rate of newly detected holes and marginal contribution of added retrievers both converge toward zero as the retriever pool expands, averaged over 10 runs with 25 systems added in random order.
Methodology in Plain English
The task is framed as deciding whether a candidate text chunk supports the answer to a given query. DREAM sets up two LLM agents with deliberately opposite starting positions. In each round, each agent sees the query, the chunk, the answer set, and the previous round's debate history; it critiques the opponent, extracts supporting evidence sentences, and issues a new label with reasoning. If the two agents agree in any round, the debate stops and that label is accepted automatically. If they still disagree after a maximum number of rounds (R=2 by default), the case is routed to humans — and the humans receive the debate history so they can see exactly where the argument stalled. This replaces confidence thresholds with agreement as the escalation signal.
To validate, the researchers built a 700-query evaluation set (100 each from MS MARCO and NQ in BEIR, and from five RobustQA domains) with expert-adjudicated ground-truth labels, and compared DREAM against LLMJudge, LARA at four escalation ratios, and a Human-Only MTurk baseline. All automatic methods used Llama3.3-70B-Instruct at temperature 0.0.
For BRIDGE, they sampled 3,657 queries (550 per dataset, except Science with 357 valid queries), built a candidate pool from 25 retrieval systems taking top-10 chunks each, applied an LLM-based filter to remove typical irrelevant cases, and obtained 116,622 query-answer-chunk triplets out of 296,053 initially retrieved. DREAM resolved 112,566 automatically within two rounds; the remaining 4,056 disagreement cases went to three MTurk annotators each with majority voting, with Fleiss' Kappa of 0.62. Total human cost was about $506.
Why This Matters
- Impact on research: The paper shows that a widely assumed property of IR benchmarks — that unlabeled chunks are safely treated as irrelevant — systematically biases retriever evaluation, especially penalizing newer fusion and rewriting systems. It also reframes retrieval–generation gaps in RAG as partly a benchmark artifact rather than purely a knowledge-conflict phenomenon.
- Search engine evaluation: Teams comparing sparse, dense, and reranked retrieval pipelines can use BRIDGE-style labels to avoid underrating systems that retrieve relevant chunks the original benchmark never labeled.
- RAG system development: Developers diagnosing why better retrieval fails to improve generation can check whether the retrieval metric itself is miscalibrated, using RAGAlign as a diagnostic.
- Annotation operations: Organizations building labeled datasets can adopt agreement-based escalation and debate-history handoff to cut human labeling cost — reported as 200× cheaper and 3.5×–7.0× faster than non-expert crowd labeling.
- Dataset maintenance: Because the bias analysis shows growth and marginal contribution converging to near zero, benchmark maintainers get a practical stopping rule for when a candidate pool is diverse enough.
- Industry relevance: Retrieval quality underpins enterprise search, question answering, and agentic systems built on RAG; reducing benchmark bias changes which retrievers companies would purchase or deploy.
Future Directions
- Reducing spurious consensus: The paper reports that increasing debate rounds does not consistently reduce spurious agreement, leaving open how to detect and suppress agreement that is wrong.
- Explaining why debate diversity fails: Adding agents, mixing model families, and raising temperature all hurt or failed to help; understanding when diversity helps versus harms would inform debate design.
- Extending beyond the seven subsets: BRIDGE covers two BEIR subsets and five RobustQA domains; whether the same refinement pipeline scales to the full BEIR collection or other IR resources is untested.
- Further closing the human loop: 3.5% of cases still escalate, and humans outperform LLM adjudication on those cases, so better ways to handle persistent disagreement remain an open problem.
- Broadening RAGAlign: The alignment metric is instantiated with Hit@10 and GPT-4o-based binary evaluation; its behavior with other retrieval and generation metrics is not established.
Target Audience
IR and RAG researchers who build or use retrieval benchmarks; practitioners designing LLM-based annotation pipelines who care about escalation cost and label quality; and dataset curators interested in multi-agent debate as an alternative to confidence-based human-in-the-loop labeling. Readers with some familiarity with retrieval metrics such as Hit@10, nDCG@10, and Recall@10 will get the most from the evaluation sections.
Authors’ abstract
Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies reduce costly human effort, they remain prone to LLM overconfidence and ineffective AI-to-human escalation. To address this, we propose DREAM, a multi-round debate-based relevance assessment framework with LLM agents, built on opposing initial stances and iterative reciprocal critique. Through our agreement-based debate, it yields more accurate labeling for certain cases and more reliable AI-to-human escalation for uncertain ones, achieving 95.2% labeling accuracy with only 3.5% human involvement. Using DREAM, we build BRIDGE, a refined benchmark that mitigates evaluation bias and enables fairer retriever comparison by uncovering 29,824 missing relevant chunks. We then re-benchmark IR systems and extend evaluation to RAG, showing that unaddressed holes not only distort retriever rankings but also drive retrieval-generation misalignment. The relevance assessment framework is available at https: //github.com/DISL-Lab/DREAM-ICLR-26; and the BRIDGE dataset is available at https://github.com/DISL-Lab/BRIDGE-Benchmark.