AI agents
Research Agents and Evidence Synthesis
Design research agents that plan queries, collect sources, reconcile conflicts, and produce claim-level evidence.
By the end you can
- Define research-agent systems as an operational contract rather than a capability label
- Contrast Answer-first research with Evidence-first research in “A market-research agent delivered a comprehensive report built on one data lineage”
- Trace “Citations can decorate a claim without supporting it” through a concrete execution path
- Produce “Design a research-agent evaluation” with evidence for “Material claims have source-linked evidence with dates and definitions”
Comparison
Competing strategies for research-agent systems
Answer-first research, Evidence-first research, and a Parallel research team differ in when the conclusion is allowed to exist. Answer-first writes it and then looks for support. Evidence-first is not allowed to write it yet. The risk that separates them is one sentence long. Citations can decorate a claim without supporting it. A reference collected after the conclusion was already written sits beside the sentence without ever having tested it.
That risk has a measured rate. In 2023 human evaluators went through the output of four commercial generative search engines — Bing Chat, NeevaAI, perplexity.ai and YouChat. On average only 51.5% of generated sentences were fully supported by their citations. Only 74.5% of citations supported the sentence they were attached to. Roughly one generated sentence in two was not fully supported. One citation in four did not support the sentence beside it. None of those systems was choosing between the three strategies above in any visible way. The output looked the same either way.
So the check on all three is the same. Take a claim out of the report and ask whether its source actually says that, on a stated date and under a stated definition.
Answer-first research
Draft a conclusion, then find supporting sources.
- Fast narrative
- Confirmation bias
- Weak contradiction search
Evidence-first research
Build claim-level evidence before synthesis.
- Stronger auditability
- More structured work
- Good default
Parallel research team
Several agents explore independent claims or source classes.
- Broad coverage
- Deduplication cost
- Needs coordinator
Visual
Teams skip the evidence ledger, one row per claim
A Research plan, an Acquisition stage, an Evidence ledger, and a Synthesis stand between the question and the report. Synthesis and Review should not share an owner, and should not share a test. The ledger is the part teams skip: one row per claim, with the source and the date beside it.
This is not a diagram invented for agents. Medicine has run it as a numbered reporting standard since 29 March 2021, the day the PRISMA 2020 statement was published. The statement describes its own shape in its summary points: “The PRISMA 2020 statement consists of a 27-item checklist, an expanded checklist that details reporting recommendations for each item, the PRISMA 2020 abstract checklist, and revised flow diagrams for original and updated reviews.”
Twenty-seven items. The ones agents lose first are the ledger columns: the eligibility criteria, the full search strategy for every database, and the date each source was last searched. A research agent that cannot produce those three has not skipped a formality. It has skipped the part of the record that lets a second person redo the search and find out whether the answer still holds.
- 1
Research plan
Claims, definitions, exclusions, dates, and evidence requirements.
- 2
Acquisition
Search, browse, query databases, and inspect primary records.
- 3
Evidence ledger
Sources, passages, tables, calculations, and unresolved conflicts.
- 4
Synthesis
Conclusions linked to evidence strength and limitations.
- 5
Review
Fact checks, methodological checks, and domain judgment.
Value comes from wider evidence, not authoritative prose
Research agents combine question decomposition, search, retrieval, browsing, data extraction, calculation, synthesis, and citation. That list is one pipeline: break the question up, go and find things, read them, pull the numbers out, do the arithmetic, write it up, and say where each claim came from. Their value comes from widening and organizing evidence. Not from making generated prose sound authoritative.
The task contract should define source classes, time boundaries, claim units, uncertainty, and what counts as sufficient support. It also needs a name for the way a cited claim fails, and there is a published one. A preregistered set of over 200 legal queries went to three RAG-based commercial legal research tools and to GPT-4. The researchers who ran them, Magesh and colleagues, named the state that sits between right and wrong: “A response is misgrounded if key factual propositions are cited but misinterpret the source or reference an inapplicable source.” Misgrounded is exactly what a bibliography conceals. The citation is present, the source is real, and it does not carry the proposition.
The rates they measured appeared in the Journal of Empirical Legal Studies in 2025. Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinate between 17% and 33% of the time. Lexis+ AI was accurate on 65% of queries, Westlaw AI-Assisted Research on 42%. These are retrieval-augmented products sold into a profession where the citation is the deliverable.
High-stakes conclusions still need qualified review. Qualified means somebody who could have written the report themselves.
Hand the report to a reviewer who could not have produced it and they will grade the writing, which is the one part the agent was never being trusted for.
Example
115 generated references, checked one by one: 7% survived
Thirty short medical papers, each with at least three references: that was the order given to ChatGPT-3.5 in 2023. The four researchers who placed it then looked up every one of the 115 references in Medline, Google Scholar and the DOAJ. Their result, in the abstract: “Among these references, 47% were fabricated, 46% were authentic but inaccurate, and only 7% were authentic and accurate.”
Of the 115 references, 7% were both authentic and accurate. The 46% that were authentic but inaccurate are the dangerous half. The paper exists. The journal exists. A reader who recognises the name nods and moves on, and the paper does not say the thing it was cited for. The papers also carried an incorrect PMID in 93% of cases — an identifier whose only function is to be looked up, wrong in more than nine papers out of ten. Read as prose, those bibliographies looked like coverage. Counted, they were a long list resting on almost nothing.
- Decision at stake: Design research agents that plan queries, collect sources, reconcile conflicts, and produce claim-level evidence.
- Hidden assumption: A long bibliography indicates that research coverage is sufficient. 115 references looked like coverage. 7% were both authentic and accurate.
- Primary control question: Citations can decorate a claim without supporting it. 47% of those references were fabricated outright and a further 46% were authentic but inaccurate. Only opening them separates the two.
- Evidence to collect: Material claims have source-linked evidence with dates and definitions. Check them the way those 115 were checked: against Medline, Google Scholar and the DOAJ, one reference at a time, including the identifier that was wrong in 93% of papers.
Key idea
Citations can decorate a claim without supporting it
A source may mention the topic without actually supporting the claim, use a different definition, or describe an outdated period. A citation being there is not the same as the claim being supported. Somebody has to open the link.
How often it fails is known. The audit of those four generative search engines says so in its abstract: “We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence.” Fluent and appear informative is the whole difficulty. The failure has no surface signature. It is only visible from the other end of the link.
And nobody looks until the decision has gone out. A brief filed in Mata v. Avianca cited six judicial opinions that do not exist — “Varghese”, “Shaboon”, “Petersen”, “Martinez”, “Durden” and “Miller” — produced by ChatGPT. On 22 June 2023 Judge P. Kevin Castel issued a 43-page sanctions order. It is blunt: “Peter LoDuca, Steven A. Schwartz and the law firm of Levidow, Levidow & Oberman P.C. (the “Levidow Firm”) (collectively, “Respondents”) abandoned their responsibilities when they submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question.” The court fined the two attorneys and their firm $5,000, jointly and severally, payable into the Registry of the Court within 14 days. The six fake authorities passed drafting and filing untouched. What caught them was judicial orders calling their existence into question. Even that did not stop the respondents standing by them.
Evaluate source authority, passage entailment, independence, recency, and calculation lineage at claim level.
Nobody notices a citation that does not hold until the decision built on it has already gone out.
Steps
Design a research-agent evaluation
Grade a research agent on claims rather than on reports, and build that evaluation for one real research workflow. For each sampled claim, open the cited source and mark whether it states that claim, states something weaker, or does not mention it at all. The graded set is your evidence that material claims have source-linked evidence with dates and definitions. Grade twenty claims, not twenty reports.
The two numbers that grading produces already have names. ALCE, the first automatic benchmark for evaluating LLM citations, scores systems on fluency, correctness and citation quality. It splits that last one into citation recall and citation precision — the claim-level pair your ledger is implicitly computing. The team that built it reported in 2023: “Our experiments with state-of-the-art LLMs and novel prompting strategies show that current systems have considerable room for improvement—For example, on the ELI5 dataset, even the best models lack complete citation support 50% of the time.”
Make the claim units short enough to check against a reference answer, the way OpenAI's BrowseComp does. Published in 2025, it is 1,266 questions requiring persistent web navigation. Human trainers solved only 29.2% of the 1,255 questions they attempted, and gave up after two hours on 70.8% of them. GPT-4o scored 0.6%. GPT-4o with browsing, 1.9%. OpenAI Deep Research, 51.5%.
Then read the calibration column. It is the one a research agent lives or dies by. The paper states it plainly: “As shown in Table 3, we found models with browsing capabilities such as GPT-4o w/ browsing and Deep Research exhibit higher calibration error, suggesting that access to web tools may increase the model’s confidence in incorrect answers.” Deep Research posted the worst calibration error in that table, 91%. The strongest browser in the benchmark was also the most confidently wrong one. Injecting contradictory evidence and auditing calculations are how you find that out before the report is written, not after.
- 1
Create answerable claim units
Write facts and calculations that can be graded independently.
- 2
Specify source expectations
Prefer primary records and define acceptable secondary evidence.
- 3
Inject contradictory evidence
Test whether the agent preserves and explains disagreement.
- 4
Audit calculations
Recompute tables and derived values from cited inputs.
- 5
Review synthesis
Check coverage, uncertainty, omissions, and unsupported rhetoric.
An index nobody refreshes still answers confidently
A research agent should be allowed to conclude that the evidence is insufficient. Forced completeness is a major source of fabricated certainty.
The call belongs to whoever reads the report and has to act on it. They decide whether the references under a claim are load-bearing or ornamental. A research agent earns that trust when every claim that matters carries a link to the source it came from, the date that source speaks to, and the definition it uses for the terms in question.
Grounding is defined, and somebody has to keep it up to date. Retrieval-augmented generation pairs parametric with non-parametric memory; Lewis and colleagues set it out in 2020. AWS describes the operational half: the model consults an authoritative knowledge base outside its training data before it answers. That same documentation says the documents and their embeddings must be updated asynchronously.
Staleness has been benchmarked, and so has the related failure of accepting a question's premise. FreshQA is a dynamic benchmark of 600 questions that mixes fast-changing items with false-premise items, and its builders collected more than 50,000 human judgments. Their result, reported in 2024: “Through human evaluations involving more than 50K judgments, we shed light on limitations of these models and demonstrate significant room for improvement: for instance, all models (regardless of model size) struggle on questions that involve fast-changing knowledge and false premises.”
Regardless of model size is the part to keep. A larger model does not refresh an index, and it does not debunk a premise it was handed. Both failures produce the same artefact: an answer built from a snapshot, delivered at the confidence of a fact. The refresh is a job somebody owns, on a schedule somebody wrote down.
Required to answer from an index nobody refreshes, the agent still answers, and it sounds exactly like the right one.
Key takeaways
- Research agents combine question decomposition, search, retrieval, browsing, data extraction, calculation, synthesis, and citation.
- The task contract should define source classes, time boundaries, claim units, uncertainty, and what counts as sufficient support — including a name for partial failure. A response is misgrounded when key factual propositions are cited but misinterpret the source or reference an inapplicable source.
- Claims, definitions, exclusions, dates, and evidence requirements. The PRISMA 2020 statement has asked systematic reviewers for the same since 29 March 2021, across 27 checklist items, down to the date each source was last searched.
- Search, browse, query databases, and inspect primary records — then check what came back. Across four commercial generative search engines, only 51.5% of generated sentences were fully supported by their citations, and only 74.5% of citations supported the sentence beside them.
- Evaluate source authority, passage entailment, independence, recency, and calculation lineage at claim level. Of 115 references ChatGPT-3.5 produced for 30 medical papers, 47% were fabricated, 46% were authentic but inaccurate, and only 7% were both authentic and accurate.
- A research agent should be allowed to conclude that the evidence is insufficient. Forced completeness is a major source of fabricated certainty, and BrowseComp's calibration column shows the reverse: OpenAI Deep Research scored highest on the benchmark, 51.5%, and posted the worst calibration error in the table, 91%.