AI agents
Parallel Research and Map–Reduce Patterns
Design parallel research agents with independent scopes, source diversity, and a verifiable synthesis stage.
By the end you can
- Define parallel research agents as an operational contract rather than a capability label
- Contrast Topic split with Source-method split in “Five research agents returned the same article through different URLs”
- Trace “Parallel search can multiply the same retrieval bias” through a concrete execution path
- Produce “Design a research map–reduce job” with evidence for “Parallel workers increase unique authoritative evidence, not only document count”
Visual
Evidence widens at map and narrows at merge
Evidence widens at the Research plan and the Map assignments, then narrows again at the Provenance merge and the Conflict analysis. Whoever runs the Conflict analysis should not also own the Synthesis, and each should be tested on its own.
The shape is borrowed, and the loan came with measurements attached. Dean and Ghemawat named and measured the pattern in 2004. Their benchmarks ran on a cluster of approximately 1,800 machines. At the time of writing, upwards of one thousand MapReduce jobs ran on Google's clusters every day.
What they measured is the part a research fan-out inherits and usually forgets to budget for. The paper puts it in one line: “As an example, the sort program described in Section 5.3 takes 44% longer to complete when the backup task mechanism is disabled.”
The 44% is not the cost of mapping. It is the cost of the tail at the merge. After 960 seconds all but 5 of the reduce tasks were finished. Those last stragglers took a further 300 seconds. Widening the map is the cheap half of the job. The expensive half is the stage where everything has to come back together and be reconciled. That is the stage a research fan-out is most tempted to skip.
- 1
Research plan
Claims, source classes, exclusions, and coverage requirements.
- 2
Map assignments
Independent bounded searches that produce evidence ledgers.
- 3
Provenance merge
Deduplicate common origins and preserve source relationships.
- 4
Conflict analysis
Compare dates, definitions, methods, and authority.
- 5
Synthesis
Write conclusions linked to supporting and contradicting evidence.
Optimist and critic branches diversify style, not evidence
Parallel research can divide a topic by claim, source class, geography, time period, or method. The map stage produces scoped evidence artifacts. The reduce stage deduplicates origins, resolves conflicts, and checks coverage.
Splitting by vague roles such as “optimist” and “critic” may create style diversity without evidence diversity. Assignments should force genuinely different search spaces or validation methods. The question to ask of any proposed split is not what each worker will argue. It is which records each worker can reach that the others cannot. That question has an answer before the first search runs.
Figure
The test of a branch split is whether two branches could return different origins; a persona split fails it before the first search runs.
Example
220,553 citation paths that reduce to four papers from one laboratory
An entire field already accepted that beta amyloid is involved in inclusion body myositis. Steven A. Greenberg did not count the support for that claim. He traced it to its origins, and published the result in the BMJ in 2009.
Counted as documents, the case is overwhelming. The citation network held 242 papers joined by 675 citations, and 220,553 distinct citation paths through that network supported the belief. Any coordinator tallying confirmations would stop there and write that the claim is widely supported.
Traced to origins, the same network is four papers. Only four authoritative papers in it actually reported supporting experimental data, and Greenberg records what they had in common: “All four papers were from the same laboratory, two of which probably reported mostly the same data without citing each other, a practice currently viewed as one that distorts available evidence”. The traffic was concentrated the same way. 63% of all citation paths (n=139,391) flowed through a single review paper, and 95% through four review papers by one research group.
The network also shows what happens to the evidence pointing the other way. Of the 214 citations to primary data, the supportive papers took 94%. The six papers whose data weakened or refuted the claim took 6% (P=0.01). Nothing here was fabricated. No single author did anything visibly wrong. The distortion is a property of the merge. Six hundred and seventy-five citations were counted as if they were 675 witnesses. Only four of them had seen anything at first hand, and all four sat in one room.
- Decision at stake: Design parallel research agents with independent scopes, source diversity, and a verifiable synthesis stage — so that a pooled answer is graded on origins the way Greenberg graded them, not on the 220,553 paths that led back to four papers.
- Hidden assumption: Several agents citing different URLs proves independent corroboration. In Greenberg's network the 675 citations across 242 papers reduced to four primary-data papers from the same laboratory, two of which probably reported mostly the same data.
- Primary control question: Parallel search can multiply the same retrieval bias. Ask what share of the pooled support passes through one node: Greenberg measured 63% of paths through a single review and 95% through four reviews by one group.
- Evidence to collect: Parallel workers increase unique authoritative evidence, not only document count. Count sources that report their own data — four, in the network — rather than paths that mention it, of which there were 220,553. Then check whether the contradicting evidence survived at all: the refuting six papers held 6% of the 214 primary-data citations against 94% (P=0.01).
Analogy
Survey Teams Covering Different Terrain
Survey crews take distinct regions and instruments, then merge their measurements onto one map. Sending every crew down the same road creates no new coverage.
A crew can see that it is walking a road another crew already covered. Five agents cannot. They can file five reports that all trace back to one laboratory, one agency wire, or one press release, and nothing in the five reports says so. Only the merge can find that out, and only if it is asked to look at origins rather than at counts.
Parallel research needs independent assignments and provenance-aware reduction.
Comparison
Where designs for parallel research agents diverge
Topic split, Source-method split, and Persona debate divide the work along three axes. Only two of them divide the evidence. A persona split changes who argues, not where anyone looks. A topic or source-method split sends workers to different places.
The source-method split is the one that has been measured. Searching is itself a method with a reported error rate, and systematic review is the field that reports it. Bramer and colleagues tracked 58 published systematic reviews prospectively, covering 1,746 relevant references found by database searching. Their 2017 finding: “Sixteen percent of the included references (291 articles) were only found in a single database”. Those 291 articles are unreachable by any number of workers who all query the same place, however differently they phrase themselves.
The recall figures put the same point on a scale. Embase alone reached 85.9% recall. Embase plus MEDLINE reached 92.8%. Embase, MEDLINE, Web of Science Core Collection and Google Scholar together reached 98.3% overall recall, with 100% recall in 72% of the reviews. That climb from 85.9% to 98.3% was bought by adding source-methods, not by adding readers to one source. That is the difference between a split that pays for itself and a split that only thickens the report. Parallel workers should increase unique authoritative evidence, not only document count.
Topic split
Each worker investigates a different subtopic.
- Simple coverage
- Cross-topic dependencies
- Potential gaps
Source-method split
Workers use different primary records, datasets, or verification methods.
- Stronger independence
- Higher coordination
- Good corroboration
Persona debate
Workers adopt argumentative roles over the same evidence.
- Exposes rhetoric
- Correlated sources
- Weak sole strategy
Key idea
Parallel search can multiply the same retrieval bias
Agents may share search engines, query phrasing, cached results, and source-ranking assumptions. More workers then create volume without broader evidence.
The cleanest measurement of that failure was taken on human reporters. A 2008 study content-analysed 2,207 domestic news items in the Guardian, The Times, Independent, Daily Telegraph and Daily Mail, plus 402 broadcast items, across two week-long samples in 2006.
On the surface this is five independent investigations of the same day. 72% of press stories were written by named journalists. Only 1% were directly attributed to the Press Association or another agency service. Then the stories were set against the copy the agency services had actually produced. 30% replicated it almost verbatim, and a further 19% were largely dependent on it — nearly half of all press stories coming wholly or mainly from agency services. Counting PR, agency and other-media material together, the authors report that “60 per cent of press stories rely wholly or mainly on pre-packaged information, a further 20 per cent are reliant to varying degrees on PR and agency materials.” Only 12% were without any discernible pre-packaged content, and in a further 8% the presence of PR content was unclear.
Distinct outlets, distinct URLs, distinct bylines, one origin. A coordinator counting five confirmations in that corpus would have been counting the same wire five times. The byline is what hid it. Measure unique origins and method diversity, and assign at least one branch to seek contradiction or missing coverage.
Workers that all query the same way produce a thicker report and not one additional fact.
Steps
Design a research map–reduce job
A map–reduce research job is worth specifying end to end for one real question rather than sketching for many. Write down what each worker searches, what it hands back, and how the reduce step decides which of those returns survive into the answer. Specify it far enough to see whether the branches are genuinely looking in different places, or repeating one search under three names.
The evidence ledger does not have to be invented. PRISMA 2020 is a published reporting standard for exactly this problem: a checklist of seven sections with 27 items, written by 26 authors and published in the BMJ in 2021. Item 16a requires authors to describe the results of the search and selection process, from the number of records identified to the number of studies included, ideally using a flow diagram. It also settles, in advance and in writing, the question a reduce stage otherwise decides ad hoc — when two records are the same evidence. Its glossary states: “Records that refer to the same report (such as the same journal article) are “duplicates”; however, records that refer to reports that are merely similar (such as a similar abstract submitted to two different conferences) should be considered unique.”
Note where deduplication sits in the template flow diagram. The first box is “Records identified from: Databases / Registers”. The very next box, the first attrition step and before any screening at all, is “Records removed before screening: Duplicate records removed”. Origins are resolved before anything is judged, not after.
And deduplication is itself a step with a measured failure rate, so it should not be assumed away. Rathbone and colleagues hand-annotated a 1,988-citation search to build a benchmark, then compared tools: “The sensitivity (84%) and specificity (100%) of the SRA-DM was superior to EndNote (sensitivity 51%, specificity 99.83%).” A default author/year/title merge caught about half the duplicates. Between the two tools that was a 42.86% increase in duplicates detected. EndNote also wrongly discarded some unique records, losing evidence in the same pass in which it failed to collapse copies. A merged document count is therefore an overstatement of independent support twice over. The job earns its cost only if the pooled result holds sources a single agent would have missed, rather than more copies of the sources it would have found anyway.
- 1
Define claim units
List propositions the final answer must support or qualify.
- 2
Assign independent paths
Split by source type, jurisdiction, dataset, or verification method.
- 3
Require evidence ledgers
Capture origin, passage, date, authority, and uncertainty.
- 4
Deduplicate origins
Trace syndication and circular citation before counting support.
- 5
Reduce with conflict rules
Preserve disagreement and grade claim-level evidence.
Fan-out bills fifteen times the tokens of a chat
Multi-agent research is justified when it widens evidence or reduces latency. If all branches see the same context and tools, one well-evaluated agent may be better, because parallel search can multiply the same retrieval bias rather than escape it. Before widening the fan-out, ask what each additional branch would see that the others cannot. Then judge the result by the distinct authoritative sources it added, not by how many documents came back.
Fan-out has a published price. In 2025 Anthropic reported that “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens as chats.” That is a vendor's instrumentation of its own product, reported as approximate multiples.
The coordination price was measured by other people, on other people's frameworks, and it was measured rather than estimated. Cemri and colleagues state the method behind their failure taxonomy: “We develop MAST through rigorous analysis of 150 traces, guided closely by expert human annotators and validated by high inter-annotator agreement (kappa = 0.88).” That derivation yielded 14 unique failure modes in 3 categories: system design issues, inter-agent misalignment, and task verification. The taxonomy was then applied to MAST-Data, 1,600+ annotated traces across 7 multi-agent frameworks and models including GPT-4, Claude 3, Qwen2.5 and CodeLlama. Two of the three categories are failures of the merge rather than of any worker.
Widening the search does not by itself widen the evidence. The step that does is the one where somebody reconciles what came back — the step Dean and Ghemawat measured at 44%, the step PRISMA puts before screening, and the step Greenberg had to perform by hand before 220,553 supporting paths turned out to be four papers.
Fan-out is easy to buy and easy to mistake for rigor; the token bill arrives whether or not the branches disagree about anything.
Key takeaways
- Parallel research can divide a topic by claim, source class, geography, time period, or method; only a split that sends workers to different origins divides the evidence.
- Splitting by vague roles such as “optimist” and “critic” may create style diversity without evidence diversity.
- Greenberg's citation network held 242 papers, 675 citations and 220,553 supporting paths, yet only four papers reported supporting experimental data, and all four came from the same laboratory.
- Sixteen percent of included references (291 articles) turned up in only one database, and recall rose from 85.9% on Embase alone to 98.3% across four databases.
- Deduplicate origins before counting support: PRISMA 2020 removes duplicate records before screening, and a hand-annotated benchmark measured EndNote's default merge at 51% sensitivity against 84% for the SRA-DM.
- Multi-agent research is justified when it widens evidence or reduces latency. The bill is about 4× the tokens of a chat for an agent and about 15× for a multi-agent system, plus the 14 failure modes in 3 categories derived from 150 traces at kappa = 0.88.