Skip to content
AI.info

Research

Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts

Overview Research area: Natural Language Processing — benchmarking large language models on causal discovery from text, with emphasis on biomedical and clinical language. Technical level: Intermediate

Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
arXiv
2601.15479
Published
2026-01-21
Authors
Sydney Anuyah, Sneha Shajee-Mohan, Ankit-Singh Chauhan, Sunandan Chakraborty

AI summary

Overview

  • Research area: Natural Language Processing — benchmarking large language models on causal discovery from text, with emphasis on biomedical and clinical language.
  • Technical level: Intermediate. Readers should be comfortable with LLM prompting strategies (zero-shot, few-shot, Chain-of-Thought) and basic causal-inference vocabulary.
  • Scope: The paper builds a unified benchmark of 13 open-source LLMs across 12 datasets and 12 task configurations to measure how well models detect and extract pairwise causal relations from text.

What This Paper Is About

For LLMs to be safe in high-stakes settings such as medicine, they must distinguish causation from mere correlation — for example, whether a treatment actually caused an improvement or was simply administered at the same time. The paper tests this with pairwise causal discovery (PCD): given a text, decide whether it contains a true causal link, and if so, extract the exact cause and effect phrases. The authors benchmark 13 open-source models across 12 datasets spanning biomedical, financial, news, and abstract domains to see where these models succeed and where they break down.

Key Contributions

  1. A comprehensive PCD benchmark spanning multiple dataset types: intra-sentential vs. inter-sentential text, explicit vs. implicit causal markers, and single vs. multiple causal pairs. The authors state that their coverage matrix (Table 3) shows prior work each addressed only a single part of the problem, while their evaluation covers all 12 configurations.
  2. Twelve experimental setups organized as 4 Causal Detection tasks (1a–1d) and 8 Causal Extraction tasks (2a–2h), used to test robustness under different linguistic conditions.
  3. Evaluation of thirteen open-source LLMs spanning roughly 3B–70B parameters, with dense and MoE architectures and both base and instruction-tuned variants.
  4. A dataset validated through a strict annotation protocol, with mean inter-annotator agreement of κ ≥ 0.758, plus publicly released data, code, and prompts at https://github.com/sydneyanuyah/CausalDiscovery.

Main Findings

  • Overall performance is weak. The best detection model, DeepSeek-R1-Distill-Llama-70B, reached a mean score of 49.57% (C_detect). The best extraction model, Qwen2.5-Coder-32B-Instruct, reached just 47.12% (C_extract).
  • Explicit, single-sentence relations are easiest. Models generally performed best on simple, explicit, single-sentence causal relations (Task 1a, Task 2a).
  • Implicit and multi-sentence cases collapse. Performance dropped significantly for implicit relations within a single sentence (Task 1b), for cause–effect pairs spanning multiple sentences, and for texts containing multiple causal pairs.
  • Detection and extraction are separable skills. DeepSeek models excelled at detection but were deficient at extraction; Qwen2.5-Coder-32B-Instruct occupies the high-extraction, lower-detection region of the trade-off plot (Figure 2). The authors attribute the DeepSeek pattern partly to long reasoning traces: DeepSeek models recorded approximately 60% of their errors as missing values, because reasoning consumes tokens and the model sometimes stops or fails to emit an answer.
  • Instruction tuning matters greatly. Smaller, non-instruction-tuned models (e.g., DeepSeek-R1-0528-Q3-8B and Llama-3.2-3B) performed near zero, while their instruction-tuned counterparts improved substantially.
  • No single model dominates end-to-end. DeepSeek-R1-Distill-Qwen-32B leads in four tasks (2a, 2d, 2e, 2g), while Mistral-7B-Instruct-v0.3 leads in three (2b, 2f, 2h). Qwen2.5-Coder-32B-Instruct leads on the extraction side, coming first in five of eight extraction tasks. Mixtral-8x7B-Instruct-v0.1 reached 68.06% on X_{0,1,1}, which the authors link to MoE routing of queries to specialized experts.
  • Human annotators far outperform models. Annotators achieved a 95% true positive rate on detection tasks with κ ≥ 0.758 agreement, versus an average of 30–40% for LLMs.
  • Statistical tests confirm real differences. The Friedman test yielded χ² = 157.018, p < 0.0001 for detection and χ² = 229.624, p < 0.0001 for extraction. Across 78 t-comparison tests on I[A_2x = 1], DeepSeek-R1-0528-Q3-8B was demonstrably different from all other models; Qwen2.5-Coder-32B-Instruct and DeepSeek-R1-Distill-Qwen-32B were statistically indistinguishable (p ≈ 0.45); Llama-3.2-3B vs. Llama-3.2-3B-Instruct differed significantly (p ≈ 0.0436), as did Mistral-7B-Instruct vs. Mixtral-8x7B-Instruct (p ≈ 0.015).
  • Error profile for the top extraction model. For Qwen2.5-Coder-32B-Instruct on extraction, the largest buckets were Missing (35.70%) and Correct (45.55%), followed by Incorrect (6.07%), Hallucinated data (4.47%), Switched cause/effect (4.06%), and False Negative (3.22%). False Positive, New sentence, Rewrite cause/effect, Mixed, and Others were each at or below 0.31%.

Methodology in Plain English

Datasets and tasks. The authors assembled 12 datasets (Table 2) covering biomedical/medical, financial, general/news, and multi-domain/abstract text — including MedCaus (15K sentences), FinCausal (29K segments), SemEval2010Task8 (10.7K sentences), COPA (1K instances), CRASS (274 premise–counterfactual tuples), and PubMed (over 30M abstracts). Every instance is described by three axes of difficulty: marker presence (explicit vs. implicit), textual scope (within one sentence vs. across sentences), and cardinality (single vs. multiple cause–effect pairs). Detection tasks (1a–1d) ask whether a causal relation exists; extraction tasks (2a–2h) ask for all gold cause–effect pairs.

Annotation. Eight annotators were recruited through a screening test scoring ≥ 83% and split into two groups of four. They applied strict rules: only factual assertions count; cause and effect spans must be unambiguous; spans must be the smallest contiguous phrase and must exclude causal cue words such as "because"; multi-pair sentences are decomposed into separate relations; and ambiguous cases default to non-causal. Disagreements were adjudicated by a senior reviewer and the authors, and only instances with 100% agreement were retained.

Prompting. Detection used four prompting styles: instruction-only (zero-shot), few-shot in-context learning (FICL) with k = 3 examples, Chain-of-Thought (CoT), and hybrid CoT+FICL — the hybrid approach proved most effective. Extraction used five styles: general instruction-only, FICL, CoT, Least-to-Most (which extracts the cause first, then the effect), and ReAct (alternating thought/action/observation traces, used for the hardest multi-pair, inter-sentence tasks).

Scoring. Detection was scored with a simple binary script. Extraction spans were too variable for exact-string matching, so the authors used SentenceLM cosine similarity mapped into 10 bands, each calibrated with 400 annotated samples per band (4,000 total, seed=4000), producing a probabilistic score map (Table 6). Hungarian 1:1 matching pairs predicted spans to gold spans. On a 1,000-item human-checked test set, approximation error was 3.28% (96.72% accuracy) at the 5% significance level, dropping to 1.62% at corpus scale. Experiment 2 involved 8,925 instances, 13 models, and 5 prompting styles, yielding 580,125 generations. Reproducibility settings included seed=4000, temperature=0.7, batch size=64, and top-k=0.7.

Why This Matters

Impact on research. The paper provides the first unified evaluation framework the authors are aware of that covers all 12 combinations of marker presence, sentence scope, and causal cardinality in one benchmark. It shows that reported aggregate PCD scores hide large variation across linguistic conditions, and it supplies validated data, code, and prompts so results can be compared directly. It also reframes causal discovery as a "level-0" capability relative to Peal's association rung, arguing models must clear this bar before more complex causal inference is meaningful.

Real-world applications:

  • Automated literature review and evidence synthesis, where findings often connect ideas across multiple sentences in dense biomedical papers.
  • Adverse drug event detection and pharmacovigilance, replacing or assisting the manual annotation that currently underlies drug–side-effect databases.
  • Clinical decision support, where confusing a treatment that caused an outcome with one merely administered alongside it could have severe consequences.
  • Healthcare data mining over clinical notes and electronic health records, where marker-less, implicit causal language is common.

Industry relevance. The benchmark highlights practical cost/accuracy trade-offs across model sizes and architectures. DeepSeek reasoning models lead on detection but waste tokens on extraction; MoE models show selective-activation advantages for complex extraction; and instruction tuning produces very large gains over base models — all directly relevant to teams choosing models for text-mining pipelines.

Future Directions

  • Human-LLM hybrid workflows. The authors suggest blending model speed with annotator precision, since annotators reached a 95% true positive rate versus 30–40% for LLMs.
  • Better handling of inter-sentential and implicit causality. This is the clearest bottleneck, and it maps directly onto automated evidence extraction for systematic reviews and drug discovery.
  • Context-window extensions and MoE-inspired fine-tuning, which the authors propose after observing that failures stem from how models manage context and computation, not only from training data.
  • A fuller account of extraction behavior on multi-pair tasks, using the error bucketization shown for Qwen2.5-C-32B-Instruct (which still produced missing outputs in 35.70% of cases) to drive targeted improvements.

Target Audience

This paper is most useful to NLP and clinical-informatics researchers building or evaluating causal reasoning systems, benchmark designers who need a validated multi-domain PCD test suite, and applied machine learning engineers selecting open-source LLMs for biomedical text mining. Practitioners in pharmacovigilance, evidence synthesis, and clinical decision support will find the human-versus-model comparison and the failure analysis directly actionable. The study design mentions an expectation that fine-tuning would yield the largest gains while clever prompting offers modest improvements; the results reported in the content provided focus on the prompting strategies rather than on fine-tuned models.

Authors’ abstract

The safe deployment of large language models (LLMs) in high-stakes fields like biomedicine, requires them to be able to reason about cause and effect. We investigate this ability by testing 13 open-source LLMs on a fundamental task: pairwise causal discovery (PCD) from text. Our benchmark, using 12 diverse datasets, evaluates two core skills: 1) \textbf{Causal Detection} (identifying if a text contains a causal link) and 2) \textbf{Causal Extraction} (pulling out the exact cause and effect phrases). We tested various prompting methods, from simple instructions (zero-shot) to more complex strategies like Chain-of-Thought (CoT) and Few-shot In-Context Learning (FICL). The results show major deficiencies in current models. The best model for detection, DeepSeek-R1-Distill-Llama-70B, only achieved a mean score of 49.57\% ($C_{detect}$), while the best for extraction, Qwen2.5-Coder-32B-Instruct, reached just 47.12\% ($C_{extract}$). Models performed best on simple, explicit, single-sentence relations. However, performance plummeted for more difficult (and realistic) cases, such as implicit relationships, links spanning multiple sentences, and texts containing multiple causal pairs. We provide a unified evaluation framework, built on a dataset validated with high inter-annotator agreement ($κ\ge 0.758$), and make all our data, code, and prompts publicly available to spur further research. \href{https://github.com/sydneyanuyah/CausalDiscovery}{Code available here: https://github.com/sydneyanuyah/CausalDiscovery}

Read the original paper