Skip to content
AI.info

Research

Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse

Overview Research area: Evaluation of large language model reasoning, specifically mechanistic causal reasoning in NLP, combined with argumentation studies and climate-change discourse analysis. Techn

Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse
arXiv
2510.13417
Published
2025-10-15
Authors
Liesbeth Allein, Nataly Pineda-Castañeda, Andrea Rocci, Marie-Francine Moens

AI summary

Overview

  • Research area: Evaluation of large language model reasoning, specifically mechanistic causal reasoning in NLP, combined with argumentation studies and climate-change discourse analysis.
  • Technical level: Intermediate. The task setup is easy to follow, but the diagnostic evaluation metrics (Jaccard dissimilarity, Hamming distance, Fleiss' kappa) and the causal-reasoning terminology assume some familiarity with NLP evaluation practice.
  • Scope: The paper introduces the task of implicit causal chain discovery, provides a zero-shot baseline across nine LLMs on climate-change cause-effect pairs, and diagnostically tests whether these models perform genuine causal reasoning or merely associative pattern matching.

What This Paper Is About

When someone says "climate change causes flooding," they rarely spell out the intermediate steps that make that connection work — steps like excessive rainfall. The paper asks whether large language models can reconstruct those unstated intermediate causal steps, producing full causal chains that link an initial cause to a final effect. The goal is both to build a baseline method for this "implicit causal chain discovery" task in argumentative discourse and to diagnose whether LLM outputs reflect real causal reasoning or superficial pattern matching.

Key Contributions

  1. Task formulation and baseline method. The paper formally defines implicit causal chain discovery: given a cause-effect pair, generate a set of N causal chains, each passing through intermediate events, in a single zero-shot prompting step. The prompt is designed to be domain-agnostic and transferable across models.
  2. A diagnostic evaluation framework. The authors run five evaluation setups (A1–A5) probing self-consistency and confidence, directionality of causality, susceptibility to position heuristics, self- and cross-model chain integrity, and human expert judgment of chain validity and coherence.
  3. A benchmark dataset of generated causal chains. A publicly released dataset of structurally consistent causal chains and intermediate CE pairs generated by multiple LLMs, together with the reproduction code and human-evaluation setup, hosted at https://github.com/laallein/implicit-causal-chain-discovery.
  4. Human validation by argumentation experts. A controlled human evaluation in which domain experts judged the integrity and logical coherence of LLM-generated chains, showing where model judgments and human judgments diverge.

Main Findings

  • Models differ sharply in productivity and depth. Across the nine LLMs, the number of chains generated for a given CE pair ranges widely. On PolarIs3CAUS, o1 produced 865 chains in total (mean 9.2 per CE pair, max 25) while Phi 4-mini produced 346 (mean 4.12). On PolarIs4CAUS, Llama 3.1 Nemotron was the outlier with 3,714 chains (mean 20.52 per CE pair, max 173) versus Mistral Nemo with 744 (mean 4.13).

  • Neither productivity nor verbosity distinguishes reasoning models from general-purpose models. The variation in how many chains models generate, and in how long those chains are, does not cleanly separate reasoning LLMs (o1, o1-mini, DeepSeek R1) from general-purpose LLMs (GPT4o, Llama 3 70b, Mistral Nemo, Llama 3.1 Nemotron, Phi 4-mini, Mixtral).

  • More chains correlates with shorter chains. There is a statistically significant negative correlation (p < .01) between the number of chains generated for a CE pair and their average length, with Pearson's r ranging from −.11 to −.42 across models. The authors note this is somewhat counter-intuitive, since generating more chains might be expected to produce more elaborate, longer ones.

  • Models are self-consistent and confident about their own generated links. In setup A1, the majority of intermediate CE pairs were classified as causal, with minimal differences between general-purpose and reasoning LLMs, suggesting the models understood the task and the prompt was suitably transferable.

  • Directionality of causality is poorly understood. In setup A2, roughly 50% of reversed CE pairs were still judged as causal. The authors attribute part of this to evaluating each pair in isolation without context, and note that climate systems involve cyclic causality, such as the ice-albedo feedback loop.

  • LLMs are vulnerable to position heuristics. In setup A3, when the same causal content was rephrased from active to passive voice, disagreement was relatively low for the forward pairs, with Jaccard dissimilarity and Hamming distance values from 0.09 to 0.24. For the reversed (incorrect) pairs, disagreement rose substantially, to values from 0.25 to 0.51. The authors read this as a marked drop in consistency and evidence of reliance on surface-level patterns rather than genuine causal inference.

  • Reasoning models tend to generate higher-integrity chains, with one exception. In setup A4, reasoning models produced a higher proportion of chains with preserved integrity than general-purpose LLMs, except for GPT4o.

  • Chain length and chain quantity affect quality inconsistently. Contrary to the hypothesis that longer chains degrade causal discovery performance, the correlation between chain length and the proportion of causal intermediate pairs was not consistently negative across models. Whether generating more chains improves or worsens quality also differed by model.

  • Human experts rated the chains relatively highly but disagreed with the LLMs. In setup A5, participants confirmed the integrity of 27 out of 36 chains and marked 24 as logically coherent. However, LLM judgments disagreed with human assessments on the integrity of half of the chains. Inter-annotator agreement was low (Fleiss' κ = .084 for integrity; κ = .035 for coherence), despite participants expressing confidence in their own ability to construct valid, coherent chains.

Methodology in Plain English

The researchers took annotated cause-effect pairs from two existing resources: PolarIs3CAUS (95 pairs) and PolarIs4CAUS (181 pairs), both manually extracted from English climate-change discussions on Reddit (PolarIS-3) and X (PolarIS-4). Argumentation experts had standardized each cause and effect into a consistent noun-phrase form, and the pairs come from two polarized belief groups — climate change believers and skeptics — which makes them well suited to argumentative analysis. Because the resources were published after most models' training cut-offs, the authors expect they are not in the training data.

Nine LLMs were given a single zero-shot prompt containing a definition of a causal chain, the two slots for cause and effect, and formatting instructions requiring steps separated by a <step> token and chains by a <chain> token. No few-shot examples were used, to avoid biasing the number or granularity of chains; no chain-of-thought prompting was added; and no retrieval augmentation was used, since the aim was to test knowledge already encoded in the model parameters. The prompt was refined through manual prompt engineering, and most model outputs still required post-processing to parse the chains.

Evaluation then proceeded in stages. Intermediate links were extracted from each chain and tested through yes/no questions: whether the link is causal (A1), whether the reversed link is causal (A2), and whether judgments hold when the same question is rephrased in passive voice (A3). Chain integrity (A4) was judged by models evaluating their own generations and those of other models. Finally (A5), argumentation experts evaluated 36 chains covering 18 CE pairs from PolarIs4CAUS — two chains per pair, all generated by o1 — in a controlled in-person setting with 10 master's and PhD students from a communication faculty, each chain judged by four participants who were not told the chains were LLM-generated. The authors note that due to the high number of experiments and computing cost, they report results from a single inference pass per model and evaluation configuration.

Why This Matters

The paper's central argument is that knowing how a cause is thought to lead to an effect is what makes causal debate tractable. Making unexpressed causal links explicit can expose where speakers rely on shared assumptions, where explanations have gaps, and where rhetoric rests on faulty causal logic — which matters most in polarized domains like climate change.

Impact on research: The work reframes causal reasoning evaluation away from judging explicit causal relations already stated in text, and toward reconstructing mechanisms that were never stated. It provides a baseline, a diagnostic framework, and a released dataset for benchmarking, and it connects NLP causal reasoning to argumentation studies and belief polarization research.

Real-world applications:

  • Argument assessment and fallacy detection: Surfacing hidden causal steps lets analysts identify missing links, unsupported assumptions, and manipulative reasoning in public debate.
  • Counterargument generation and debate support: Detailed chains can supply the evidence and intermediate detail needed to strengthen or challenge a causal claim.
  • Media and platform analysis: The framework could help trace how the same causal claim is framed differently across polarized communities, where framing effects could otherwise change the causal chains a model infers.
  • Critical thinking education: The authors note participants expected their own chains to differ in length and detail from the ones evaluated, highlighting that clear guidance on desired depth is needed for consistent causal explanation.

Industry relevance: Any application that depends on language models producing reliable causal explanations — content moderation, fact-checking, policy analysis, scientific communication — is affected by the finding that models can be confident and self-consistent while still remaining sensitive to minor rephrasings of the same causal content.

Future Directions

  • Domain transfer. The authors expect different causal discovery behavior and quality across domains, since some domains have more complex or implicit mechanisms and therefore thinner coverage in training data.
  • Contested-causality domains. In areas where causality is heavily disputed, LLMs and causal knowledge graphs are more likely to encode conflicting or invalid causal associations, which complicates any evaluation of causal validity.
  • Retrieval-augmented approaches. RAG is flagged as a promising follow-up baseline, though the authors caution it may struggle to retrieve, align, and evaluate causal relations because of subtle mismatches in how cause-effect pairs are linguistically formulated.
  • Better chain-quality criteria. The paper notes that assessing integrity through intermediate pairs does not account for complexities specific to causal chains such as transitive inference, scene drift, and threshold effects — for example, the sequence cold → vasoconstriction → increased blood pressure → stroke → death.

Target Audience

This paper is most useful for NLP and AI researchers working on causal reasoning, commonsense reasoning, or LLM evaluation; for argumentation mining and computational linguistics researchers interested in how causality operates in polarized discourse; and for climate communication scholars studying belief polarization. Practitioners building systems that must produce or audit causal explanations — fact-checking, content moderation, policy analysis — will also find the diagnostic findings on prompt sensitivity directly relevant, though they should note the evaluation uses a single inference pass per configuration and is scoped to climate-change discourse.

Authors’ abstract

How does a cause lead to an effect, and which intermediate causal steps explain their connection? This work scrutinizes the mechanistic causal reasoning capabilities of large language models (LLMs) to answer these questions through the task of implicit causal chain discovery. In a diagnostic evaluation framework, we instruct nine LLMs to generate all possible intermediate causal steps linking given cause-effect pairs in causal chain structures. These pairs are drawn from recent resources in argumentation studies featuring polarized discussion on climate change. Our analysis reveals that LLMs vary in the number and granularity of causal steps they produce. Although they are generally self-consistent and confident about the intermediate causal connections in the generated chains, their judgments are mainly driven by associative pattern matching rather than genuine causal reasoning. Nonetheless, human evaluations confirmed the logical coherence and integrity of the generated chains. Our baseline causal chain discovery approach, insights from our diagnostic evaluation, and benchmark dataset with causal chains lay a solid foundation for advancing future work in implicit, mechanistic causal reasoning in argumentation settings.

Read the original paper