Skip to content
AI.info

Research

I-CAM-UV: Integrating Causal Graphs over Non-Identical Variable Sets Using Causal Additive Models with Unobserved Variables

Overview Research area: Causal discovery from observational data — specifically, integrating causal graphs learned from multiple datasets whose variable sets are not identical, in the presence of unob

I-CAM-UV: Integrating Causal Graphs over Non-Identical Variable Sets Using Causal Additive Models with Unobserved Variables
arXiv
2603.03207
Published
2026-03-03
Authors
Hirofumi Suzuki, Kentaro Kanamori, Takuya Takagi, Thong Pham, Takashi Nicholas Maeda, Shohei Shimizu

AI summary

Overview

Research area: Causal discovery from observational data — specifically, integrating causal graphs learned from multiple datasets whose variable sets are not identical, in the presence of unobserved (latent) variables.

Technical level: Advanced. The paper assumes familiarity with directed acyclic graphs (DAGs), causal discovery families (constraint-based, score-based, functional model-based), additive noise and causal additive models, and combinatorial search.

One-sentence scope: The paper proposes I-CAM-UV, an enumeration-based method that merges Causal Additive Model with Unobserved Variables (CAM-UV) results from several datasets with non-identical variable sets into a set of consistent full DAGs, including directed edges for variable pairs never observed together.

What This Paper Is About

Most causal discovery methods assume a single dataset and no unobserved confounders, but in practice researchers often have several datasets that share a goal yet measure different subsets of variables. Simply estimating a graph per dataset and overlapping the results leaves many causal relationships unidentified, because variables missing from a dataset can act as confounders, and some variable pairs never appear together in any dataset. The paper's goal is to recover a full causal DAG over the union of all variables by exploiting the structural information that CAM-UV provides about unobserved causal paths (UCPs) and unobserved backdoor paths (UBPs).

Key Contributions

  1. An enumeration approach named I-CAM-UV. It integrates CAM-UV results from multiple datasets with non-identical variable sets by enumerating all causal graphs that are "consistent" with those results, where consistency is defined through the (non-)existence of UCPs and UBPs on each dataset's variable set.

  2. A theoretical guarantee for ideal situations. The authors prove (Theorem 1) that the ground truth causal graph is consistent — and therefore appears in the enumerated set — whenever the union of observed variable sets equals the full variable set (no remaining confounders) and the CAM-UV results contain no estimation error.

  3. A relaxed problem for realistic settings. Because CAM-UV can err and unobserved variables may remain, the paper defines an inconsistency cost and reformulates the task as enumerating all DAGs whose inconsistency cost is at most C* + b, where C* is the minimum cost and b is a user parameter.

  4. An efficient best-first search algorithm. The method uses a lower-bound cost function with a proven monotonicity property (Theorem 2) to enumerate DAGs in ascending order of inconsistency cost, with polynomial-time procedures for searching UCPs and UBPs (each O(|A|) via breadth-first search).

Main Findings

  • Recall on identified variable pairs (Q1): I-CAM-UV achieved higher recall than all compared competitors — CAM-UV-OVL, PC-OVL, k-Nearest Neighbor imputation with k = 5 followed by CAM, and CD-MiNi — at recovering causal relationships that CAM-UV missed.

  • Recall on unobserved variable pairs (Q2): I-CAM-UV's recall was also superior on variable pairs that were never simultaneously observed in any dataset (the set E_uno).

  • Precision trade-off: I-CAM-UV showed lower precision than CAM-UV-OVL, and overall F1 scores showed almost no difference between the two. The authors caution that the output DAGs may contain a similar number of false discoveries.

  • Number of output DAGs (Q3): The number of DAGs enumerated by I-CAM-UV ranged widely — sometimes very few, sometimes huge — and is difficult to predict because it depends on the CAM-UV results. The accuracy distributions of the enumerated DAGs generally formed a single cluster with similar accuracies across many instances, suggesting random sampling could narrow a large output to a small alternative set.

  • Computation time (Q4): The enumeration process was fast enough compared with CAM-UV-OVL in many instances, and total computation times were comparable to the other methods, though some instances were significant outliers. The authors state I-CAM-UV runs in realistic time for relatively sparse DAGs with ten variables, but warn that it is exponential in general, with 3^|E| worst-case search states.

  • Experimental setup: Findings are based on 100 synthetic datasets generated from CAM with Erdős–Rényi random graphs of ten variables and edge probability 0.3, resampled into instances with m in {2, 3} datasets and |U| in {3, 4} unobserved variables per dataset, using 1,000 observations per dataset and b = 0.

  • Recoverability conditions: The paper states that recovering the ground truth DAG requires  ⊆ A*, A* \  ⊆ {(v_i, v_j) | {v_i, v_j} ∈ E}, and C(G*) ≤ C* + b. Example 4 shows a case where the ground truth cannot be recovered even under the relaxed problem.

Methodology in Plain English

The researchers start from CAM-UV, a method that fits non-linear additive causal functions and, when some variables are unobserved, returns a mixed graph: directed edges it is confident about, plus undirected edges marking pairs whose relationship is unidentified because of an unobserved causal path or an unobserved backdoor path.

They treat each dataset as a view in which all variables outside that dataset are latent, run CAM-UV on each, and then combine the results. Combining is framed as a constraint satisfaction problem: the final DAG over the union of variables must reproduce the right pattern of UCPs and UBPs for every dataset. Since many DAGs can satisfy this, they enumerate all of them rather than returning one.

To keep enumeration tractable, they assign a cost to each candidate — the number of variable pairs whose UCP/UBP status contradicts the CAM-UV findings — and design a best-first search that expands partial graphs in increasing order of a lower-bound cost. A monotonicity theorem guarantees that this lower bound never decreases as the search deepens, so the search can stop early once costs exceed a threshold. Searching for UCPs and UBPs on a given graph each takes O(|A|) time via breadth-first search, which makes evaluating each state cheap.

Why This Matters

Impact on research: The paper extends causal discovery beyond the single-dataset, no-confounding assumption, and unlike prior multi-dataset methods such as ION, IOD, and COmbINE (which output partial ancestral graphs) or CD-MiNi (which is restricted to linear non-Gaussian models), I-CAM-UV outputs explicit full DAGs for non-linear additive models. It also shows how latent-variable information can be used to orient edges between variables that no dataset observes together.

Real-world applications:

  • Integrating multi-cohort biomedical or clinical studies that measure overlapping but different sets of biomarkers.
  • Environmental and climate science, where monitoring stations record different subsets of variables (the paper cites environmental science among active causal discovery domains).
  • Materials and drug discovery pipelines, where different assays cover different measurements (also cited by the authors as an active area).
  • Any small-to-medium system (roughly ten variables) where hidden common causes are plausible and experiments are too costly, unethical, or technically infeasible.

Industry relevance: The work is a collaboration involving Fujitsu Limited, Shiga University, Gakushuin University, The University of Osaka, and RIKEN AIP, and the core search engine was implemented in C++ wrapped with Cython — indicating an intent toward practical, deployable tooling rather than purely theoretical analysis.

Future Directions

  • Beyond CAM-UV: Investigating whether causal graphs with unobserved variables can be integrated using other methods, such as RCD.
  • A Markov-equivalence-like notion: Developing a characterization of equivalence classes for the enumerated DAG sets, and deriving conditions under which I-CAM-UV yields a unique DAG.
  • Compressed representations: Building interpretable, compact summaries of enumerated DAGs so humans do not have to inspect every graph.
  • Fitness evaluation and upstream accuracy: Creating an algorithm to score how well a given DAG fits multiple datasets with non-identical variable sets, and improving CAM-UV itself, since I-CAM-UV's accuracy depends heavily on the accuracy of CAM-UV results.

Target Audience

Researchers and practitioners in causal discovery and causal inference who work with multiple heterogeneous observational datasets, particularly those in statistics, machine learning, epidemiology, environmental science, and the biomedical or materials sciences. Readers need a background in graphical causal models (DAGs, PAGs, additive noise models) and comfort with combinatorial algorithms to follow the theoretical sections; the experimental takeaways, however, are accessible to applied scientists evaluating whether to adopt the method.

Authors’ abstract

Causal discovery from observational data is a fundamental tool in various fields of science. While existing approaches are typically designed for a single dataset, we often need to handle multiple datasets with non-identical variable sets in practice. One straightforward approach is to estimate a causal graph from each dataset and construct a single causal graph by overlapping. However, this approach identifies limited causal relationships because unobserved variables in each dataset can be confounders, and some variable pairs may be unobserved in any dataset. To address this issue, we leverage Causal Additive Models with Unobserved Variables (CAM-UV) that provide causal graphs having information related to unobserved variables. We show that the ground truth causal graph has structural consistency with the information of CAM-UV on each dataset. As a result, we propose an approach named I-CAM-UV to integrate CAM-UV results by enumerating all consistent causal graphs. We also provide an efficient combinatorial search algorithm and demonstrate the usefulness of I-CAM-UV against existing methods.

Read the original paper