Skip to content
AI.info

Research

Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses

Overview Research area: Statistical machine learning — specifically conformal inference (distribution-free uncertainty quantification) applied to verifying the factuality of Large Language Model respo

Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
arXiv
2602.01285
Published
2026-02-01
Authors
Kangjun Noh, Seongchan Lee, Ilmun Kim, Kyungwoo Song

AI summary

Overview

Research area: Statistical machine learning — specifically conformal inference (distribution-free uncertainty quantification) applied to verifying the factuality of Large Language Model responses.

Technical level: Advanced. The paper combines conformal prediction theory (oracle filtering rules, exchangeability, Mondrian/group-conditional calibration, margin conditions) with an applied LLM system, and its central claims are theorems rather than purely empirical results.

Scope: The paper proposes Multi-LLM Adaptive Conformal Inference (MACI), a method that filters false claims out of LLM responses by treating factuality as a product of claim-level scores, calibrating thresholds separately per group, and improving score quality through an ensemble of three LLMs.

Note on completeness: the paper content supplied here is truncated. The Appendix material (including Appendices A, B.1, D and E) and the final rows of Table 1 are not visible, so some details — dataset sizes, the full group breakdown for ExpertQA, and any numeric timing measurements — are not reported in what follows because they do not appear in the available content.

What This Paper Is About

LLM responses are broken down into individual claims, and a system must decide which claims to keep so that the retained set is almost certainly factual. Existing conformal methods for this either apply one global threshold that throws away large amounts of true information, or use adaptive error rates and simple linear threshold functions that cannot capture the complex group structure of LLM responses and that relax the guarantee that high-stakes use requires. MACI's goal is to hold a fixed user-specified error rate while keeping as many true claims as possible.

Key Contributions

  1. A multiplicative filtering framework. The authors reformulate conformal inference so that document-level factuality is modelled as the cumulative product of claim-level factuality-scores, rather than as a single worst-case score. They derive an oracle filtering rule in this setting and show it retains the maximum number of claims subject to a target coverage constraint, while generic finite-sample guarantees (Theorem 1 for marginal coverage, Theorem 2 for group-conditional coverage) are preserved.

  2. A retention theory for conformal inference. The paper states that it provides the first retention theoretical analysis in conformal inference. Theorem 3 bounds the gap between the retention ratio achieved by an estimated score and by the oracle score, showing the gap is controlled by the mean squared error of the estimator via a polynomial rate that depends on a margin parameter β. This links oracle–estimator deviation directly to true-claim preservation and motivates the ensemble design.

  3. Group-conditional calibration plus a multi-LLM ensemble. MACI extends adaptive conformal inference with Mondrian-style calibration, computing a separate conformal quantile within each group, and uses an ensemble of multiple LLM factuality-scorers with weights chosen to minimize false-positive rate subject to a true-positive-rate tolerance. In the experiments this combination yields retention ratios substantially higher than the BCI and CCI baselines.

Main Findings

  • MACI reaches the target coverage while retaining far more claims. On MedLFQA at 80% target coverage, marginal retention is 0.71 for MACI versus 0.56 for CCI and 0.06 for BCI; at 90% it is 0.50 versus 0.31 and 0.02; at 95% it is 0.30 versus 0.18 and 0.01. All three methods are marked as hitting coverage within 1−α ± 0.01 on this dataset.

  • The pattern holds on WikiBio. At 80% target coverage, marginal retention is 0.43 for MACI, 0.19 for CCI and 0.02 for BCI; at 90% it is 0.25, 0.11 and 0.01; at 95% it is 0.13 for MACI, while CCI is flagged as under-covering (coverage 0.93) and BCI retains 0.01.

  • ExpertQA reveals a coverage-versus-retention difference. The baselines over-cover on this dataset (BCI coverage 0.91 where 0.80, 0.90 and 0.95 were targeted; CCI coverage 0.85), spending that excess on very low retention (BCI 0.13, CCI 0.17–0.18). MACI lands on the target coverage at every level tested while retaining 0.45 at 80%, 0.15 at 90% and 0.10 at 95%.

  • Group-conditional results mirror the marginal ones. Within MedLFQA's Medical Content subgroups (Info, Interpret, Action) and False-Claim Risk subgroups (Low, Medium, High), and within WikiBio's View Count and False-Claim Risk subgroups, MACI is reported to be the highest-retention method in almost all groups while staying at target coverage. BCI under-covers in several subgroups — for example WikiBio View Count "Low" at 80% (coverage 0.74) and MedLFQA False-Claim Risk "High" at 80% (coverage 0.73) — while CCI exceeds the target in several others.

  • The baselines are conservative in the way the paper predicts. BCI's retention is near zero on MedLFQA and WikiBio (0.01–0.06 in the rows shown) because it relies on a single worst-case conformity score, while CCI achieves higher retention but repeatedly violates the fixed-coverage target at 95% on WikiBio and across ExpertQA.

  • Ensembling is argued to reduce the retention gap. The paper states that Figure 3 shows the ensemble also reduces MSE in practice, consistent with Theorem 3's prediction that smaller estimation error yields a smaller retention gap.

  • Reported cost advantage without numbers. The abstract claims MACI achieves "lower time cost than baselines," but the available content does not report specific timing measurements.

Methodology in Plain English

The authors start by imagining a perfect factuality scorer that knows exactly how likely each claim is to be true. With this oracle, they ask: what is the best way to keep claims while guaranteeing that almost everything kept is true? The answer they derive is to sort a response's claims from most to least likely true and keep a prefix of that list, where the prefix is the longest one whose multiplied probability stays above a threshold. Multiplying probabilities across a whole document (rather than looking only at the single worst claim, as prior work does) makes the decision less sensitive to any one score being wrong.

Because no perfect scorer exists, the real system substitutes an estimated score and calibrates the threshold using held-out data. For each calibration document, the method computes the smallest threshold that would have made the filtered set contain only true claims; this single number summarizes the document's filtering event, and a conformal quantile of those numbers gives a threshold with a finite-sample coverage guarantee. To guarantee coverage separately for each group, the quantile is computed within each group instead of pooled across all documents.

Score quality matters for how much the method keeps, and the paper proves that the retention gap shrinks as the squared error of the score estimator shrinks. So the final ingredient is an ensemble: three LLMs (Llama-3.3-70B-Instruct, Qwen-2.5-72B-Instruct, DeepSeek-V3) are each asked to output a factuality probability in [0,1] for every prompt–claim pair, and the weights combining them are tuned to reduce the false-positive rate while keeping the true-positive rate above a tolerance. The authors report that directly minimizing MSE is impractical because the oracle score is unobservable and binary labels push predictors toward overconfidence.

Evaluation uses three datasets with different characteristics — MedLFQA, WikiBio and ExpertQA — each represented as (Prompt, Response, Claim Set, Ground Truth) with atomically decomposed claims and binary factuality labels. The underlying responses are fixed to those released with each dataset (mainly generated by GPT-4 and GPT-3.5-turbo), and all methods are applied to exactly the same pool of responses. Each dataset gets a dataset-specific grouping criterion plus a general False-Claim Risk grouping. Reported values are means over 30 repeated trials, at target coverage levels of 80% (α = 0.2), 90% (α = 0.1) and 95% (α = 0.05).

Why This Matters

Impact on research. The paper argues that previous work in this space forces a choice between rigorous but wasteful guarantees (BCI's global threshold) and more flexible but weaker ones (CCI's adaptive error rate and limited threshold function). By proving group-conditional validity alongside a retention bound tied to estimator error, it gives a concrete theoretical reason why better factuality scores translate into better filtering, and it positions retention — not just coverage — as a first-class quantity to analyse in conformal inference.

Real-world applications:

  • Clinical decision support, where medical answers must be filtered for factual reliability without discarding so much content that the output becomes useless.
  • Legal research and drafting, where an incorrect claim in a generated brief can have serious consequences.
  • High-throughput content pipelines that need to verify large volumes of generated text with per-group reliability (for example by topic or user population) rather than only on average.
  • Any deployment that treats an LLM as a black box: MACI needs only per-claim scalar scores, so it can be used as a plug-and-play filter on top of arbitrary generators, not just the models that produced the benchmark responses.

Industry relevance. The explicit trade-off MACI targets — a fixed, user-chosen error rate with as little discarded true information as possible — is what operational teams need when deciding how much machine-generated content to auto-publish versus route to human review. Group-conditional guarantees also matter for compliance and fairness auditing, since marginal guarantees can hide systematic under-coverage of specific subpopulations.

Future Directions

  • Group definitions. The paper instantiates grouping with high-level, dataset-specific categories (medical question types, entity groups, view-count buckets, risk tiers); how to define and validate groups that are meaningful in deployment remains open.
  • Score quality. Theorem 3 makes retention a function of estimator MSE, so improvements in factuality scoring — better ensembles, different verification prompts, or stronger verifier models — are a direct route to higher retention.
  • Sample size in small groups. The paper notes that each group is covered based on its own calibration size, so smaller groups get more conservative thresholds and reduced retention; handling small groups more efficiently is a natural extension.
  • Scope beyond the reported setup. The abstract and Appendix references point to additional analyses (MultiValid Conformal Inference and group-clustering comparisons, conformity-score variants, joint probability modelling, and behaviour under covariate shift) that are not visible in the truncated content; reproducing and extending those comparisons, and quantifying the claimed time savings numerically, are obvious next steps.

Target Audience

Researchers and graduate students working on conformal prediction, uncertainty quantification, or LLM factuality and hallucination detection; applied scientists and ML engineers building verification or guardrail layers on top of black-box LLMs in high-stakes domains; and statistically literate practitioners who need distribution-free coverage guarantees with per-group validity and care about how much content survives filtering. Readers without a background in conformal inference will find the theorems and the oracle-filtering construction demanding, though the motivation and empirical tables are accessible.

Authors’ abstract

Ensuring factuality is essential for the safe use of Large Language Models (LLMs) in high-stakes domains such as medicine and law. Conformal inference provides distribution-free guarantees, but existing approaches are either overly conservative, discarding many true-claims, or rely on adaptive error rates and simple linear models that fail to capture complex group structures. To address these challenges, we reformulate conformal inference in a multiplicative filtering setting, modeling factuality as a product of claim-level scores. Our method, Multi-LLM Adaptive Conformal Inference (MACI), leverages ensembles to produce more accurate factuality-scores, which in our experiments led to higher retention, while validity is preserved through group-conditional calibration. Experiments show that MACI consistently achieves user-specified coverage with substantially higher retention and lower time cost than baselines. Our repository is available at https://github.com/MLAI-Yonsei/MACI

Read the original paper