Research
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
Overview Research area: Multimodal machine learning, specifically Audio-Visual Question Answering (AVQA), and selective prediction / uncertainty estimation. Technical level: Intermediate. The paper us

- arXiv
- 2602.04924
- Published
- 2026-02-04
- Authors
- Dinh Phu Tran, Jihoon Jeong, Saad Wazir, Seongah Kim, Thao Do, Cem Subakan, Daeyoung Kim
AI summary
Overview
- Research area: Multimodal machine learning, specifically Audio-Visual Question Answering (AVQA), and selective prediction / uncertainty estimation.
- Technical level: Intermediate. The paper uses formal notation for calibration, risk-coverage curves, and Bayes-optimality, but the core idea is intuitive: teach a model to say "I don't know" when it is likely wrong.
- Scope: The paper formalizes a "Reliable AVQA" (ℛ-AVQA) setting, benchmarks four existing confidence baselines against three AVQA backbones on three datasets, and proposes a lightweight learnable confidence-refinement method called Adaptive Confidence Refinement (ACR).
What This Paper Is About
AVQA models answer questions about combined audio and video, but they always answer, even when they are guessing. For users who cannot independently verify an answer, a confidently wrong response is worse than no response. This paper reframes AVQA as a selective prediction problem where the model may abstain, and it asks how well current models can tell when they are likely wrong. The authors show that standard confidence signals (like the maximum softmax probability) are poor at this, and they propose a small add-on module that learns to correct those confidence scores.
Key Contributions
- First reliability analysis of AVQA under selective prediction. The authors state this is the first study to analyze and evaluate AVQA model reliability under a selective prediction framework, where the model can answer or abstain.
- Adaptive Confidence Refinement (ACR), a learnable confidence estimator. ACR keeps MSP as the primary confidence signal and applies an input-adaptive residual correction, combining MSP with a learned residual risk signal via a simple linear fusion modulated by an input-adaptive weighting mechanism.
- Two learned heads. A Residual Risk Head (RRH) predicts low-magnitude correctness residuals that MSP does not capture, and a Confidence Gating Head (CGH) decides when MSP should be trusted, i.e., when to apply the residual correction.
- New benchmarks and baselines across three backbones and three datasets. The authors report that ACR consistently outperforms existing methods on in-distribution, out-of-distribution, and data-bias settings for QA-TIGER, ST-AVQA, and TSPM.
Main Findings
- Existing AVQA models are unreliable despite high accuracy. On MUSIC-AVQA, the state-of-the-art QA-TIGER method can answer only about 14% of questions while maintaining a 1% error rate, despite an overall AVQA accuracy of over 77%.
- Low-risk coverage is far from the oracle. In Table 1, the C@1% gap between the selection functions and the Oracle exceeds 60%.
- ACR beats MSP, MCD, VS, and Doctor on MUSIC-AVQA. With QA-TIGER, ACR reaches C@1% = 17.21, C@5% = 43.07, C@10% = 61.39, and C@20% = 91.96, versus MSP's 14.29, 41.53, 60.47, and 91.83. ACR's AURC is 8.41 and ECE is 4.88, compared with MSP's AURC of 8.67 and ECE of 8.83.
- Gains hold across backbones. With ST-AVQA, ACR reaches C@1% = 3.67 and ECE = 1.35 (MSP: 1.71 and 6.23). With TSPM, ACR reaches C@5% = 35.12 and C@10% = 54.50 (MSP: 30.04 and 49.21), with AURC 9.76 and ECE 1.51.
- Largest gains appear in the low-risk regime and on the bias-focused dataset. ACR achieves its largest C@1% gains on MUSIC-AVQA-v2.0 with QA-TIGER: +8.26 percentage points on the balance test set and +7.43 percentage points on the biased test set.
- ACR generalizes across distribution shifts. The authors report that ACR surpasses other methods on in-distribution, out-of-distribution (MUSIC-AVQA-R), and biased (MUSIC-AVQA-v2.0) settings, while other methods vary in performance across settings.
- ACR improves calibration. ACR achieves markedly lower ECE than the other methods across the reported tables.
- Accuracy correlates with reliability. Higher standard AVQA accuracy tends to accompany better selective prediction behavior; QA-TIGER is the most accurate model in standard AVQA and also surpasses TSPM and ST-AVQA in reliability evaluations across settings.
- Validation-selected thresholds transfer to test. The authors report that discrepancies between target risk and realized test risk are consistently small when thresholds chosen on a validation set are applied to the test set, a finding they align with prior selective-prediction literature.
- Theory supports the design. Theorem 3.3 states MSP equals the Bayes-optimal selection function under strong calibration; Theorem 3.5 gives a condition on the error cross-moment, σ_MR < min(σ²_M, σ²_R), under which the fused confidence has strictly lower MSE than either estimator alone; Theorem 3.6 shows that when the optimal fusion weight varies across inputs, any fixed weight is suboptimal in both MSE and ranking.
- Fusion is genuinely input-adaptive. The authors report that no distribution of the gating weight α collapses to fixed values (α = {0, 1}).
Methodology in Plain English
The authors take an AVQA model that already works reasonably well and freeze it. They then train two small multi-layer perceptrons on top of it, using the multimodal fusion features from the backbone's final layer and the pre-softmax logits as inputs.
The first head (Residual Risk Head) tries to predict whether the frozen model's answer is correct, producing a second opinion about correctness. The second head (Confidence Gating Head) looks at the same inputs and decides, per example, how much to weight the original MSP confidence versus the new residual-risk estimate. The final confidence is a weighted average of the two, with a per-example weight.
Both heads are trained with binary cross-entropy against a simple 0/1 label indicating whether the frozen model's prediction was correct. Because the backbone is frozen, the confidence estimation is decoupled from the base prediction task. The authors argue that binary cross-entropy is a strictly proper scoring rule, so minimizing it pushes the fused confidence toward the true posterior probability of correctness — which, under Theorem 3.6, is the Bayes-optimal ranking for selective prediction.
For evaluation, they use risk-coverage analysis: coverage is the fraction of questions the model answers, risk is the error rate among those answered, C@R is the maximum coverage achievable at a target risk (1%, 5%, 10%, 20%), AURC summarizes behavior across all operating points, and ECE measures calibration. They train the selector on validation data and split test sets into 20% for validating selection functions and 80% held out for final evaluation.
Why This Matters
- Research impact: The paper reframes AVQA as a reliability problem rather than a pure accuracy problem, and provides a formal problem definition, theorems, benchmarks, and baselines that other researchers can build on. It also extends selective prediction from unimodal and bimodal settings to a trimodal audio-visual-text setting, where modalities can be individually in-distribution but conflicting when combined.
- Real-world applications:
- Assistive systems for users with sensory impairments, who cannot independently verify a wrong answer.
- Any deployment where a wrong multimodal answer carries real cost and "I don't know" is a safe fallback.
- Data-bias-sensitive deployments, since ACR is evaluated on balanced and biased test sets.
- Systems that must handle rare or out-of-distribution inputs, as tested on MUSIC-AVQA-R.
- Industry relevance: ACR is described as lightweight, adds only two small MLP heads, and leaves the pretrained backbone untouched — making it plausible as a post-hoc addition to existing AVQA pipelines. It also has a practical advantage over Monte Carlo Dropout, whose K-fold inference the authors call prohibitive for latency-sensitive AVQA.
Future Directions
- Closing the gap to the oracle. The authors note the C@1% gap between practical selection functions and the oracle exceeds 60%, and explicitly call for improving both AVQA models and selection functions.
- Improving the underlying predictor, not just the selector. The paper reports that higher base accuracy correlates with better reliability and that backbones with higher accuracy tend to learn higher α values, but the authors state they focus only on improving the selector in this study.
- Extending the learnable confidence estimation scheme. The authors describe ACR as the first learnable confidence estimation scheme in trimodal AVQA, leaving room for other architectures, fusion strategies, or larger-scale backbones.
- Broader reliability evaluation. The reported study covers MUSIC-AVQA, MUSIC-AVQA-R, and MUSIC-AVQA-v2.0 with three backbones; the paper does not report evaluation on other AVQA domains or datasets.
Target Audience
Researchers and practitioners working on multimodal question answering, selective prediction, uncertainty estimation, and model calibration — especially those building AVQA systems where abstention is preferable to error. It will also interest engineers who need to add a reliability layer on top of a frozen pretrained multimodal model, and readers who want a formal treatment of when a simple confidence baseline like MSP is theoretically justified (strong calibration) and how to correct it when it is not.
Authors’ abstract
We present a formal problem formulation for \textit{Reliable} Audio-Visual Question Answering ($\mathcal{R}$-AVQA), where we prefer abstention over answering incorrectly. While recent AVQA models have high accuracy, their ability to identify when they are likely wrong and their consequent abstention from answering remain underexplored areas of research. To fill this gap, we explore several approaches and then propose Adaptive Confidence Refinement (ACR), a lightweight method to further enhance the performance of $\mathcal{R}$-AVQA. Our key insight is that the Maximum Softmax Probability (MSP) is Bayes-optimal only under strong calibration, a condition usually not met in deep neural networks, particularly in multimodal models. Instead of replacing MSP, our ACR maintains it as a primary confidence signal and applies input-adaptive residual corrections when MSP is deemed unreliable. ACR introduces two learned heads: i) a Residual Risk Head that predicts low-magnitude correctness residuals that MSP does not capture, and ii) a Confidence Gating Head to determine MSP trustworthiness. Our experiments and theoretical analysis show that ACR consistently outperforms existing methods on in- and out-of-disrtibution, and data bias settings across three different AVQA architectures, establishing a solid foundation for $\mathcal{R}$-AVQA task. The code and checkpoints will be available upon acceptance \href{https://github.com/PhuTran1005/R-AVQA}{at here}