Research
Explanatory Interactive Machine Learning for Bias Mitigation in Visual Gender Classification
Overview Research area: Explainable AI (XAI) and interactive machine learning, applied to bias mitigation in visual gender classification. Technical level: Intermediate. Readers need some familiarity

- arXiv
- 2602.13286
- Published
- 2026-02-07
- Authors
- Nathanya Satriani, Djordje Slijepčević, Markus Schedl, Matthias Zeppelzauer
AI summary
Overview
- Research area: Explainable AI (XAI) and interactive machine learning, applied to bias mitigation in visual gender classification.
- Technical level: Intermediate. Readers need some familiarity with convolutional neural networks, attribution/explanation methods, and loss-function terminology, but the paper's framing is accessible.
- Scope: A case study evaluating whether two explanatory interactive learning (XIL) strategies — CAIPI and Right for the Right Reasons (RRR) loss — and a proposed hybrid of both can steer an image classifier away from spurious visual cues and toward relevant image regions, using a manually curated subset of MS COCO.
What This Paper Is About
Image classifiers can learn to make predictions from the wrong evidence — background objects, clothing, or context rather than the person themselves — which is a form of data bias. This paper asks whether letting a human give feedback on a model's explanations during training can push the model toward the features that actually matter, using gender classification as a deliberately bias-prone test case. The authors compare CAIPI, RRR loss, and a new combination of the two, measuring both how well the model attends to relevant regions and how evenly it misclassifies male and female examples.
Key Contributions
- An investigation of whether XIL methods can mitigate data bias in image classifiers, tested on visual gender classification.
- A proposed hybrid strategy that integrates CAIPI (augmentation-based) and RRR (loss-based) to optimize iteratively trained models for both accuracy and fairness.
- A qualitative and quantitative study of whether the steered model learns to focus on relevant regions, evaluated with a post-hoc explainability method (GradCAM) and a self-explaining method (Bounded Logit Attention, BLA).
- A systematic experimental comparison across three feedback strategies, three counterexample counts (k = 1, 3, 5), two sample selection strategies, and two explainability methods, plus six ImageNet-pretrained baseline architectures.
Main Findings
- XIL shifts attention to relevant regions. Explaining-guided feedback improved focus on the person and reduced attention to the background relative to the baseline, which scored FFP 0.35, BFP 0.31, BSR 0.74 and DICE 0.316. The effect was much more pronounced for CAIPI and the hybrid approach than for RRR.
- CAIPI produced the strongest alignment with ground truth. CAIPI with high-confidence sampling and k = 5 reached the best overall DICE score of 0.502, with FFP 0.49, BFP 0.18 and BSR 0.58. Performance improved as k increased.
- The hybrid approach suppressed background attention most effectively. Hybrid with high-confidence sampling, GradCAM and k = 1 yielded FFP 0.49, BFP 0.19 and the lowest reported BSR of 0.51; the same configuration with BLA reached 74.85% accuracy with FFP 0.47, BFP 0.18, BSR 0.54 and DICE 0.463.
- RRR's bias mitigation was less pronounced. RRR at its best (high-confidence sampling, BLA) reached FFP 0.39, BFP 0.19, BSR 0.64, DICE 0.356 and 72.35% accuracy, an improvement over the baseline but smaller than CAIPI or the hybrid.
- Accuracy usually dropped slightly. Steering generally cost a few percentage points of accuracy compared to the 74.58% baseline, which the authors describe as the price of improved explainability and bias mitigation.
- CAIPI was the exception to the accuracy trade-off. CAIPI with uncertainty sampling and k = 1 slightly improved accuracy from 74.58% to 75.42%, with FFP 0.34, BFP 0.22, BSR 0.70 and DICE 0.343.
- Misclassification imbalance was reduced. At baseline, males accounted for 46.2% and females 53.8% of misclassifications. CAIPI with high-confidence sampling and k = 5 balanced this to 48.0% versus 52.0%.
- Sampling strategy mattered, and the usual intuition was reversed. High-confidence sampling generally outperformed uncertainty sampling for XIL — the opposite of what is typically observed in Active Learning.
- BLA slightly outperformed GradCAM on accuracy. Differences between the two explainability methods were small, but BLA gave better classification results in most cases; the authors attribute the visual differences to GradCAM being post-hoc while BLA is learned jointly with the model.
- Baseline architecture choice. Of the six ImageNet-pretrained CNNs tested (DenseNet121, EfficientNet-B0, GoogLeNet, MobileNet-V2, ResNet50, VGG16), EfficientNet-B0 and VGG16 performed best; EfficientNet-B0 was reported in detail as the more recent architecture.
Methodology in Plain English
The authors built a gender classification dataset by manually annotating images from MS COCO, keeping only images with a single person in the foreground and excluding ambiguous cases. The result was 1,830 images, split evenly between 915 male and 915 female labels, and divided into a training set of 1,198, a validation set of 257 and a test set of 257.
Six CNNs pretrained on ImageNet were fine-tuned for 20 epochs with a learning rate of 1×10⁻⁴ using the Adam optimizer, establishing a baseline. Selected samples were then explained using GradCAM (which computes gradients of the target class on the last convolutional layer to produce a heatmap) or BLA (a trainable module integrated into the architecture that produces explanations faithful by design).
Feedback was delivered two ways. CAIPI generates counterexamples by applying a fixed sequence of image transformations — random inversion, posterization, equalization, color jittering and solarization — to the regions marked irrelevant by the ground-truth mask, while leaving the label unchanged, and adds these to the training set. RRR adds an auxiliary loss term that penalizes large gradients in irrelevant regions, combined with the standard cross-entropy term and L2 weight regularization. The hybrid approach first augments the data with CAIPI, then trains with the RRR loss.
Ground-truth segmentation masks from MS COCO simulated user feedback, allowing automated and quantitative experiments. Each iteration selected five samples using either uncertainty sampling (predictions close across classes) or high-confidence sampling (correct predictions above 90% confidence), and k ∈ {1, 3, 5} counterexamples per sample were generated. Evaluation used accuracy plus four explanation-derived metrics: DICE, Foreground Focus Proportion (FFP), Background Focus Proportion (BFP) and Background Saliency Ratio (BSR), with GradCAM maps discretized by thresholding at the 25% quantile and BLA providing hard explanations directly.
Why This Matters
The paper tests whether explanation-level human feedback — not just label correction — can address bias that is baked into a dataset's spurious correlations. It reports both a fairness benefit (more balanced misclassification rates between male and female predictions) and a transparency benefit (models attending to the person rather than the background), while documenting the accuracy cost that usually accompanies them. The finding that CAIPI can improve accuracy while reducing bias challenges the assumption that fairness and performance must trade off.
Real-world applications:
- Fairness auditing of face and person analysis systems, where classifiers may key on clothing, background or context rather than the subject.
- Medical and diagnostic imaging, where models can latch onto imaging artifacts or scanner-specific cues instead of pathology, and where explanation-guided correction is directly relevant.
- Content moderation and recommendation pipelines that use person attributes, where spurious shortcuts can produce systematically uneven error rates across groups.
- Human-in-the-loop annotation and model-refinement tools, where domain experts correct a model's reasoning rather than only its labels.
Industry relevance: the reported code and dataset are publicly released, and the paper is framed as directly comparable to prior work using standard pretrained architectures (DenseNet121, EfficientNet-B0, GoogLeNet, MobileNet-V2, ResNet50, VGG16), which lowers the barrier to reproducing or adapting the pipeline. The finding that high-confidence sampling works better than uncertainty sampling is a concrete, actionable guideline for teams building interactive labeling interfaces, and the accuracy-versus-fairness results give practitioners grounded expectations about the cost of explanation-based steering.
Future Directions
- Human user studies. All user interactions in this work were simulated with ground-truth masks; the authors state that evaluating the methods with real users is an important next step.
- Relaxing the annotation requirement. The evaluation depends on precise segmentation masks, which the authors note would need to be relaxed in practice, for example by using bounding boxes or coarse annotations.
- Testing beyond a single dataset. All experiments used one manually curated dataset, offering controlled conditions but limiting generalizability.
- Extending to multi-label and multi-class problems, which the authors explicitly name as future work.
Target Audience
Researchers and practitioners in explainable AI, interactive machine learning and algorithmic fairness who want evidence on whether explanation-level feedback can reduce spurious correlations in vision models. It is also useful for engineers building human-in-the-loop training or annotation systems, since it provides concrete guidance on sampling strategy, counterexample counts and the choice between post-hoc and self-explaining methods. Readers seeking a fully worked example of quantifying bias mitigation with segmentation-based explanation metrics (DICE, FFP, BFP, BSR) will find the experimental design and public code release particularly relevant.
Authors’ abstract
Explanatory interactive learning (XIL) enables users to guide model training in machine learning (ML) by providing feedback on the model's explanations, thereby helping it to focus on features that are relevant to the prediction from the user's perspective. In this study, we explore the capability of this learning paradigm to mitigate bias and spurious correlations in visual classifiers, specifically in scenarios prone to data bias, such as gender classification. We investigate two methodologically different state-of-the-art XIL strategies, i.e., CAIPI and Right for the Right Reasons (RRR), as well as a novel hybrid approach that combines both strategies. The results are evaluated quantitatively by comparing segmentation masks with explanations generated using Gradient-weighted Class Activation Mapping (GradCAM) and Bounded Logit Attention (BLA). Experimental results demonstrate the effectiveness of these methods in (i) guiding ML models to focus on relevant image features, particularly when CAIPI is used, and (ii) reducing model bias (i.e., balancing the misclassification rates between male and female predictions). Our analysis further supports the potential of XIL methods to improve fairness in gender classifiers. Overall, the increased transparency and fairness obtained by XIL leads to slight performance decreases with an exception being CAIPI, which shows potential to even improve classification accuracy.