Research
ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution
Overview Research area: Computer vision / explainable AI (XAI) — specifically the intersection of pixel-level saliency methods and concept-based interpretability. Technical level: Intermediate to Adva

- arXiv
- 2610.04605
- Published
- 2026-10-03
- Authors
- Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad, Noam Koenigstein
AI summary
Overview
Research area: Computer vision / explainable AI (XAI) — specifically the intersection of pixel-level saliency methods and concept-based interpretability.
Technical level: Intermediate to Advanced. Readers need some familiarity with classifier architectures (CNNs and Vision Transformers), attribution methods such as LRP and Grad-CAM, and the concept activation vector (CAV) literature.
Scope: The paper introduces ConEx (Concept-based Explanations), a fully automatic post-hoc framework that discovers class-specific visual concepts, grounds them spatially, builds concept activation vectors in a model's latent space, and produces saliency maps that are both faithful to the model and human-interpretable, evaluated across 3 main datasets plus ImageNet-Segmentation, 5 pretrained models, and 8 saliency baselines and 3 concept-based baselines.
What This Paper Is About
Existing explanation methods split into two camps that do not talk to each other: saliency methods show where a model looks but not what it perceives, while concept-based methods (like CBMs) reveal what semantic attributes matter but lack spatial grounding. ConEx aims to unify both perspectives by decomposing a model's prediction into spatially grounded, human-interpretable visual concepts (e.g., "yellow beak") and aggregating them into a concept-based explanation map. The framework is designed to work post-hoc, without manual concept curation or model retraining.
Key Contributions
- A unified automatic framework (ConEx) that bridges saliency-based visualization and concept-based reasoning so explanations are simultaneously spatial and semantic, requiring no manual annotation or retraining.
- A CAV construction pipeline that automatically discovers class-specific concepts using vision-language priors, grounds them with zero-shot segmentation, and eliminates the need for manual image curation per concept.
- Architecture-specific embedding strategies — layer-wise masking with neighborhood padding for CNNs and patch-based token selection and aggregation for ViTs — intended to increase concept purity and enable robust CAV construction for both architecture families.
- Two new validation metrics, VCM (Vector-Concept Match) and CCM (Concept-Class Match), which quantify alignment between learned CAVs and their visual meaning, and between discovered concepts and the model's decision boundaries.
Main Findings
- Faithfulness (perturbation): ConEx achieves state-of-the-art Insertion (INS) and Deletion (DEL) AUC across CUB and ImageNet on ResNet50, DenseNet121, and ConvNeXt-Base. On CUB with RN, ConEx scores INS 59.47 and DEL 8.19, versus the next-best baseline IIA at INS 58.13 and DEL 9.21. On ImageNet with RN, ConEx scores INS 59.11 and DEL 8.49. Results for the Stanford Dogs dataset are reported in the Appendix.
- Part-based benchmark: On FunnyBirds (500 images, 50 classes, predefined parts beak, wings, feet, eyes, and tail as concepts) with the provided RN model, ConEx achieves Completeness 0.78, Correctness 0.64, and Contrastivity 0.92, exceeding RISE (0.68/0.54/0.63), GC (0.70/0.56/0.68), AC (0.76/0.61/0.83), SC (0.72/0.57/0.85), and IIA (0.74/0.59/0.88).
- Concept quality: ConEx leads on Concept Insertion and Deletion and on the new metrics across all five architectures. On RN, ConEx reaches CINS 51.20, CDEL 8.74, VCM 0.68, and CCM 28.01, compared with ACE (44.89/10.82/0.53/23.04), ICE (43.74/11.28/0.49/21.38), and MCD (48.16/9.56/0.56/20.22).
- Vision Transformers score lower than CNNs: ViT-B reaches CINS 35.33, CDEL 10.35, VCM 0.53, and CCM 18.43; ViT-S reaches 34.30, 10.51, 0.51, and 17.92. The authors attribute this to the difficulty of operating on patch-level representations in ViTs.
- Human interpretability: In a study with 58 participants, ConEx achieves Selection Accuracy 72.13, PRC 70.48, INNS 0.52, and INTS 0.28, versus ACE (58.47 / 49.51 / 0.41 / 0.34) and MCD (34.32 / 62.28 / 0.47 / 0.39). ANOVA p-values are reported as less than 0.001 for all four metrics.
- Ablation — CAV construction: Replacing ConEx's CAV builder with those of ACE, ICE, Visual-TCAV, or MCD yields lower faithfulness. ConEx scores INS 59.21 / DEL 8.34, versus Con-MCD (58.16 / 10.22), Con-ICE (57.92 / 9.83), Con-ACE (57.45 / 9.58), and Con-Vis (56.98 / 10.61).
- Ablation — attribution fusion: LRP-based multiplicative fusion outperforms substitutes. Standalone GC, SHAP, and IIA score 54.92/12.68, 47.54/15.49, and 58.39/9.48 respectively; their ConEx-integrated variants reach 58.68/8.92, 56.99/9.84, and 58.96/8.65; ConEx itself is 59.21/8.34.
- Further ablations (reported in the Appendix): CAVs built from concept-specific segments are reported as decisively superior to those built from full images containing the concept, and the authors analyze sensitivity to N, the number of samples used for CAV creation.
- Design choices: Concept discovery used GroundedSAM (Ren et al., 2024) with an occurrence rate threshold of at least 15% and spatial coverage threshold of at least 20%, yielding 348 concepts for CUB, 189 for Stanford Dogs, and 3,602 for ImageNet. Initial concepts were created with GPT-4o following Oikarinen et al. (2023). CAVs used N = 100 positive and negative segments; CNN embeddings came from the final convolutional layer with 7x7 neighborhood padding, and ViT embeddings from layer 6.
Methodology in Plain English
ConEx works in stages, all post-hoc on an already-trained classifier.
- Find candidate concepts. For each class, the framework uses vision-language priors (following Oikarinen et al., 2023, with GPT-4o generating initial terms) to produce a list of linguistically interpretable, class-discriminative textual attributes — things like "yellow beak" or "floppy ears".
- Ground and validate them. Those attributes are passed to the zero-shot segmentation model GroundedSAM, which returns masks where the concept is visible. Concepts are kept only if they occur in at least 15% of the class images and if their union covers at least 20% (IoU) of the class's segmentation region. This is the automatic substitute for a human choosing concept example images.
- Build concept vectors. For each surviving concept, the authors collect N = 100 segments where it appears and N = 100 segments of other concepts, embed them, and take the difference of the two mean embeddings (
CAV_k = mu_k+ - mu_k-). The authors argue this centroid difference is more robust to outliers than classifier-based approaches such as linear SVMs. - Embed without masking artifacts. For CNNs, the image and its binary segment mask are propagated together and only activations from unmasked pixels are kept (layer-wise masking), with neighborhood padding applied at the first convolutional layer only to avoid boundary erosion. For ViTs, patch embeddings overlapping the segment are retained from an intermediate layer and averaged.
- Localize concepts. Each CAV is globally average-pooled into a channel-weighted vector (CWV), which is combined with a latent representation of the input image. A ReLU removes uncorrelated regions, and the map is normalized by the maximum activation of the positive concept centroid so results are comparable across concepts and images.
- Attribute relevance. Concept importance maps are produced by multiplying the concept localization map with an LRP attribution (epsilon-rule, epsilon = 0.01). Because multiplication requires both concept presence and model relevance to be high, irrelevant regions are suppressed. Each concept's contribution is measured by masking its region and recording the normalized drop in the predicted class probability.
- Aggregate. Summing the weighted concept importance maps over a class's concepts produces the final saliency map; averaging these over a dataset produces global, class-level concept rankings.
The two new metrics validate the vectors directly: VCM measures the cosine similarity between a CAV and the latent representation of held-out segments of the same concept, and CCM measures the cosine similarity between a CAV and the gradient of the associated class logit with respect to the latent representation.
Why This Matters
The paper's core claim is that faithfulness and interpretability need not be traded off: pixel-accurate localization can be combined with semantically named concepts, and grounding explanations in coherent concepts may also reduce the influence of spurious correlations and background artifacts. If explanations can be decomposed into named concepts, a human reviewer can check whether the reasons given for a prediction match domain knowledge, rather than inspecting a heatmap with no labels.
Real-world applications implied by this work:
- Auditing high-stakes classifiers: The paper frames opaque vision models as informing high-stakes decisions, where understanding the rationale is described as essential for trust and responsible deployment.
- Fine-grained visual domains: CUB-200-2011 and Stanford Dogs are the paper's evaluation datasets, so the framework maps naturally onto tasks where discrimination depends on small parts (beaks, eyes, fur patterns).
- Part-level analysis for controlled settings: The FunnyBirds evaluation uses predefined anatomical parts (beak, wings, feet, eyes, tail) as concepts, showing the pipeline can be driven by a supplied concept vocabulary rather than automatic discovery.
- Large-scale model inspection: ConEx is reported to scale to ImageNet (3,602 concepts) without human supervision or retraining, which matters for teams auditing models trained on datasets too large to annotate by hand.
Industry relevance: The framework is post-hoc and requires no retraining or manual annotation, so it can be applied to already-deployed models — an important practical constraint. It is also architecture-agnostic in terms of CAV construction: the paper reports high-quality global explanations for both CNNs and ViTs, though the complete spatial explanation pipeline is currently designed for CNNs. The authors note that extending it to ViTs is valuable in itself because it provides a mechanism to quantitatively audit the global semantic knowledge encoded in transformers. Code is released at https://github.com/yonisGit/conex.
Future Directions
- Spatial grounding for ViTs. The authors state that because ViTs lack inherent spatial correspondence, they require a dedicated formulation to enable localized concept-based explanations — described as an open challenge beyond the scope of this work, with the current CAV pipeline offered as a foundation for future ViT localization research.
- Closing the CNN–ViT gap. The consistently lower VCM, CCM, CINS, and CDEL scores for ViT-B and ViT-S are attributed to the difficulty of working with patch-level representations; improving embedding strategies for patch-based architectures is a clear next step.
- Layer and sample selection. The paper mentions an Appendix ablation on the ViT extraction layer and an analysis of sensitivity to N, the number of samples used for CAV creation — both suggest headroom in tuning the construction of concept vectors.
- Comparison with manually curated pipelines. The paper notes that a direct quantitative comparison with Visual-TCAV is not straightforward because Visual-TCAV explains one user-specified concept at a time while INS/DEL measure completeness of a full explanation; the authors work around this by evaluating Visual-TCAV with ConEx's aggregation mechanism, leaving a fully controlled head-to-head comparison as an open question.
Target Audience
- XAI researchers working on concept-based interpretability, CAVs, or attribution methods, who will care about the VCM and CCM metrics and the centroid-difference CAV construction.
- Practitioners auditing deployed vision models who need post-hoc explanations without retraining or manual concept annotation, and who need to know whether the framework covers their architecture (CNNs for spatial explanations, both CNNs and ViTs for global explanations).
- Computer vision engineers working with CNNs and Vision Transformers who want tooling that handles architecture-specific embedding issues such as masking artifacts for CNNs and patch-token selection for ViTs.
- Human-computer interaction and evaluation researchers interested in the human participant protocol (58 participants, 300 distinct samples, metrics SA / PRC / INNS / INTS) as a template for measuring explanation understandability.
Authors’ abstract
Many visual explanation methods in computer vision highlight pixel importance but struggle to link these low-level cues to semantically meaningful concepts, limiting their interpretability and trustworthiness. We introduce Concept-based Explanations (ConEx), a novel framework that bridges saliency visualization with concept-based reasoning to provide both faithfulness and interpretability. ConEx automatically discovers class-specific concepts and represents them through concept activation vectors (CAVs), learned without manual supervision using an architecture-specific masking mechanism that reduces noise introduced by the segmentation masks to enhance concept purity. ConEx generates faithful saliency maps that reveal where each concept appears in the image and how it contributes to the prediction. To evaluate the reliability of these learned concepts, we propose two complementary metrics, Vector-Concept Match (VCM) and Concept-Class Match (CCM), that quantify concept alignment and enable direct comparison with existing methods. Extensive experiments across diverse settings demonstrate that ConEx achieves state-of-the-art performance on faithfulness, segmentation, and concept-quality benchmarks. Overall, ConEx advances the field toward truly interpretable and concept-grounded explanations in vision models.