Research
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Overview Research area: Explainable AI (XAI) for video action recognition, combining concept bottleneck models, human pose analysis, and large language models. Technical level: Advanced — assumes fami

- arXiv
- 2511.03725
- Published
- 2025-11-05
- Authors
- Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim, Jinwoo Choi
AI summary
Overview
Research area: Explainable AI (XAI) for video action recognition, combining concept bottleneck models, human pose analysis, and large language models.
Technical level: Advanced — assumes familiarity with concept bottleneck models, video backbones, vision-language dual encoders, and attribution-based explainability.
Scope: The paper proposes DANCE, an ante-hoc framework that explains video action recognition predictions by disentangling them into three concept types — motion dynamics (represented as human pose sequences), objects, and scenes — and evaluates it on four action recognition datasets plus a user study.
What This Paper Is About
Existing explanations for video action recognition models, such as saliency maps or spatio-temporal tubelets, highlight regions of a video in a tangled way, so a user cannot tell whether a prediction came from how the body moved or from the surrounding objects and background. Language-based explanations add structure but struggle with motion, because motion is "tacit knowledge" — intuitively understood but hard to put into words. DANCE addresses this by forcing a model to make its prediction through three explicitly separated concept types, using pose sequences to convey motion in an appearance-invariant way and text to convey objects and scenes.
Key Contributions
-
A structured, motion-aware video XAI framework. DANCE provides explanations by disentangling motion dynamics concepts from spatial context concepts (objects and scenes), producing organized explanations rather than entangled saliency maps.
-
A label-free concept discovery pipeline. Motion dynamics concepts are discovered automatically by clustering human pose sequences extracted from key clips, while object and scene concepts are discovered by querying GPT-4o. No manual concept annotation is required.
-
An ante-hoc concept bottleneck architecture. A concept layer is inserted between a frozen pretrained video backbone and a final linear classifier, with the concept layer's parameters partitioned into disjoint motion, object, and scene blocks, so the model predicts concept activations before the action label.
-
Comprehensive evaluation and practical utility. The paper reports experiments on four datasets, a user study, qualitative comparisons, sanity checks (including reversed videos), failure case analysis, and demonstrations of model debugging and editing without retraining.
Main Findings
-
Interpretability beats baselines in user studies. In pairwise comparisons against three baselines — (i) a CBM using entangled spatio-temporal concepts generated by GPT-4o, (ii) VTCD, a spatio-temporal saliency method, and (iii) a CBM using UCF-101 attributes — more than 70% of responses fell into "ours is much better" or "ours is slightly better" across all three comparisons.
-
Pose-based motion concepts are rated most interpretable. In a second user study comparing three temporal concept types, DANCE's motion dynamics concepts achieved the highest average score of 4.3, with 89.7% of participants rating them 4 or 5. Language-based GPT-4o concepts scored 2.3 and expert-defined UCF-101 attribute concepts scored 3.4.
-
Recognition performance is largely preserved, and sometimes improved. Using the same backbone encoder across methods, DANCE reached 91.1% Top-1 on KTH, 98.1% on Penn Action, 70.7% on HAA-100, and 87.5% on UCF-101. The baseline without interpretability scored 89.7, 97.8, 73.5, and 88.4 respectively — so DANCE is slightly higher on KTH and Penn Action, with a 2.8 point drop on HAA-100 and a 0.9 point drop on UCF-101.
-
Clearer concepts outperform entangled or expert-defined ones. DANCE consistently outperformed (i) a CBM with UCF-101 attributes (86.8 on UCF-101, other datasets not reported), (ii) LF-CBM with entangled language concepts (87.4, 96.3, 66.5, 85.5), and (iii) LF-CBM with disentangled language concepts (89.9, 97.7, 65.3, 83.7).
-
The model is sensitive to temporal direction. In a sanity check, DANCE correctly predicted Bowing FullBody for the original video, and predicted Burpee for the same video played backward based on "standing up"-like motions, confirming it uses motion dynamics as intended.
-
Sample-level intervention can fix mispredictions. When a model predicted Table Tennis Shot largely due to the scene concept "Table tennis club," deactivating that concept caused the model to use motion dynamics concepts and correctly classify the input as Cricket Shot.
-
Class-level intervention recovers accuracy under severe domain shift without retraining. On UCF-101-SCUBA, where test video backgrounds are altered, assigning a weight of 1.0 to an ignored motion dynamics concept for Volleyball Spiking corrected 98 misclassifications at the cost of one additional error, improving accuracy by 2.5 points (84.0% to 86.5%). Further adjusting weights for Golf Swing and Tennis Swing produced a 4.3 point improvement (77.7% to 82.0%).
-
Failure analysis is supported. DANCE explained that a misprediction of Push up stemmed from high contributions from downward body movement.
Methodology in Plain English
The authors start by freezing a pretrained video backbone encoder and extracting a video-level feature vector from each input video.
To build the motion dynamics concepts, they detect keyframes using an off-the-shelf method based on pixel value differences, cut a fixed-length clip centered on each keyframe, and run a 2D pose estimator on every frame to get pose sequences of shape L (clip length) by J (joints) by 2 coordinates. Low-confidence or jumpy pose sequences are filtered out. All pose sequences from training videos are pooled and clustered with an algorithm such as FINCH; each cluster becomes one motion dynamics concept. A video is labeled with a concept if any of its key clips falls into that cluster — an entirely unsupervised step.
To build object and scene concepts, they prompt GPT-4o with the action class names, asking for important physical objects and for common places or backgrounds. Post-processing removes overly long phrases, near-duplicates, and concepts too similar to the action name. Since no human annotates concept labels, they generate soft pseudo-labels by multiplying a text embedding matrix of the concepts with a video embedding from a vision-language dual encoder.
Training happens in two stages. First, the concept layer is trained: the motion branch uses binary cross-entropy because multiple motion concepts can be present at once, while the object and scene branches use a cosine cubed loss against the pseudo-labels. Second, the concept layers are frozen and a final linear classifier is trained with cross-entropy plus a regularization term combining a Frobenius norm and an element-wise ℓ1 norm. Because prediction must pass through the concept layer, explanations come from a single forward pass rather than requiring a post-hoc optimization step. Concept contribution is defined as the product of a concept activation and the concept weight for the predicted class, and motion concepts are visualized using the pose sequence closest to the cluster medoid.
Why This Matters
Impact on research. The paper establishes that disentangling temporal dynamics from spatial context is achievable in video XAI without sacrificing recognition accuracy, and it challenges the assumption that interpretability must come at a performance cost. It also introduces a label-free concept discovery route for video that other researchers can reuse, and it provides evidence through user studies that pose sequences communicate motion better than text descriptions of motion.
Real-world applications (implied by the paper's framing and demonstrations):
- Model debugging and auditing. Practitioners can inspect which concept types drive a prediction and trace misclassifications to specific concepts, as demonstrated in the failure analysis of Push up.
- Post-deployment correction under domain shift. The class-level intervention results suggest systems can recover accuracy on shifted data by adjusting concept-to-class weights without retraining, as shown on UCF-101-SCUBA.
- Sports and coaching analytics. Motion dynamics concepts give an appearance-agnostic view of how a body moves over time, which is directly relevant to skill analysis in actions like Baseball Swing or Basketball Shoot.
- Trustworthy deployment in sensitive settings. Structured, human-readable explanations separating motion from context support transparency and accountability claims in high-stakes video monitoring.
Industry relevance. The architecture is lightweight relative to full end-to-end retraining: the backbone stays frozen and only a concept layer plus a linear classifier are trained. Explanations require no extra optimization pass, which matters for latency-sensitive deployment. The ability to edit model behavior after the fact without retraining is directly useful for production systems that face distribution shift.
Future Directions
-
Extending disentanglement to more concept types or finer granularity. The current framework uses three types; whether tools, interactions, or audio-derived concepts could be added as separate, interpretable channels remains open.
-
Scaling the concept discovery to larger and more diverse datasets. Concepts are discovered per target dataset, and the paper notes concepts can be shared across classes; how the pipeline behaves on very large action vocabularies and long-tail classes is unresolved.
-
Reducing the residual accuracy gap on some benchmarks. DANCE drops 2.8 points on HAA-100 relative to the non-interpretable baseline, so improving concept coverage for fine-grained classes is a clear next step.
-
Broadening the intervention tooling. The demonstrated sample-level and class-level interventions are manual; automated or semi-automated procedures for selecting which concept weights to adjust under domain shift would extend practical usefulness.
Target Audience
Researchers and graduate students in computer vision, video understanding, and explainable AI; engineers building interpretable video recognition systems for sports, surveillance, or healthcare; and anyone interested in concept bottleneck models or in how pose-based representations can substitute for text when describing motion. Readers should already be comfortable with standard deep learning terminology, attribution methods, and vision-language models.
Authors’ abstract
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods based on saliency produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature -- intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets -- KTH, Penn Action, HAA500, and UCF-101 -- demonstrate that DANCE significantly improves explanation clarity with competitive performance. We validate the superior interpretability of DANCE through a user study. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis.