Research
Towards Safer and Understandable Driver Intention Prediction
Towards Safer and Understandable Driver Intention Prediction Overview Research area: Computer vision for autonomous driving, specifically interpretable driver intention prediction (DIP) and explainabl

- arXiv
- 2510.09200
- Published
- 2025-10-10
- Authors
- Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar
AI summary
Towards Safer and Understandable Driver Intention PredictionOverview
Research area: Computer vision for autonomous driving, specifically interpretable driver intention prediction (DIP) and explainable video understanding.
Technical level: Advanced. The paper assumes familiarity with deep learning, vision transformers, concept bottleneck models, and feature-visualization techniques such as GradCAM and t-SNE.
Scope: The paper introduces the DAAD-X dataset of multimodal driving videos with hierarchical textual explanations, and proposes the Video Concept Bottleneck Model (VCBM), a framework that predicts driving maneuvers while generating spatio-temporally coherent explanations without post-hoc techniques.
What This Paper Is About
Existing driver intention prediction systems can forecast maneuvers such as turning, lane changing, or stopping, but they behave as black boxes: they do not say why an action was predicted, which makes failures hard to diagnose and erodes trust in safety-critical driving systems. The authors address this by building a dataset that pairs driving maneuvers with human-written explanations drawn from both the driver's eye-gaze and the ego-vehicle's front view, and by designing a model that predicts maneuvers and their justifications jointly. The goal is a driving system whose reasoning is legible to a human, so that overlooked obstacles and flawed decisions can be traced back to specific evidence.
Key Contributions
- DAAD-X dataset: A new multimodal, ego-centric driving action anticipation video dataset with hierarchical explanations from both in-cabin eye-gaze and out-cabin ego-vehicle perspectives, providing human-understandable justifications for driving maneuvers.
- Video Concept Bottleneck Model (VCBM): A multimodal, video-aware concept bottleneck framework with learnable token merging (LTM) and a localised concept bottleneck model (LCBM). The authors state this is the first work to propose a concept-based interpretability method explicitly tailored for video-based models.
- Benchmarking across backbones: Qualitative and quantitative evaluation of VCBM on DAAD-X across multiple backbone models, showing that transformer-based models exhibit greater interpretability than conventional CNN-based models.
- Label-anchored multi-label t-SNE: A visualization technique that uses explanations as anchor points in the latent space to illustrate the disentanglement and causal correlation among multiple explanations in a video.
Main Findings
- Transformers are more interpretable than CNNs: In the bottleneck comparison, transformer-based MViTv2 with bottleneck reached 25.35 accuracy and 37.1 F1 on ego-vehicle explanations versus CNN-based I3D with bottleneck at 25.26 accuracy and 36.73 F1, despite I3D having higher raw action accuracy (74.09 versus 63.29). The authors attribute this to video explanation tasks requiring strong temporal understanding across frames.
- Proposed VCBM improves explanation prediction: MViTV2 + LTM with bottleneck achieved the best ego-vehicle explanation results at 31.22 accuracy, 43.86 F1, 29.17 F1(mac), and 49.11 F1(mic). I3D + LTM with bottleneck reached 28.31 accuracy, 39.43 F1, 24.1 F1(mac), and 44.06 F1(mic), improving over plain I3D with bottleneck (25.26 accuracy, 36.73 F1).
- Full token aggregation beats CLS-token summarization: For I3D with bottleneck, using all tokens gave 28.31 accuracy / 39.43 F1 versus 23.15 accuracy / 33.7 F1 when summarizing with the CLS token, and 16.9 versus 24.1 F1(mac).
- Cluster count matters, with an optimum around five: For I3D + LTM with bottleneck, one cluster (equivalent to a global CLS representation) gave 70.78 action accuracy and 24.64 explanation accuracy; five clusters gave 73.21 action accuracy and 28.31 explanation accuracy; seven clusters raised action accuracy to 74.28 but dropped explanation accuracy to 24.64; ten clusters fell to 70.35 action accuracy and 23.92 explanation accuracy. More clusters were said to introduce noisy cluster centers.
- Both proposed modules contribute: In the component ablation on I3D, removing both LTM and LCBM gave 68.1 action accuracy and 11.22 explanation accuracy; LCBM alone gave 74.09 and 25.26; LTM alone gave 72.8 and 26.03; combining both gave 73.21 action accuracy and 28.31 explanation accuracy — the best explanation result, with a slight drop in action accuracy attributed to averaging during token merging.
- Explanation loss helps, up to a point: With the explanation scaling factor λ set to 0, action accuracy was 72.14 and explanation accuracy 0; at λ = 0.1 action accuracy peaked at 74.28 with explanation accuracy 5.71; at λ = 0.5 action accuracy was 73.21 with explanation accuracy 28.31; at λ = 1 action accuracy fell to 70.35 while explanation accuracy rose to 30. Excessive emphasis on explanations therefore slightly degrades action prediction.
- Cropped gaze regions beat overlaid gaze: Without gaze, action accuracy was 68.11 and explanation accuracy 8.77; overlaying gaze gave 71.57 and 14.03; cropping a circular gaze region at radius R = 350 pixels gave the best explanation result of 74.09 action accuracy and 26.42 explanation accuracy (versus 25.26 at R = 450 and 24.64 at R = 550). Larger crops were said to dilute gaze-relevant information.
- Temporal cues are essential for explanations: Under increasing severity of frame reshuffling, MViTv2's action and explanation accuracy dropped significantly compared to I3D, and explanation accuracy dropped more than action accuracy. The reshuffling procedure divides the video into 16 equal segments of ℓ = T/16 frames, merges every s consecutive segments to form M = 16/s merged segments, and uniformly samples s frames from each, always totalling 16 sampled frames.
- Dataset scale and distribution: DAAD-X contains 1,568 video clips of 7 to 15 seconds each, annotated with 17 ego-vehicle explanations and 15 gaze explanations, totalling 2,536 explanations. The distribution is highly unbalanced: the most frequent gaze explanation was "towards the forward direction" (223 occurrences) and the least common was "to the left side" (10 occurrences); the most frequent ego-vehicle explanation was "a left turn coming ahead" (199 occurrences), while "nearing an intersection and traffic light is green" occurred only 7 times for go straight, 6 times for a left turn, and 2 times for a right turn. Stratified sampling was used for a 70% training / 20% validation / 10% test split.
- Annotation quality check: Annotations were shuffled 3 times among annotators, ambiguous cases were resolved by consensus among 10 annotators, and less than 1% of the total videos were found to be incorrectly annotated and subsequently corrected.
Methodology in Plain English
The authors first build a new dataset from the existing DAAD dataset, which already contains multimodal driving video with eye-gaze recordings. They select 1,568 clips and have human annotators watch each one and attach reasoning: a single gaze label chosen from 15 predefined options, plus one or more of 17 ego-vehicle labels describing scene attributes. Because some explanations occur far more often than others, the dataset is split with stratified sampling to keep the classes balanced.
For the model, they start with an existing concept bottleneck idea: instead of letting a network map raw features directly to a driving decision, force an intermediate layer where each unit corresponds to a human-understandable explanation, then combine those explanations with a simple linear layer to make the final prediction. The problem is that this design assumes a single image, not a video.
Their fix has three parts. A dual video encoder processes the gaze video and the front-view video separately, producing token embeddings which are concatenated along the channel dimension. A Learnable Token Merging module then groups similar tokens into a small number of representative cluster centers, using a composite distance that combines feature similarity (cosine distance), spatial distance, and temporal distance, with soft clustering so each token is assigned softly to clusters rather than hard-assigned. Finally, a Localised Context Bottleneck Model feeds all the merged tokens into the bottleneck instead of averaging them first — a "late averaging" strategy — so that each explanation unit sees fine-grained spatial and temporal detail. The training objective sums the maneuver classification loss with a weighted sum of explanation losses, where λ controls how much emphasis explanations receive.
Evaluation uses I3D (pretrained on ImageNet RGB images), VideoMAE with a ViT-B/16 backbone, and MViTv2-B (both pretrained on Kinetics-400). Beyond accuracy and F1 scores, the authors visualize where the model looks using GradCAM and introduce a variant of t-SNE in which each explanation is placed in the 2D space as an anchor point, computed as the average of the features of all samples where that explanation is active.
Why This Matters
Research impact: The work reframes driver intention prediction as an interpretability problem and supplies both a dataset with categorical, temporally grounded explanations and a video-specific concept bottleneck method. It contrasts with BDD-OIA, which offers only single-frame explanations, and BDD-X, whose freeform contextual explanations the authors argue cannot support interpretable models because categorical annotations with precise, distinct mappings are needed. It also extends concept bottleneck models, which the authors note fail to model temporal inputs, into the video domain.
Real-world applications:
- Autonomous vehicles that can explain a maneuver to a passenger or investigator, helping diagnose cases where, for example, a parked vehicle in a blind spot was missed during a left turn at an intersection.
- Driver monitoring and handover systems in semi-autonomous cars, where the vehicle must infer whether the human driver has noticed a hazard based on gaze and scene context.
- Safety auditing and post-incident analysis of near-misses and collisions, where explanations of intent form part of the evidence trail.
- Fleet and insurance telematics that need to distinguish a deliberate maneuver from a mistaken one, using explanations tied to specific gaze and scene cues.
Industry relevance: The paper contends that trust in autonomous driving technology depends not only on performance but on the ability to scrutinize, explain, and refine decisions over time. The authors note the approach relies on a pretrained text encoder in some prior label-free methods but instead uses human-annotated categorically defined concepts, which are more direct to validate; the dataset, code, and models are released publicly, and the work was supported by iHub-Data and Mobility at IIIT Hyderabad.
Future Directions
- Cross-view alignment: The authors report as a limitation that GradCAM activations are predominantly observed in the forward direction, because the method assumes important objects are only considered if visible from both views, i.e. the driver's gaze guides where the model should look in the front view. They propose investigating explicit alignment of both views before token merging, for example using homography, since token merging relies on video frames being at least partially aligned.
- Long-tail explanation coverage: The dataset's explanation distribution is highly unbalanced, with some explanations appearing only a handful of times; expanding and balancing rare explanations such as "nearing an intersection and traffic light is green" is an open direction.
- Temporal ablation as a diagnostic: The frame-reshuffling analysis shows explanation accuracy collapses much faster than action accuracy as temporal order is destroyed. Understanding how to make explanation generation more robust to disrupted temporal ordering remains unresolved.
- Balancing action versus explanation accuracy: The component ablation shows that token merging slightly lowers action accuracy while improving explanations, and increasing λ trades action accuracy for explanation accuracy. Finding a configuration that does not trade one for the other is an open question.
Target Audience
Researchers and graduate students in computer vision, explainable AI, and autonomous driving who are interested in concept-based interpretability and video understanding; dataset and benchmark builders looking for annotation schemes that pair actions with causal explanations; and practitioners in automotive safety, perception engineering, and human-machine interaction who need maneuver prediction systems whose reasoning can be inspected rather than merely trusted.
Authors’ abstract
Autonomous driving (AD) systems are becoming increasingly capable of handling complex tasks, mainly due to recent advances in deep learning and AI. As interactions between autonomous systems and humans increase, the interpretability of decision-making processes in driving systems becomes increasingly crucial for ensuring safe driving operations. Successful human-machine interaction requires understanding the underlying representations of the environment and the driving task, which remains a significant challenge in deep learning-based systems. To address this, we introduce the task of interpretability in maneuver prediction before they occur for driver safety, i.e., driver intent prediction (DIP), which plays a critical role in AD systems. To foster research in interpretable DIP, we curate the eXplainable Driving Action Anticipation Dataset (DAAD-X), a new multimodal, ego-centric video dataset to provide hierarchical, high-level textual explanations as causal reasoning for the driver's decisions. These explanations are derived from both the driver's eye-gaze and the ego-vehicle's perspective. Next, we propose Video Concept Bottleneck Model (VCBM), a framework that generates spatio-temporally coherent explanations inherently, without relying on post-hoc techniques. Finally, through extensive evaluations of the proposed VCBM on the DAAD-X dataset, we demonstrate that transformer-based models exhibit greater interpretability than conventional CNN-based models. Additionally, we introduce a multilabel t-SNE visualization technique to illustrate the disentanglement and causal correlation among multiple explanations. Our data, code and models are available at: https://mukil07.github.io/VCBM.github.io/