Research
Predict and Resist: Long-Term Accident Anticipation under Sensor Noise
Overview Research area: Computer vision for autonomous driving safety — specifically traffic accident anticipation from dashcam video under degraded sensor conditions. Technical level: Advanced. The p

- arXiv
- 2511.08640
- Published
- 2025-11-10
- Authors
- Xingcheng Liu, Bin Rao, Yanchen Guan, Chengyue Wang, Haicheng Liao, Jiaxun Zhang, Chengyu Lin, Meixin Zhu, Zhenning Li
AI summary
Overview
Research area: Computer vision for autonomous driving safety — specifically traffic accident anticipation from dashcam video under degraded sensor conditions.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning (actor-critic methods, policy gradients, value functions), diffusion models, GRU-based sequence modeling, and standard accident-anticipation metrics (AP, TTA, mTTA).
Scope: The paper proposes a unified framework combining diffusion-based feature denoising with a time-aware actor-critic decision module to predict traffic accidents earlier and more reliably than existing methods, both on clean video and on video corrupted by synthetic Gaussian and impulse noise.
What This Paper Is About
Traffic accident anticipation systems try to warn a driver or vehicle seconds before a collision happens, but real deployments face two coupled problems: camera inputs are degraded by rain, glare, blur, and hardware faults, and it is unclear when a warning should fire, since late alerts are useless while premature ones erode trust. This paper treats anticipation as a long-horizon sequential decision problem rather than a frame-by-frame classification task, and adds denoising at both the image and object level so that predictions stay stable when the input is corrupted. The goal is a system that issues early warnings while suppressing false alarms, on both clean and noisy video.
Key Contributions
-
Reframing anticipation as long-horizon credit assignment. The authors formulate accident anticipation as a sequential decision problem and optimize not just whether an accident is predicted, but the timing of the alert, using an actor-critic reinforcement learning framework.
-
Dual-level diffusion denoising. They design image-level and object-level diffusion modules that reconstruct noise-resilient features, allowing the model to retain critical temporal cues under sensor degradation.
-
A unified pipeline that couples robustness with timing. A self-adaptive object-aware attention module, a time-weight layer, and the actor-critic module are combined so that noise resilience and temporal credit assignment reinforce each other rather than being solved separately.
-
Benchmark and noise-robustness evaluation. Across three benchmarks (DAD, CCD, A3D) and their noise-augmented variants, the framework reports state-of-the-art Average Precision (AP) and mean Time-to-Accident (mTTA), with additional ablations isolating each component.
Main Findings
-
Best reported AP and mTTA on all three benchmarks. On DAD the model reaches 91.2% AP and 4.59 s mTTA; on CCD 99.8% AP and 4.29 s mTTA; on A3D 95.7% AP and 4.60 s mTTA. The paper reports AP gains of +0.3% on CCD and +0.6% on A3D over the best baseline, alongside larger mTTA balance improvements.
-
Baselines compared against. The strongest prior model cited, LATTE (IF), reports 89.7% AP / 4.49 s mTTA on DAD, 98.8% / 4.53 s on CCD, and 92.5% / 4.52 s on A3D. Other baselines include DSA, ACRA, AdaLEA, UString, DSTA, GSC, and AccNet. On A3D, LATTE's mTTA (4.52 s) is slightly higher than the proposed model's 4.60 s is reported as a gain, while on CCD LATTE's mTTA (4.53 s) is higher than the proposed model's 4.29 s.
-
Robustness to Gaussian noise. Under moderate corruption (σ = 5.0), the model holds 99.6% AP on CCD and 95.2% AP on A3D. Degradation is gradual at extreme levels: σ = 10.0 gives 98.0% AP / 3.43 s mTTA on CCD and 92.9% / 3.97 s on A3D; σ = 20.0 gives 91.6% / 3.05 s on CCD and 92.9% / 3.96 s on A3D.
-
Robustness to impulse noise. Up to 20% pixel corruption the model stays near baseline (99.6% AP / 4.33 s on CCD; 95.7% / 4.37 s on A3D). At 30% it reaches 99.2% / 4.22 s on CCD and 93.1% / 3.60 s on A3D; even at 50% it produces usable output (98.0% / 3.38 s on CCD; 91.6% / 3.79 s on A3D).
-
Anticipation loss is the single most critical component. Ablating it on CCD collapses AP to 33.3% (mTTA 5.00 s), which the authors note shows that a higher mTTA alone is not necessarily better. Removing the value loss drops performance to 92.8% AP / 3.03 s mTTA.
-
Object-aware and time-weight modules help early warning. Removing the self-adaptive object-aware module gives 99.3% AP / 4.61 s mTTA, and removing the time-weight layer gives 99.5% AP / 4.47 s, versus the full model's 99.8% / 4.29 s on CCD. Removing the policy gradient loss gives 99.6% / 4.47 s.
-
Diffusion modules matter most at moderate noise. Under σ = 5 on CCD, dual diffusion best preserves performance (99.6% AP) while removing modules causes slight drops. Under severe noise (σ ≥ 10), all variants degrade, and the paper reports that omitting image diffusion sometimes improves AP, suggesting over-denoising may harm heavily corrupted inputs.
-
Reward and penalty scaling trades accuracy against earliness. On A3D, raising the reward weight from ×1 to ×50 steadily reduces AP from 95.7% to 92.7% while giving small mTTA gains; reducing it to ×0.1 / ×0.02 raises AP to 96.2% / 95.8% but lowers mTTA to 4.47 s / 4.46 s. A strong ×10 penalty yields the highest mTTA (4.92 s) but the lowest AP (91.2%).
-
Long-horizon beats frame-level in qualitative comparisons. Comparing a history window of 10 against a frame-level window of 0 on three DAD scenarios (threshold 0.5), the long-horizon model produces shorter and less frequent false alarms in a rainy multi-agent scene, predicts nearly one second earlier in a typical collision, and still alerts slightly earlier in a sudden complex crash.
Methodology in Plain English
The system watches a video frame by frame and, at each step, decides whether to raise an accident alert.
Seeing the scene. Each frame goes through a Cascade R-CNN object detector, and up to the top-K dynamic agents are encoded with a VGG-16 backbone. Global image features are also extracted with VGG-16 plus an MLP.
Deciding which objects matter. A self-adaptive object-aware module computes attention weights from the previous hidden state and the object features, so the model focuses on the traffic participants most likely to matter right now, rather than treating all objects equally.
Cleaning up noisy features. Before anything else happens, both the image features and the object features pass through diffusion-based denoising modules. During training, noise is injected into the features at a randomly sampled diffusion step using a variance-preserving schedule with a linear beta schedule from 0.001 to 0.02. A small two-layer feedforward network learns to reverse that noise, and the denoised output is added back to the original feature as a residual, scaled by λ = 0.15. This residual design keeps the original semantics intact while correcting for corruption.
Reasoning over time. The enhanced image and object features are concatenated and fed into a 256-unit GRU that tracks how risk evolves. A rolling buffer of the most recent hidden states is averaged into a summary vector, which smooths out frame-to-frame jitter.
Deciding when to warn. That summary vector feeds an actor-critic module. The actor outputs a distribution over actions and samples one; the critic estimates the expected return from the same state. The reward gives a positive, exponentially discounted value for correct predictions at time t and a fixed negative penalty for mistakes, and rewards are normalized by batch mean and standard deviation to reduce variance. Training combines a supervised anticipation loss (with a time-weight derived from the GRU hidden state and a temporal penalty that rewards earlier correct predictions) with the actor and critic losses.
Training setup. The framework is implemented in PyTorch 2.0 and trained for 30 epochs on an NVIDIA RTX 3050 with batch size 10, using Adam at an initial learning rate of 3×10⁻⁴ with a ReduceLROnPlateau scheduler. Each frame includes up to 19 objects with 4096-D features from the VGG-16 backbone.
Why This Matters
Impact on research. The paper argues that robustness to sensor degradation and the timing of alerts are not separate problems but coupled ones — noisy perception increases the need for long-horizon reasoning, and without good temporal credit assignment a model cannot exploit redundancy across frames to overcome noise. Framing anticipation as reinforcement learning with a time-weighted reward gives the community a way to optimize earliness and reliability jointly rather than trading one off implicitly. The ablation showing that AP collapses to 33.3% without the anticipation loss also serves as a caution: high mTTA by itself is not evidence of a good anticipation system.
Real-world applications:
- Advanced driver assistance and autonomous driving stacks that must decide when to brake or swerve, where a lead time of nearly one second in typical collisions — as reported in the DAD visualization — can enable evasive action.
- Fleet safety and insurance telematics for commercial vehicles operating in rain, glare, or with aging cameras.
- Robust perception testing, using the Gaussian and impulse noise protocols as a stress test for any anticipation model.
- Traffic infrastructure and roadside monitoring systems that must operate under adverse weather.
Industry relevance. Deployable anticipation requires functioning under the imperfect sensors that real vehicles have, not under clean benchmark video. The reported degradation curve — near-baseline performance up to moderate noise and usable output at 50% pixel corruption — is the kind of characterization an engineering team needs before integrating a warning system. The finding that over-denoising can hurt heavily corrupted inputs is also directly actionable for anyone tuning such pipelines.
Future Directions
-
Address over-denoising at severe corruption. The paper observes that omitting image diffusion sometimes improves AP under σ ≥ 10, which raises the question of whether denoising strength should adapt to estimated noise level rather than being fixed.
-
Improve performance on heavily degraded inputs. Performance still declines meaningfully at σ = 20.0 (91.6% AP on CCD, 92.9% on A3D) and at 50% impulse corruption, so robustness is improved but not solved.
-
Reconcile the mTTA trade-off. On CCD the model's mTTA (4.29 s) is below LATTE's 4.52 s, and the reward-scaling experiments show AP and mTTA move in opposite directions; determining the principled operating point for safety-critical deployment remains open.
-
Validate beyond benchmark clips. The evaluation uses 5-second clips from DAD, CCD, and A3D with synthetic noise; whether the framework holds up on longer sequences, real (rather than simulated) sensor degradation, and closed-loop driving is not established in the paper. The paper also points to appendices covering algorithm details, noise visual examples, long-horizon credit assignment robustness, visualizations, window-size comparisons, and inference time, which are not included in the provided content.
Target Audience
This paper is most useful to researchers and graduate students working on accident anticipation, video-based risk prediction, or safety-critical perception; to autonomous driving and ADAS engineers concerned with robustness to sensor degradation; and to readers interested in applying reinforcement learning and diffusion models to temporal decision problems. Readers without background in actor-critic methods or diffusion processes will find the methodology section demanding.
Authors’ abstract
Accident anticipation is essential for proactive and safe autonomous driving, where even a brief advance warning can enable critical evasive actions. However, two key challenges hinder real-world deployment: (1) noisy or degraded sensory inputs from weather, motion blur, or hardware limitations, and (2) the need to issue timely yet reliable predictions that balance early alerts with false-alarm suppression. We propose a unified framework that integrates diffusion-based denoising with a time-aware actor-critic model to address these challenges. The diffusion module reconstructs noise-resilient image and object features through iterative refinement, preserving critical motion and interaction cues under sensor degradation. In parallel, the actor-critic architecture leverages long-horizon temporal reasoning and time-weighted rewards to determine the optimal moment to raise an alert, aligning early detection with reliability. Experiments on three benchmark datasets (DAD, CCD, A3D) demonstrate state-of-the-art accuracy and significant gains in mean time-to-accident, while maintaining robust performance under Gaussian and impulse noise. Qualitative analyses further show that our model produces earlier, more stable, and human-aligned predictions in both routine and highly complex traffic scenarios, highlighting its potential for real-world, safety-critical deployment.