Research
DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
Overview Research area: End-to-end autonomous driving (E2E-AD), specifically trajectory planning for the ego vehicle using computer vision and generative models. Technical level: Advanced. The paper a

- arXiv
- 2511.17150
- Published
- 2025-11-21
- Authors
- Liuhan Yin, Runkun Ju, Guodong Guo, Erkang Cheng
AI summary
Overview
Research area: End-to-end autonomous driving (E2E-AD), specifically trajectory planning for the ego vehicle using computer vision and generative models.
Technical level: Advanced. The paper assumes familiarity with bird's-eye-view (BEV) perception, transformer decoders, cross-attention/deformable attention, and diffusion models with DDIM sampling.
Scope: The paper proposes DiffRefiner, a two-stage coarse-to-fine planner that generates anchor-based trajectory proposals and then refines them with a semantically conditioned diffusion model, evaluated on the open-loop NAVSIM v2 and closed-loop Bench2Drive benchmarks.
What This Paper Is About
Predicting where a self-driving car will drive next is inherently multimodal — at any moment several maneuvers may be reasonable. Discriminative planners that regress a single trajectory average across these options and generalize poorly, while diffusion-based planners that start from random noise or fixed anchors need many denoising iterations and are not adapted to the current scene. DiffRefiner's goal is to combine the strengths of both: use a discriminative module to produce a good coarse trajectory first, then let diffusion refine it while explicitly attending to road semantics and other traffic participants.
Key Contributions
- A coarse-to-fine planning framework in which a transformer-based Proposal Decoder produces anchor-based trajectory proposals that act as strong priors, and a Diffusion Refiner then optimizes them through generative denoising.
- A fine-grained denoising decoder containing a Scene-Aware Semantic Interaction Module (called the Fine-Grained Semantic Interaction Module, FGSIM) that aligns trajectories with drivable areas and dynamic agents during denoising, using global cross-attention, local deformable attention, and adaptive gating.
- State-of-the-art results on both an open-loop real-world benchmark (NAVSIM v2) and a closed-loop simulation benchmark (Bench2Drive), with ablations validating each component.
- An efficiency argument: because diffusion starts from offset-corrected anchors rather than raw anchors or noise, near-optimal accuracy is reached with a single denoising step.
Main Findings
- NAVSIM v2 record: DiffRefiner achieves 87.4 EPDMS on NAVSIM v2 with a V2-99 backbone, and 86.2 EPDMS with a ResNet34 backbone. The paper reports this surpasses the previous best method by margins of 3.7% (ResNet34) and 1.6% (V2-99). For context, the table lists a Human Agent reference at 90.3 EPDMS.
- Bench2Drive record: DiffRefiner reaches 87.1 DS and 71.4 SR, improving over the previous best learning-based method HiPAD by 0.3 DS and 2.3 SR without model ensembling. Its Multi-Ability Mean is 69.0, versus 66.0 for HiPAD.
- Ablation on the planning framework: Adding the refiner raises EPDMS from 85.0 (proposal only, 57.2M parameters, 12 ms latency) to 86.2 (74.8M parameters, 27 ms latency), an improvement of 1.2 EPDMS. A discriminative refiner instead of a generative one drops performance to 78.3 EPDMS at 27 ms.
- Ablation on refiner components: EPDMS climbs progressively from 82.4 (planning token only) to 85.0 as the agent token, BEV modulation, drivable area, and traffic participant inputs are added one at a time.
- Ablation on FGSIM: Global cross-attention alone and local deformable attention alone each reach 85.9 EPDMS; naive additive fusion also gives 85.9, whereas the gated fusion reaches 86.2, showing the gating mechanism resolves conflicting global and local information.
- Ablation on denoising steps: 86.20 EPDMS with 1 step, 86.22 with 2 steps, and 86.17 with 5 steps — i.e., a single denoising step is already near-optimal.
- Qualitative results: Visualizations show better collision avoidance and stricter lane-constraint compliance than DiffusionDrive in complex interactive scenarios.
Methodology in Plain English
The system processes synchronized front, left-front, and right-front camera images through a BEV encoder that produces a bird's-eye-view feature map. Two perception heads sit on top of it: a dense segmentation head that labels road elements, dynamic agents, and static obstacles, and a sparse agent head that detects individual objects. The ego vehicle's own velocity, acceleration, and navigation commands are encoded and mixed into the scene context, and a transformer decoder produces two kinds of tokens — a planning token for trajectory generation and an agent token for detection.
For planning, the model starts with 20 offline-clustered trajectory anchors. A lightweight transformer takes each anchor, position-encodes it, and cross-attends it against the planning token, predicting an offset that adjusts the anchor into a coarse proposal.
The diffusion stage then refines those proposals. During training, Gaussian noise is added to the proposal over T steps using a standard schedule; at test time the proposal itself is the starting point rather than pure noise. A denoising network predicts the noise to remove, and a DDIM update produces the next, less noisy trajectory. What makes the denoiser distinctive is the Fine-Grained Semantic Interaction Module: it first extracts semantically meaningful regions from the segmentation output, uses global cross-attention between the trajectory queries and those regions to capture scene-wide context, then uses deformable attention anchored on the trajectory endpoint to focus on locally relevant geometry. A learned sigmoid gate blends the global and local representations. Separate heads then output a refined trajectory and a confidence score, and the highest-scoring trajectory is selected as the final plan.
Training happens in two stages for stability: first the perception network is trained with a Transfuser-style perception loss, then perception and planning are fine-tuned jointly end to end. A winner-takes-all strategy selects the proposal closest to ground truth for the regression loss, which is combined with a classification loss. Training used a batch size of 384 and a learning rate of 4e-4 for 100 epochs on a cluster of 8 NVIDIA RTX 4090 GPUs.
Evaluation used the NAVSIM v2 Navtest split of 12,146 frames and the Bench2Drive closed-loop benchmark with 220 routes spanning 44 interactive scenarios, following the TF++ preprocessing pipeline.
Why This Matters
The paper argues that the choice of initialization for diffusion-based planners is the bottleneck: starting from unstructured noise or unadapted anchors requires many denoising iterations and adds latency that safety-critical driving systems cannot afford. By showing that a discriminative proposal gives the generative refiner a better starting point — good enough that one denoising step suffices — DiffRefiner points to a practical way to get generative multimodality without generative cost. It also argues that explicit semantic grounding (drivable area and agent regions inside the denoising loop) reduces unsafe behaviors such as collisions and map violations, rather than leaving perception-to-planning interaction implicit.
Real-world applications:
- Autonomous driving software stacks, where low-latency multimodal planning is required and probabilistic planners are hard to deploy.
- Robotaxi and advanced driver-assistance systems (ADAS) operating in interactive urban scenarios with pedestrians and other vehicles.
- Closed-loop simulation and validation pipelines such as CARLA-based Bench2Drive, used to certify driving policies before road deployment.
- Transferable motion planning for other embodied agents (warehouse robots, delivery vehicles) that must respect drivable regions and avoid collisions.
Industry relevance: the paper comes from a collaboration between Zhejiang University and Nullmax, an autonomous driving company, and the code is released at https://github.com/nullmax-vision/DiffRefiner. The emphasis on single-step refinement and reported planning latencies of 12 ms to 27 ms (and 40 ms for one configuration) speaks directly to deployment constraints.
Future Directions
- Extending beyond camera-only input. The NAVSIM v2 configuration uses only camera modality, and Bench2Drive follows the TF++ camera pipeline. Whether LiDAR or radar would improve the semantic interaction module is not explored.
- Scaling and adapting the anchor vocabulary. The framework fixes 20 clustered anchors; the paper does not study how anchor count or clustering quality interacts with refinement quality or latency.
- Reducing the refinement overhead further. The refiners add latency (12 ms for a proposal-only model versus 27 ms with the generative refiner), and the paper does not report whether a more targeted refinement can recover that cost.
- Closing the gap to the human reference. The Human Agent row sits at 90.3 EPDMS versus DiffRefiner's 87.4, and the paper does not characterize which scenario categories account for the remaining difference.
Target Audience
Researchers and engineers working on end-to-end autonomous driving, diffusion-based planning, and generative motion prediction. It is most useful to readers who already understand BEV perception and transformer decoders and who want a concrete design for combining discriminative proposals with generative refinement on standard AD benchmarks. Practitioners focused on latency-sensitive deployment will find the denoising-step and latency ablations particularly relevant.
Authors’ abstract
Unlike discriminative approaches in autonomous driving that predict a fixed set of candidate trajectories of the ego vehicle, generative methods, such as diffusion models, learn the underlying distribution of future motion, enabling more flexible trajectory prediction. However, since these methods typically rely on denoising human-crafted trajectory anchors or random noise, there remains significant room for improvement. In this paper, we propose DiffRefiner, a novel two-stage trajectory prediction framework. The first stage uses a transformer-based Proposal Decoder to generate coarse trajectory predictions by regressing from sensor inputs using predefined trajectory anchors. The second stage applies a Diffusion Refiner that iteratively denoises and refines these initial predictions. In this way, we enhance the performance of diffusion-based planning by incorporating a discriminative trajectory proposal module, which provides strong guidance for the generative refinement process. Furthermore, we design a fine-grained denoising decoder to enhance scene compliance, enabling more accurate trajectory prediction through enhanced alignment with the surrounding environment. Experimental results demonstrate that DiffRefiner achieves state-of-the-art performance, attaining 87.4 EPDMS on NAVSIM v2, and 87.1 DS along with 71.4 SR on Bench2Drive, thereby setting new records on both public benchmarks. The effectiveness of each component is validated via ablation studies as well.