Research
TrajDiff: End-to-end Autonomous Driving without Perception Annotation
TrajDiff: End-to-end Autonomous Driving without Perception Annotation Overview Research area: Computer vision and autonomous driving — specifically end-to-end trajectory planning with generative (diff

- arXiv
- 2512.00723
- Published
- 2025-11-30
- Authors
- Xingtai Gui, Jianbo Zhao, Wencheng Han, Jikai Wang, Jiahao Gong, Feiyang Tan, Cheng-zhong Xu, Jianbing Shen
AI summary
TrajDiff: End-to-end Autonomous Driving without Perception AnnotationOverview
Research area: Computer vision and autonomous driving — specifically end-to-end trajectory planning with generative (diffusion) models.
Technical level: Advanced. The paper assumes familiarity with bird's-eye-view (BEV) representations, transformer encoders, and denoising diffusion probabilistic models.
Scope: The paper introduces a planning framework that trains an end-to-end driving system using only raw sensor inputs and future trajectories, with no perception labels and no predefined trajectory anchors, and evaluates it on the NAVSIM benchmark.
What This Paper Is About
Most end-to-end driving systems secretly depend on perception supervision — object detection or online HD map construction — which requires enormous manual labeling effort. The authors note that a single dataset of 1000 scenarios requires over 7000 hours of manual labeling. TrajDiff's goal is to remove that dependency entirely: the system learns to plan from raw camera and LiDAR inputs plus the ground-truth future trajectory, using a Gaussian BEV heatmap as an internal self-supervised target instead of perception annotations, then generates trajectories through a diffusion process without any handcrafted motion anchors.
Key Contributions
-
Gaussian BEV heatmap as a self-supervised target. Future trajectories are converted into Gaussian BEV heatmaps that encode position and velocity characteristics, allowing a BEV feature to be learned without any perception annotation.
-
A Trajectory-oriented BEV Encoder. A lightweight encoder fuses raw BEV features with ego status (velocity, acceleration, navigation command) to produce both a predicted BEV heatmap and a Trajectory-oriented BEV feature (TrajBEV).
-
Trajectory-oriented BEV Diffusion Transformer (TB-DiT). A diffusion transformer that denoises continuous trajectories conditioned on an ego query and the TrajBEV feature, operating in an anchor-free manner and eliminating handcrafted motion priors such as trajectory anchors or goal points.
-
A demonstration of data scaling in the annotation-free setting. Because no perception labels are needed, the authors show that simply increasing trajectory quantity — even without additional scenario diversity — improves planning performance.
Main Findings
-
Annotation-free state of the art: TrajDiff reaches 87.5 PDMS on the NAVSIM navtest split with camera and LiDAR input, described as state-of-the-art among all annotation-free methods.
-
Data scaling pushes it further: With trajectory scaling up (initial-point variation resampling plus the full navtrain dataset), TrajDiff reaches 88.5 PDMS, which the authors describe as comparable to advanced perception-based approaches.
-
Camera-only variant is competitive: The camera-only TrajDiff scores 86.4 PDMS, improving over the annotation-free methods LAW (83.8 PDMS) by 2.6 PDMS and World4Drive (85.1 PDMS) by 1.3 PDMS.
-
Outperforms perception-dependent methods: TrajDiff exceeds UniAD by 4.1 PDMS (83.4), PARA-Drive by 3.5 PDMS (84.0), VADv2 by 6.6 PDMS (80.9), Hydra-MDP by 4.5 PDMS (83.0), the Transfuser baseline by 3.5 PDMS (84.0), and the diffusion baseline Transfuser-DP by 1.8 PDMS (85.7).
-
Comparable to anchor-based diffusion: With scaling, TrajDiff's 88.5 PDMS matches DiffusionDrive (88.1) and WoTE (88.3) while showing some superior sub-metrics.
-
Every TB-DiT component matters: Removing all three components (TrajBEV, Ego-BEV Interaction, BEV cross-attention) drops performance to 81.6 PDMS. Removing only the Gaussian heatmap supervision drops it to 85.3 PDMS. The Ego-BEV Interaction module contributes 1.4 PDMS and BEV cross-attention contributes 0.4 PDMS on top of the partial configurations.
-
Heatmap radius design matters: A constant radius of 5 yields 87.2 PDMS; a constant radius of 25 degrades to 85.7 PDMS with notable losses in drivable area compliance and ego progress; the velocity-aware radius achieves 87.5 PDMS.
-
Trajectory diversity is real: In best-of-K rollouts, PDMS rises from 88.5 (±0.000) at K=1, to 89.3 (±0.010) at K=3, 89.7 (±0.013) at K=5, and 92.0 (±0.036) at K=10, with the mean standard deviation growing as K increases.
-
Faster inference than the anchor-based baseline at equal steps: On an NVIDIA L20 GPU with 20 denoising steps, TrajDiff (65M parameters) runs at 17.1 FPS versus DiffusionDrive (60M parameters) at 13.5 FPS. TrajDiff with 12 steps runs at 25.6 FPS with 88.3 PDMS; with 40 steps it runs at 9.8 FPS with 88.5 PDMS. With 3 steps it collapses to 38.1 PDMS, showing a minimum step requirement.
-
Robustness to noisy trajectories: Injecting Gaussian noise into 10% of training trajectories barely changes results: 88.5 PDMS without noise, 88.4 with N(0, 0.01), and 88.1 with N(0, 1).
-
Data scaling ablation: Initial-point variation resampling alone gives +0.4 PDMS (87.9); using the full navtrain dataset alone gives +0.6 PDMS (88.1); both together give 88.5 PDMS.
Methodology in Plain English
The system takes camera images and LiDAR at the current timestep, plus ego status (velocity, acceleration, and a navigation driving command). It does not take any future sensor frames — a deliberate contrast with prior annotation-free methods like LAW and World4Drive, which need them.
From the raw sensor inputs, the encoder builds an initial BEV feature using a Transfuser-style backbone. A set of learnable "heatmap queries," initialized at 4× spatial downsampling (N = H×W/16) for efficiency, attends over the concatenated BEV and ego features and is then upsampled into a one-channel heatmap.
The training target for that heatmap comes from the future trajectory itself. Each future ego position becomes the center of a Gaussian, with a standard deviation that scales with the vehicle's velocity at that timestep. The largest Gaussian value at each pixel forms the target map. A Gaussian focal loss penalizes predictions near the positive location less than predictions far from it, on the reasoning that areas close to the trajectory are still drivable.
The predicted heatmap is fused with the raw BEV feature through a channel-fusion network to produce TrajBEV, which becomes the conditioning signal for the diffusion model.
The diffusion transformer (TB-DiT) adds Gaussian noise to continuous trajectories and learns to reverse it. Its noise-prediction network takes four inputs: a learnable ego query, a diffusion timestep embedding, the noisy trajectory, and the TrajBEV feature. The ego query is enriched by an Ego-BEV Interaction module that cross-attends against TrajBEV, and the result is added to the timestep embedding to form the conditioning used for the adaptive layer-norm (adaLN-Zero) modulation. Each TB-DiT block applies temporal self-attention along the trajectory, BEV cross-attention against a compressed set of BEV queries (a Q-former structure), and an MLP. Training loss is a weighted sum of the BEV heatmap loss and the standard diffusion MSE loss, with weights of 200 and 10 respectively.
Training uses the NAVSIM navtrain split's 978 training scenarios for 200 epochs with AdamW on 8 NVIDIA A100 GPUs at batch size 512, a learning rate of 6×10e−4 with cosine scheduling, and inputs of three forward-facing camera images cropped and concatenated to 1024×256 plus a rasterized LiDAR point cloud. Inference uses DDIM with 20 sampling steps by default.
Why This Matters
Impact on research. The paper argues that end-to-end driving's core promise — optimizing planning directly by backpropagation — is naturally suited to training without perception labels. Prior annotation-free methods still depend on multi-frame inputs and dedicated architectures that predict latent BEV or image features, and they do not use future trajectories to build planning-specific targets. TrajDiff removes both dependencies, and the resulting label-free setting unlocks data scaling experiments that were previously expensive to run.
Real-world applications:
- Cheaper autonomous driving development. Reducing reliance on 7000+ hours of labeling for 1000 scenarios lowers the cost barrier for training planning systems.
- Scaling with fleet data. Because only trajectories are needed, driving logs from large vehicle fleets can be put to use without a parallel annotation pipeline.
- Resilience to localization error. The robustness test with Gaussian noise at standard deviation 1 meter into 10% of training trajectories suggests the system tolerates imperfect ground-truth trajectories, which matters in deployment.
- Generative planners for uncertain scenes. Best-of-K sampling (92.0 PDMS at K=10) points toward planners that propose multiple plausible maneuvers rather than a single deterministic output.
Industry relevance. The inference-speed comparison on an NVIDIA L20 GPU (17.1 FPS for TrajDiff versus 13.5 FPS for DiffusionDrive at 20 steps, with similar parameter counts of 65M versus 60M) addresses the practical on-vehicle compute constraints that generative planners face.
Future Directions
-
Integration with other functional modules. The conclusion states the authors will further explore end-to-end paradigms and investigate enhanced integration strategies with other functional modules.
-
Extending data scaling beyond trajectory quantity. The paper shows gains from trajectory resampling and the full navtrain dataset but does not demonstrate performance from adding new scenario diversity, leaving that question open.
-
Reducing the minimum denoising steps. The 3-step configuration drops to 38.1 PDMS while 12 steps reach 88.3 PDMS, so closing the gap between fast and accurate sampling remains an open problem.
-
Turning best-of-K diversity into a single-pass decision. With 10 rollouts reaching 92.0 PDMS, an unresolved question is how to select the best candidate without paying the cost of ten forward passes.
Target Audience
Researchers and engineers working on end-to-end autonomous driving, diffusion-based planning, or BEV perception. The paper is also relevant to practitioners who care about reducing annotation cost and to anyone tracking self-supervised representation learning applied to real-time robotics. Readers without background in diffusion models or BEV representations will need to consult the cited prior work (Transfuser, DiffusionDrive, DiT) first.
Authors’ abstract
End-to-end autonomous driving systems directly generate driving policies from raw sensor inputs. While these systems can extract effective environmental features for planning, relying on auxiliary perception tasks, developing perception annotation-free planning paradigms has become increasingly critical due to the high cost of manual perception annotation. In this work, we propose TrajDiff, a Trajectory-oriented BEV Conditioned Diffusion framework that establishes a fully perception annotation-free generative method for end-to-end autonomous driving. TrajDiff requires only raw sensor inputs and future trajectory, constructing Gaussian BEV heatmap targets that inherently capture driving modalities. We design a simple yet effective trajectory-oriented BEV encoder to extract the TrajBEV feature without perceptual supervision. Furthermore, we introduce Trajectory-oriented BEV Diffusion Transformer (TB-DiT), which leverages ego-state information and the predicted TrajBEV features to directly generate diverse yet plausible trajectories, eliminating the need for handcrafted motion priors. Beyond architectural innovations, TrajDiff enables exploration of data scaling benefits in the annotation-free setting. Evaluated on the NAVSIM benchmark, TrajDiff achieves 87.5 PDMS, establishing state-of-the-art performance among all annotation-free methods. With data scaling, it further improves to 88.5 PDMS, which is comparable to advanced perception-based approaches. Our code and model will be made publicly available.