Skip to content
AI.info

Research

WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

Overview Research area: Computer Vision / Autonomous Driving — specifically datasets and evaluation metrics for vision-based end-to-end (E2E) driving. Technical level: Intermediate. The paper assumes

WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
arXiv
2510.26125
Published
2025-10-30
Authors
Runsheng Xu, Hubert Lin, Wonseok Jeon, Hao Feng, Yuliang Zou, Liting Sun, John Gorman, Ekaterina Tolstaya, Sarah Tang, Brandyn White, Ben Sapp, Mingxing Tan, Jyh-Jing Hwang, Dragomir Anguelov

AI summary

Overview

Research area: Computer Vision / Autonomous Driving — specifically datasets and evaluation metrics for vision-based end-to-end (E2E) driving.

Technical level: Intermediate. The paper assumes familiarity with autonomous driving pipelines, open-loop versus closed-loop evaluation, and metrics such as ADE and PDMS.

Scope: This paper introduces WOD-E2E, a Waymo Open Dataset release of 4,021 long-tail driving segments, together with a new human-preference-based open-loop metric called the Rater Feedback Score (RFS).

What This Paper Is About

Existing end-to-end driving benchmarks mostly contain nominal, everyday driving scenarios, so they do not stress-test E2E systems on the rare, safety-critical situations that matter most. Their standard open-loop metrics, such as Average Distance Error (ADE) and PDMS, also struggle in these settings: ADE compares a prediction to a single logged future even though driving is inherently multi-modal, and PDMS depends on annotated agent positions and trajectories that are impractical for novel obstacles such as a flock of birds. The authors address both gaps by releasing a dataset curated specifically for long-tail events occurring with a frequency of less than 0.03%, and by proposing a metric that scores a predicted trajectory against multiple expert-rated trajectory preferences.

Key Contributions

  1. WOD-E2E dataset. A new open dataset for benchmarking end-to-end autonomous driving, containing 4,021 challenging driving segments totaling approximately 12 hours, each drawn from real-world long-tail scenarios that occur with a frequency of less than 0.03% in daily driving. Every segment includes high-level routing information, ego states, and 360-degree camera views from 8 surrounding cameras.

  2. Rater Feedback Score (RFS). A novel, human-aligned open-loop metric designed to assess E2E driving performance in long-tail scenarios by measuring how closely a predicted trajectory matches rater-annotated trajectory preference labels, addressing what the authors describe as the limitations of ADE and PDMS.

  3. Released rater preference labels. Human preference labels are provided for all WOD-E2E validation set segments, while the held-out test set labels were used for the 2025 WOD-E2E Challenge.

  4. Baseline and leaderboard analysis. The paper presents a baseline E2E model (NaiveEMMA) and a comparison and analysis of multiple methods submitted to a public leaderboard, categorized into MLLM-based, diffusion-based, and MLP-based approaches.

Main Findings

  • Dataset composition and splits: WOD-E2E contains 4,021 driving segments, each 20 seconds long, partitioned into 2,037 training segments, 479 validation segments, and 1,505 test segments. Training data spans 20 seconds, whereas testing data covers 12 seconds with the subsequent 8 seconds of future data hidden for evaluation.

  • Sensor and input specification: Each segment provides images from eight cameras (front, front left, front right, side left, side right, rear, rear left, and rear right) at 10Hz, camera intrinsics and extrinsics, ego past trajectory over 4 seconds at 4Hz, velocity and acceleration, a future 5-second trajectory (training and validation only), and a high-level routing command encoded as one of {GO_STRAIGHT, GO_LEFT, GO_RIGHT}.

  • Rarity quantified by an LLM scorer: Gemini 2.5 Pro was used to assign each test set scene a rarity score from 0-100. The WOD-E2E rarity curve is described as significantly higher than all other datasets across all percentage tiles, maintaining an average score of around 93 for the most extreme 10% of the data.

  • Mining funnel: From a driving log corpus of 6,391,012 miles, automated mining identified only 6,888 miles (0.1%) fitting the long-tail criteria. A subsequent human filtering step with a 30% conversion rate reduced the overall portion of long-tail scenarios to 0.03%.

  • Eleven scenario clusters: Mining categories are Construction, Intersection, Pedestrians, Cyclists, Multi-Lane Maneuvers, Single-Lane Maneuvers, Cut-ins, Foreign Object Debris, Special Vehicles, Spotlight, and Others. The Intersections, Foreign Object Debris, and Pedestrians clusters account for the largest share of the dataset.

  • Behavior distribution: Turning behaviors at intersections (left and right, in roughly equal proportions) make up approximately 30% of the data, lane changes account for 10.3%, and on-ramp behaviors account for 1.7%. The majority of behavior is moving straight, which includes lane-following as well as hard braking and swerving.

  • Rater labeling protocol: For each critical moment, raters rate three distinct 5-second future trajectories on a 0-10 scale across five dimensions: Safety, Legality, Reaction Time, Braking Necessity, and Efficiency. Each trajectory starts at a base score of 10, with 2-point deductions for major infractions (safety, reaction time, legal violations) and 1-point deductions for minor infractions (braking necessity, efficiency). At least one rater-specified trajectory is guaranteed a score higher than 6.

  • Rank-1 separation in the label distribution: The Rank 1 trajectory has a strong bias toward optimal behavior with its lowest observed score being 6, while Rank 2 and Rank 3 trajectories span a much wider range, with substantial Rank 3 data falling below a score of 6.

  • RFS metric definition: A trust region is defined around each rater trajectory at t in {3, 5} seconds. Base thresholds are τ_lat = 1.0 and τ_lng = 4.0 at t = 3, and τ_lat = 1.8 and τ_lng = 7.2 at t = 5, with the longitudinal threshold always 4 times the lateral threshold. These are scaled by a piece-wise linear function of initial speed v: 0.5 for v < 1.4, 0.5 + 0.5 × (v − 1.4)/(11 − 1.4) for 1.4 ≤ v < 11, and 1 for v ≥ 11. A prediction inside the trust region receives the flat rater score; outside, it receives the rater score exponentially decayed by 0.1 raised to the excess error ratio. The final score takes the maximum over rater trajectories, averages over t = 3 and t = 5, and is floored at 4.

  • RFS validation on an internal long-tail test split: Scores increase monotonically as model capabilities are added — baseline 7.14, plus WOD-E2E finetuning 7.22, plus multi-camera inputs 7.30, plus test-time scaling (multi sampling) 7.39.

  • Baseline model: NaiveEMMA is a highly simplified version of EMMA, fine-tuned directly from Gemini Flash and trained exclusively on the released WOD-E2E training split. It concatenates all eight camera images into a single 768 × 768 resolution image, takes 3 seconds of past ego-status history and the high-level routing input, and does not use past camera frames. It omits generalist task training mixtures, Chain-of-Thought reasoning, and test-time scaling.

  • Leaderboard results: Selected representative submissions include Swin-Trajectory (MLP-based, Swin Transformer, 36M parameters) with RFS 7.543 and ADE 2.814; DiffusionLTF (diffusion-based, DiffusionDrive, 60M) with RFS 7.717 and ADE 2.977; UniPlan (diffusion-based, DiffusionDrive, 60M) with RFS 7.779 and ADE 2.986; Baseline (Gemini1 Nano, 3B) with RFS 7.528 and ADE 3.018; AutoVLA (Qwen2.5, 3B, SFT+RL) with RFS 7.556 and ADE 2.958; HMVLM (Qwen2.5, 3B) with RFS 7.736 and ADE 3.071; and Poutine (Qwen 2.5, 3B, SFT+RL) with RFS 7.986 and ADE 2.741, the highest RFS reported among these.

  • RFS versus ADE correlation: Using 19 leaderboard submissions, the authors observe only a mild positive correlation between RFS and ADE.

  • Qualitative RFS behavior: Example cases show perfect RFS of 10.0 when the prediction aligns with the best-rated trajectory inside the trust region, RFS of 8.0 when it matches the preferred path at a complex intersection, and decayed or floor scores (RFS = 4) when predictions fall outside the trust regions or diverge entirely from rater-specified maneuvers.

Methodology in Plain English

The authors start from a very large corpus of real driving logs spanning millions of miles, the vast majority of which is ordinary driving. To isolate the rare events they care about, they build a mining pipeline that combines rule-based heuristics with multimodal large language models. Rules defined over the dataset's existing auto-labels (3D detection, mapping, tracking, and prediction) sort logs into 11 long-tail categories, and Gemini is used to search the database for scenarios containing particular long-tail objects.

Mined candidates then pass through a human filtering round to remove scenes that are not genuinely long-tail. The surviving segments go to a labeling pipeline with three stages: raters first watch each video and select the critical moment — the earliest frame where the critical event is visually apparent in the camera images, chosen so that the vehicle has already begun responding and reaction bias from motion history is avoided — and document their rationale. A separate machine learning model, such as Wayformer, generates up to 64 diverse candidate trajectories from perception, mapping, and agent prediction inputs; these are bucketed by driving decision and sampled down to typically fewer than 12 candidates, from which raters pick three (the optimal path plus plausible alternatives and suboptimal behaviors) for grading.

Grading happens in a visualization tool where raters can step through timestamps and see how each candidate interacts with logged agent behavior and map elements. Trajectories are scored from 0 to 10 based on the five criteria, with deductions applied for infractions. Finally, the RFS converts these ratings into an automatic metric: a model's predicted trajectory is compared against the three rated reference trajectories, given the reference score if it stays inside the scaled trust region, and an exponentially decayed score if it does not.

Why This Matters

Impact on research. The paper argues that the absence of long-tail examples in existing benchmarks has hindered accurate evaluation of the robustness and generalization of E2E driving systems. By releasing a dataset explicitly mined for these scenarios plus a metric that reflects human preference rather than proximity to a single logged future, the authors provide a way to distinguish models that are merely good at nominal driving from models that handle genuinely difficult situations. The breadth of leaderboard participation across MLLM, diffusion, and MLP/Swin architectures is presented as evidence of the dataset's usefulness. The paper also frames RFS as an alternative to ADE, which cannot represent multi-modality, and PDMS, which the authors argue is impractical for novel objects and over-penalizes reasonable off-road deviation during emergency avoidance.

Real-world applications:

  • Evaluating and comparing autonomous driving stacks on safety-critical, rare situations such as unprotected intersections, construction zones, pedestrians emerging under occlusion, reckless cut-ins, and debris or animals on the road.
  • Testing whether multimodal large language model-based driving agents generalize to situations they were unlikely to encounter in nominal training data.
  • Providing human-preference supervision for training or fine-tuning trajectory planners that must choose among several defensible behaviors.
  • Serving as a benchmark for regulatory or internal safety review of how a driving system behaves in emergency maneuvers where comfort should be secondary to safety.

Industry relevance. The dataset comes from Waymo and is drawn from real fleet driving logs spanning a mixture of autonomous and manual driving, meaning the scenarios reflect operational conditions rather than simulation. The rater preference labels encode expert judgment about what counts as acceptable driving, which is directly relevant to companies that need to assess planner behavior where a single ground-truth trajectory does not exist.

Future Directions

  • Extending the RFS framework or dataset to closed-loop or reactive evaluation, since WOD-E2E is an open-loop benchmark and the paper's metric is defined on predicted trajectories rather than interactive execution.
  • Broadening rater preference labeling beyond the validation set, since the released labels cover validation only and the test set labels were reserved for the 2025 WOD-E2E Challenge.
  • Expanding the long-tail taxonomy and mining coverage across more cities and scenario types, given that city names are anonymized, the data is predominantly from cities L, K, and J, and remaining cities appear only in the test set.
  • Improving model performance on the hardest long-tail cases, particularly given that the top reported leaderboard RFS values cluster in the 7.5 to 8.0 range and the metric is floored at 4.

Target Audience

Researchers and engineers working on end-to-end autonomous driving, trajectory planning, and driving benchmarks; teams developing multimodal large language model-based driving agents who need a hard evaluation set; and practitioners who design or audit evaluation metrics for safety-critical driving behavior. Readers interested in human-in-the-loop labeling methodology and the use of LLMs for dataset curation and rarity scoring will also find the paper relevant. A working understanding of open-loop metrics such as ADE and of camera-based driving architectures is helpful before reading.

Authors’ abstract

Vision-based end-to-end (E2E) driving has garnered significant interest in the research community due to its scalability and synergy with multimodal large language models (MLLMs). However, current E2E driving benchmarks primarily feature nominal scenarios, failing to adequately test the true potential of these systems. Furthermore, existing open-loop evaluation metrics often fall short in capturing the multi-modal nature of driving or effectively evaluating performance in long-tail scenarios. To address these gaps, we introduce the Waymo Open Dataset for End-to-End Driving (WOD-E2E). WOD-E2E contains 4,021 driving segments (approximately 12 hours), specifically curated for challenging long-tail scenarios that that are rare in daily life with an occurring frequency of less than 0.03%. Concretely, each segment in WOD-E2E includes the high-level routing information, ego states, and 360-degree camera views from 8 surrounding cameras. To evaluate the E2E driving performance on these long-tail situations, we propose a novel open-loop evaluation metric: Rater Feedback Score (RFS). Unlike conventional metrics that measure the distance between predicted way points and the logs, RFS measures how closely the predicted trajectory matches rater-annotated trajectory preference labels. We have released rater preference labels for all WOD-E2E validation set segments, while the held out test set labels have been used for the 2025 WOD-E2E Challenge. Through our work, we aim to foster state of the art research into generalizable, robust, and safe end-to-end autonomous driving agents capable of handling complex real-world situations.

Read the original paper