Research
Odyssey: A Closed-Loop Benchmark for Long-Horizon Real-World Driving with Explicit Navigation Routes
Overview Research area: Autonomous driving, specifically closed-loop evaluation of end-to-end (E2E) and vision-language-action (VLA) planners using neural rendering. Technical level: Advanced. Scope:

- arXiv
- 2610.06469
- Published
- 2026-10-05
- Authors
- Jungho Kim, Hongjae Shin, Seunghoon Yu, Heecheol Yoo, Myeongjun Kim, Jiyong Oh, Donghyuk Kwak, Seunghyeop Nam, Haesung Oh, Hyunju Kim, Hyungchan Cho, Jaehyun Park, Soo Won Seo, Jun Won Choi
AI summary
Overview
- Research area: Autonomous driving, specifically closed-loop evaluation of end-to-end (E2E) and vision-language-action (VLA) planners using neural rendering.
- Technical level: Advanced.
- Scope: This paper introduces Odyssey, a closed-loop driving benchmark of 100 scenarios reconstructed from 100-second nuPlan logs that supplies planners with explicit standard-definition (SD) map routes instead of directional commands, and evaluates them with new route-compliance and lane-preparation metrics.
What This Paper Is About
Existing closed-loop driving benchmarks evaluate only short rollouts, typically around 20 seconds, so they cannot reveal whether an earlier decision such as a lane change actually enables a later turn. They also rely on ambiguous high-level commands like "turn right," which do not say which of several right exits to take, and their 3D Gaussian Splatting (3DGS) renderings degrade into artifacts when the simulated vehicle drives to viewpoints that were never recorded. Odyssey addresses all three problems at once: it evaluates planners over continuous 100-second rollouts, gives them a road-level SD-map route as the navigation objective, and refines rendered images with a diffusion model before they are fed to the planner.
Key Contributions
- Odyssey benchmark: A high-fidelity real-world closed-loop benchmark with 100 scenarios curated from nuPlan, each reconstructed from a 100-second driving log, for testing whether E2E planners sustain safe, goal-directed driving under explicit route guidance.
- SD-map route guidance: Routes derived by HMM-based map matching of the logged ego trajectory to OpenStreetMap, which resolve the ambiguity of directional commands and let planners be assessed consistently across complex road topologies. Routes are updated to the planner according to the ego vehicle's current position during the rollout.
- Hybrid rendering pipeline: A closed-loop simulation framework combining OmniRe-based 3DGS reconstruction (with ground modeling, coverage-aware sampling, and a traffic-light appearance model) with single-step diffusion refinement via Fixer, improving rendering fidelity at novel ego poses.
- New evaluation metrics: SD Route Compliance (SDC), Pre-Lane Change Accuracy (PLCA), Pre-Lane Change Score (PLCS), and RouteDS, which extends the Bench2Drive Driving Score with penalties for SD-route deviations and failed pre-lane change evaluations. Benchmark code, evaluation code, and adapted baselines are stated as planned for public release.
Main Findings
- Longer horizons change conclusions: RouteDS decreases as the evaluation horizon grows, and planner rankings reverse. DrivoR achieves higher RouteDS than SafeDrive at short horizons but falls below it as driving continues.
- Route guidance helps road-level adherence: Providing SD routes improves SDC and RouteDS across all evaluated planners in both non-reactive and reactive traffic. With guidance, DiffusionDrive reaches SDC 100 and RouteDS 35.9 (non-reactive) and SDC 100, RouteDS 41.0 (reactive); SafeDrive reaches RouteDS 44.4 and 48.9 respectively.
- Lane preparation does not follow automatically: SD-route guidance improves SDC for all planners, but PLCA decreases for four of the five planners, with only DiffusionDrive showing an increase.
- Collision behavior is not uniformly improved: RouteDS rises for every planner with route guidance, but the collision score decreases for four of the five planners (higher collision scores mean fewer collision penalties).
- Open-loop metrics do not transfer: For SafeDrive, IL scoring outperforms safety scoring (RouteDS 53.9 vs 44.4 non-reactive) despite similar NAVSIM PDMS values, and adding safety filtering raises RouteDS further to 54.4. For ReCogDrive, GRPO fine-tuning improves NAVSIM PDMS to 90.4 but lowers closed-loop performance relative to IL-only training (RouteDS 21.5 vs 37.7 non-reactive).
- NAVSIM PDMS is not a reliable proxy: DrivoR achieves higher PDMS than SafeDrive but lower RouteDS in both traffic settings.
- Diffusion refinement improves perception and driving: Compared to GS-only rendering (open-loop mIoU 31.5, closed-loop mIoU 26.0, RouteDS 53.80), adding diffusion gives 31.9 / 28.2 / 52.62, and fine-tuning on Odyssey gives 33.7 / 29.4 / 54.43. Recorded ground-truth images yield open-loop mIoU 35.4. Pretrained Fixer alone improved mIoU but slightly decreased RouteDS; only the fine-tuned version raised RouteDS above the GS-only baseline.
- Reconstruction ablation: Moving from the OmniRe baseline to the full pipeline (color correction, coverage-aware sampling, ground modeling, pretrained Fixer, Fixer fine-tuning) changes PSNR 24.34 to 25.34, SSIM 0.788 to 0.786, LPIPS 0.399 to 0.233, KID_in 29.5 to 5.6, KID_ex 52.2 to 8.9, and latency 17.8 to 39.1 ms/image on an NVIDIA B200 GPU.
- Traffic-light modeling works: The time-conditioned appearance model reproduces a red-to-green transition that OmniRe fails to recover.
- Command flips cause route deviations: Figure 7 shows SafeDrive deviating from the intended route when the driving command flips, but following the intended route under SD-route guidance.
Methodology in Plain English
The team picked 100 driving scenarios from nuPlan, using RefAV-based tagging plus human review to find sequences with complex maneuvers and interactions. Of these, 13 involve traffic lights and 47 involve pedestrians. The scenarios are chosen to include lane changes followed by turns: 49 such cases, 16 of which have more than 20 seconds between the start of the lane change and the start of the turn, which a 20-second window could not capture.
For each scenario, the full logged ego trajectory is matched to OpenStreetMap with HMM-based map matching to produce a route polyline from start to destination, which is manually checked for continuity. This road-level route is handed to the planner as the navigation objective, and the planner must combine it with camera observations to decide actual lane-level maneuvers.
The 100-second scene is reconstructed as a single 3DGS scene with OmniRe, accelerated by FastGS, from 8,000 images captured at 10 Hz across eight cameras and optimized for 140,000 iterations. Rigid vehicles, SMPL-deformed pedestrians, and a deformation field for other actors are modeled separately; traffic lights use fixed geometry with time-dependent spherical harmonics and opacity. Ground is represented by thin Gaussian disks initialized from LiDAR on a fixed height field, and coverage-aware sampling reweights training views so that rarely observed regions are seen more often.
At each simulation step, the planner predicts a 4-second trajectory at 2 Hz. The simulator drives the ego vehicle along it at 10 Hz using an LQR controller and a bicycle model, then renders new camera views, which are passed through a diffusion refinement model (Fixer, a Cosmos-backbone adaptation of Difix3D+) in a single denoising step before reaching the planner. Traffic is simulated either non-reactively (actors replay recorded motion) or reactively (IDM-based longitudinal control plus intersection-handling logic inspired by CARLA's Traffic Manager), and section-based traffic triggering aligns traffic and signal timing with the ego's actual progress rather than the original timeline.
Baselines LTF, DiffusionDrive, DrivoR, SafeDrive, and ReCogDrive were adapted to take routes: 120 route points sampled at 1-meter intervals ahead of the ego's projection on the route polyline, pooled in groups of five into 24 route tokens with Fourier and order embeddings, attended to via an added cross-attention module. ReCogDrive additionally received a single summarized route token and a route-derived navigation sentence.
Why This Matters
Impact on research: Odyssey reframes how driving planners should be judged — not on isolated trajectories or short windows, but on the downstream consequences of their own decisions over extended rollouts. Its finding that open-loop NAVSIM gains do not translate into closed-loop performance, and that planner rankings change with horizon length, is a direct challenge to common evaluation practice.
Real-world applications:
- Validating end-to-end and VLA driving stacks before road deployment, using routes rather than commands as the navigation interface.
- Testing route-following behavior for robotaxi or delivery fleets that must translate a road-level route into correct lane choices at successive intersections.
- Auditing simulation fidelity, since the paper quantifies how rendering artifacts and traffic-timing mismatches alter planner behavior.
- Benchmarking perception-plus-planning systems under traffic-light state changes and pedestrian interactions.
Industry relevance: The distinction between road-level route adherence and lane-level preparation is exactly the gap autonomous vehicle developers face when integrating consumer-grade SD maps instead of detailed HD maps, since the paper positions SD routes as reducing reliance on HD map inputs while keeping local decisions sensor-based.
Future Directions
- Applying diffusion guidance during 3D Gaussian Splatting optimization itself, to fix blurred or distorted surrounding vehicles viewed from angles not covered by the driving logs.
- Improving the temporal consistency of the diffusion refinement across frames.
- Adding adversarial agents to create more challenging traffic interactions for safety evaluation.
- Developing dedicated reasoning mechanisms that relate SD-map routes to perceived road geometry and required maneuvers, since translating a road sequence into timely lane choices before intersections remains an open problem on Odyssey.
Target Audience
Researchers and engineers working on end-to-end autonomous driving, closed-loop simulation, and neural rendering for robotics, as well as benchmark designers and teams evaluating vision-language-action driving models. Readers need familiarity with planner architectures and simulation-based driving metrics to get the most from the quantitative comparison tables, though the core argument about horizon length and navigation ambiguity is accessible to a broader audience.
Authors’ abstract
Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional commands also obscure the intended navigation objective. We introduce Odyssey, a closed-loop benchmark for long-horizon driving comprising 100 scenarios, each reconstructed from a 100-second nuPlan driving log to preserve the context of navigation maneuvers and traffic interactions. To provide a consistent navigation objective, Odyssey replaces directional commands with explicit standard-definition (SD) map routes that specify which roads to follow, while sensor-based planning determines local driving actions. Throughout these rollouts, diffusion-based refinement of 3DGS-rendered images reduces rendering artifacts along the ego trajectory. To assess how effectively planners follow these routes and prepare for upcoming maneuvers, we introduce SD Route Compliance and Pre-Lane Change Score. These assessments are complemented by RouteDS, which extends the Driving Score with penalties for SD-route deviations and failed lane preparation. We adapt state-of-the-art planners, including vision-language-action (VLA) models, and evaluate their navigation performance using these metrics. Odyssey highlights open questions in route representation and integration for E2E driving. Benchmark code and adapted baselines will be released publicly.