Research
nuCarla: A nuScenes-Style Bird's-Eye View Perception Dataset for CARLA Simulation
Overview Research area: Autonomous driving computer vision — specifically bird's-eye-view (BEV) 3D perception datasets and closed-loop simulation for end-to-end (E2E) driving systems. Technical level:

- arXiv
- 2511.13744
- Published
- 2025-11-12
- Authors
- Zhijie Qiao, Zhong Cao, Henry X. Liu
AI summary
Overview
Research area: Autonomous driving computer vision — specifically bird's-eye-view (BEV) 3D perception datasets and closed-loop simulation for end-to-end (E2E) driving systems.
Technical level: Intermediate. The paper assumes familiarity with BEV perception models, the nuScenes dataset format and detection metrics, and CARLA-based closed-loop evaluation, though it explains its dataset construction steps in accessible terms.
Scope: The paper introduces nuCarla, a 1,000-scenario, nuScenes-format camera-based BEV perception dataset generated in the CARLA simulator, validated by training four state-of-the-art BEV models and releasing their pretrained weights.
What This Paper Is About
Most autonomous driving datasets are collected in the real world under non-interactive conditions, which supports open-loop learning but is of limited use for closed-loop testing where perception, planning, and control interact. The authors argue that closed-loop E2E models underperform even simple rule-based planners largely because simulation datasets provide only raw sensor inputs and direct control outputs, leaving no room to learn meaningful intermediate representations such as BEV features. nuCarla addresses this by building a large-scale, nuScenes-compatible BEV perception dataset inside CARLA that can be used directly for closed-loop simulation development.
Key Contributions
-
A nuScenes-style BEV perception dataset in CARLA. nuCarla contains 1,000 driving scenarios (700 train, 150 validation, 150 test), each with 40 frames sampled at 0.5-second intervals, and strictly follows nuScenes naming conventions, annotation structure, file hierarchy, and API compatibility.
-
Validation through four BEV models. BEVFormer, PETR, BEVDet, and FastBEV were trained and evaluated on nuCarla using the official nuScenes detection metrics, all achieving stable convergence and competitive detection results.
-
Pretrained weights released for all evaluated BEV architectures, intended to serve as perception backbones for future E2E autonomous driving research, in the way ResNet or VoVNet serve as standard visual backbones.
-
An upgraded MMDetection3D framework. The authors patched the legacy MMDetection3D-1.0 framework to work with PyTorch 2.7, CUDA 12.8, and modern GPUs such as NVIDIA H100 and GeForce RTX 50 series, while preserving the original model codebase.
Main Findings
-
Validation performance is strong and consistent. On the nuCarla validation set (150 scenarios from seven trained maps), BEVFormer achieved the highest mAP of 0.813 and NDS of 0.778; BEVDet reached mAP 0.811 / NDS 0.753, FastBEV 0.777 / 0.728, and PETR 0.745 / 0.710. All models exceeded mAP and NDS of 0.7.
-
Test performance drops on unseen maps. On the 150 test scenarios from two unseen maps (Town10 and Mcity), BEVFormer scored mAP 0.579 / NDS 0.599, PETR 0.514 / 0.569, BEVDet 0.560 / 0.546, and FastBEV 0.509 / 0.498. The authors attribute the gap to the harder, more densely populated test environments and note that on nuScenes, validation and test performance do not differ significantly.
-
Per-class scores are more balanced than nuScenes. With BEVFormer on the validation set, nuCarla average precision (AP) ranged from 0.714 (pedestrian) to 0.867 (bicycle), compared with nuScenes where car AP was 0.618 and bicycle AP was 0.398. nuCarla scores were car 0.792, truck 0.828, bus 0.826, pedestrian 0.714, motorcycle 0.849, bicycle 0.867.
-
Velocity error is higher in nuCarla. Average velocity error (AVE) is much larger on nuCarla because all actors are actively traveling, whereas nuScenes contains many stationary participants that yield trivial zero-velocity predictions. BEVFormer's AVE for bicycles, for example, was 0.547 on nuCarla versus 0.248 on nuScenes.
-
Dataset scale is comparable to nuScenes. nuCarla contains 459,632 annotated samples across six object classes, against 417,609 actively traveling participants across the same six classes in nuScenes, but with a deliberately more balanced class distribution.
-
Class imbalance in real data hurts rare classes. The paper reports that BEVFormer's mAP on nuScenes is 0.618 for cars but only 0.398 for bicycles, and that trailer and construction vehicle classes score just 0.172 and 0.129 mAP — motivating their exclusion from nuCarla.
-
Closed-loop E2E models lag rule-based baselines. On Bench2Drive, which covers 220 evaluation routes of roughly 20 seconds each, UniAD and VAD pretrained models achieve success rates of only 16.36% and 15.00%. Later methods report MomAD 16.71%, VeteranAD 33.85%, DriveTransformer 35.01%, and Orion 54.62%, while the rule-based PDM-Lite reaches 92.27% using ground-truth perception.
-
Qualitative predictions match ground truth closely. In a Town03 visualization, BEVFormer's predictions matched ground truth in translation, rotation, and size, with the only missed detection being a partially occluded firetruck that is also barely visible to humans.
Methodology in Plain English
The authors generated the dataset rather than collecting it, using CARLA's privileged access to simulation state. Scenarios were drawn from nine maps — Town01 through Town07, Town10, and Mcity Digital Twin — with the 850 training and validation scenarios spread evenly across Town01 to Town07, while Town10 and Mcity were held out for testing in unseen environments. Town08 and Town09 were excluded because they are reserved for the CARLA leaderboard challenge.
Each scenario was assigned one of CARLA's 14 predefined weather configurations at random, covering sunny, cloudy, and rainy conditions across different times of day, producing an approximately uniform weather distribution. Traffic was generated with the default CARLA traffic manager at 125 participants per scenario: 40 cars, 10 trucks, 10 buses, 25 pedestrians, 20 motorcycles, and 20 bicycles. The authors deliberately rebalanced class frequencies away from the real-world distribution in which cars and pedestrians dominate.
The ego vehicle is a Nissan Micra (3.63 m × 1.84 m × 1.50 m), chosen to approximate the nuScenes acquisition vehicle, a Renault Zoe (4.08 m × 1.78 m × 1.56 m). Six RGB cameras at 1600 × 900 resolution — front, front-left, front-right, back, back-left, back-right — mirror the nuScenes sensor setup, with calibration taken directly from CARLA. No LiDAR or radar is included.
A key annotation challenge was visibility filtering: recording ground truth for every simulated participant would generate false positives for objects no camera can see. To solve this, an instance segmentation camera was mounted at the same position and calibration as each RGB camera, assigning unique pixel values to every object so the pipeline could record only participants visible in at least one view. This segmentation process is used solely for annotation and never at training or inference time.
A noted data generation limitation is that parked vehicles and unattended two-wheelers in prebuilt maps are static meshes without accessible ground truth, so the authors edited the CARLA source in Unreal Engine 4 to remove problematic meshes and recompiled the package, producing a custom build that differs from the official release.
All four BEV models were trained from scratch for 24 epochs on H100 GPUs using the 700 training scenarios, and evaluated with the official nuScenes detection metrics — mAP and NDS, plus ATE, ASE, AOE, AVE, and AAE.
Why This Matters
Impact on research: The paper targets a structural gap rather than a single model. By releasing both a standardized dataset and pretrained BEV backbones, it lets closed-loop E2E researchers start from a verified perception layer instead of training perception from raw sensors and control outputs. The full nuScenes format compatibility means existing camera-based BEV models built for nuScenes can be transferred to CARLA without modification, and the MMDetection3D-1.0 upgrade removes a long-standing environment conflict that has blocked adoption of these codebases on newer hardware.
Real-world applications:
- Closed-loop testing of end-to-end driving stacks in simulation before any on-road deployment, where failures are frequent and testing is otherwise risky and non-reproducible.
- Safety-focused evaluation of perception robustness across nine diverse maps, 14 weather conditions, and varied times of day, including a high-fidelity digital twin of a real autonomous vehicle test facility.
- Pre-training or fine-tuning perception models for rare, safety-critical classes such as pedestrians, motorcycles, and bicycles that are underrepresented in real-world datasets.
- Benchmarking generalization by evaluating on held-out maps with different visual and traffic characteristics.
Industry relevance: The persistent finding that rule-based PDM-Lite (92.27% success with ground-truth perception) outperforms every learned E2E model underscores that perception quality and intermediate representation learning remain bottlenecks. A verified BEV-level dataset with released weights gives industry teams a common starting point for perception modules and for diagnosing whether closed-loop failures originate in perception, prediction, or planning.
Future Directions
-
Adding sensing modalities. The authors plan to extend the dataset with LiDAR and radar and to verify their accuracy through correspondence algorithms. The current release is camera-only.
-
Measuring how nuCarla perception backbones improve closed-loop success. The paper establishes the dataset and pretrained models but does not report closed-loop driving success rates on nuCarla itself; whether these BEV backbones translate into higher SR than Bench2Drive's reported 16.36% (UniAD) and 15.00% (VAD) remains an open question.
-
Restoring the omitted object classes. Construction vehicle, trailer, barrier, and traffic cone were excluded — the first two because CARLA lacks the necessary blueprints, the latter two because no consistent placement method was found. Extending the class set would broaden coverage.
-
Improving traffic behavior realism. The authors relied on CARLA's default traffic manager and did not adopt advanced traffic control workflows, leaving open whether more realistic behavior modeling would change what perception models learn.
-
Standardizing the custom CARLA build. Because removing problematic static meshes required a recompiled CARLA package that differs from the official release, custom data collection by other users may be complicated — a reproducibility issue worth resolving.
Target Audience
Researchers and engineers working on BEV perception, 3D object detection, and end-to-end autonomous driving who need a standardized simulation dataset for closed-loop development; teams adopting or migrating MMDetection3D-1.0-based codebases to modern PyTorch and GPU hardware; and benchmarking groups interested in cross-dataset evaluation of perception models between the real world (nuScenes) and simulation (CARLA). Readers without background in BEV representations or nuScenes metrics will need supplementary reading, since the paper assumes that context.
Authors’ abstract
End-to-end (E2E) autonomous driving heavily relies on closed-loop simulation, where perception, planning, and control are jointly trained and evaluated in interactive environments. Yet, most existing datasets are collected from the real world under non-interactive conditions, primarily supporting open-loop learning while offering limited value for closed-loop testing. Due to the lack of standardized, large-scale, and thoroughly verified datasets to facilitate learning of meaningful intermediate representations, such as bird's-eye-view (BEV) features, closed-loop E2E models remain far behind even simple rule-based baselines. To address this challenge, we introduce nuCarla, a large-scale, nuScenes-style BEV perception dataset built within the CARLA simulator. nuCarla features (1) full compatibility with the nuScenes format, enabling seamless transfer of real-world perception models; (2) a dataset scale comparable to nuScenes, but with more balanced class distributions; (3) direct usability for closed-loop simulation deployment; and (4) high-performance BEV backbones that achieve state-of-the-art detection results. By providing both data and models as open benchmarks, nuCarla substantially accelerates closed-loop E2E development, paving the way toward reliable and safety-aware research in autonomous driving.