Research
R3ST: A Synthetic 3D Dataset With Realistic Trajectories
Overview Research area: Computer vision — synthetic dataset generation for traffic analysis, road-agent perception, and trajectory forecasting. Technical level: Intermediate. Scope: The paper introduc

- arXiv
- 2512.16784
- Published
- 2025-12-18
- Authors
- Simone Teglia, Claudia Melis Tonti, Francesco Pro, Leonardo Russo, Andrea Alfarano, Leonardo Pentassuglia, Irene Amerini
AI summary
Overview
- Research area: Computer vision — synthetic dataset generation for traffic analysis, road-agent perception, and trajectory forecasting.
- Technical level: Intermediate.
- Scope: The paper introduces R3ST (Realistic 3D Synthetic Trajectories), a synthetic 3D urban-intersection dataset of more than 80K frames in which vehicles follow real human-driven trajectories taken from the aerial drone dataset SinD, paired with multimodal annotations and a small benchmark of pre-trained models.
What This Paper Is About
Real-world traffic datasets capture authentic driver behavior but are expensive to collect and are typically missing precise ground-truth annotations, while synthetic datasets can be annotated perfectly and cheaply but usually move vehicles using AI models or rule-based systems that do not resemble real driving. The authors' goal is to close that gap by rendering photorealistic virtual intersections in Blender and importing actual vehicle trajectories recorded from drone footage, so that a synthetic dataset inherits real human motion while keeping exact per-frame labels.
Key Contributions
- A new synthetic dataset, R3ST, containing photo-realistic renderings of two different urban intersections, each captured from four camera views, for more than 80K frames, and covering five vehicle types: cars, trucks, buses, motorcycles, and bicycles.
- Real trajectories embedded in a synthetic environment. Instead of AI- or rule-based motion, R3ST uses real-world vehicle trajectories derived from two of the four scenarios proposed by SinD, a bird's-eye-view dataset recorded from drone footage. The two modeled scenarios are an intersection in Tianjin city and an intersection in Chongqing city. The authors state that, to the best of their knowledge, no existing work has explored synthetic datasets that integrate real trajectories.
- Multimodal per-frame annotations. Using Vision Blender, the authors compute instance segmentation and depth annotations, and derive each object's 3D bounding box from the Blender world environment and project it onto the image plane to obtain 2D boxes, organized in YOLO format (class label plus normalized center coordinates cx, cy and normalized width and height).
- Benchmarking of pre-trained models on the dataset for object detection, instance segmentation, and monocular depth estimation, including a fine-tuned YOLO11-large detector.
Main Findings
- Detection performance of fine-tuned YOLO11-large: On the R3ST test set, which is composed of 900 images, the model reaches an overall mAP@50 of 0.989 and overall mAP@50-95 of 0.965. Per-class results are Car 0.995 / 0.986, Van 0.978 / 0.925, Motorcycle 0.990 / 0.972, and Bicycle 0.994 / 0.975.
- Fine-tuning was necessary and fast. Testing YOLO11-large directly on R3ST produced uneven performance — good precision for cars but low for motorcycles, which the authors attribute to the poor quality of the mesh used. Fine-tuning for 15 epochs was enough to show how a model trained on real images can generalize quickly to R3ST.
- Detection classes are limited in the results table. The reported table covers Car, Van, Motorcycle, and Bicycle; results for trucks and buses are not reported, and numerical detection results for those vehicle types do not appear in the paper.
- Instance segmentation and monocular depth estimation are shown qualitatively only. Figure 3 presents segmentation results obtained with YOLO-Seg and the SAM2 online demo, and depth estimation performed with AnyDepth and Pixelformer Large pre-trained on KITTI. No numerical metrics for these two tasks are reported in the truncated content.
- Dataset positioning. In the comparison table, R3ST is listed as Synthetic with Street Camera point of view, providing Depth, Instance Segmentation, and Realistic Trajectories. The authors present it as the only dataset in that comparison that combines synthetic image generation with real trajectories; the table lists the dataset's year as 2024.
- Static sensor configuration. Cameras were placed on light poles with a vertical sensor of 22mm, a field of view of 35.3 degrees, and an F-Stop of 2.8, and four cameras were placed in the reconstructed scenes, each framing one of the four traffic directions at the respective intersections.
Methodology in Plain English
The authors built two virtual intersections from scratch in Blender, a 3D editor chosen for the freedom it gives in constructing environments and for how easily external trajectories can be integrated. Free realistic materials, 3D vehicle models, and building models were imported with the BlenderKit library, and vegetation was created procedurally with the free add-on Sampling Tree Generator, which produces trees dynamically from user-defined parameters.
Into those scenes they injected the annotated vehicle paths from SinD. Each SinD trajectory was matched to the corresponding vehicle type, and when the vehicle was a car, one of three different car meshes was selected at random. This is the step that distinguishes R3ST from typical synthetic datasets, whose vehicles move according to AI-driven or rule-based algorithms. The authors also note that, because the source is real drone footage, the motion patterns replicate real-world traffic, and they illustrate the variance of clustered trajectories in a crossroad.
Cameras were then positioned on light poles with the sensor settings above to mimic road-camera imagery, and vehicle animations were rendered from the perspective of each camera. Finally, Vision Blender was used to compute the additional annotations — instance segmentation masks and depth maps — while 2D bounding boxes were produced by projecting the 3D boxes known from the Blender world environment onto the image plane and stored in YOLO format. To demonstrate that the dataset is usable for real tasks, the authors evaluated pre-trained deep-learning models on detection, segmentation, and monocular depth estimation.
Why This Matters
Research impact. The paper targets a known weak point in synthetic traffic data: perfect labels but unrealistic motion. By fusing SinD's human-driven trajectories with a controllable rendered environment, it offers a resource for trajectory forecasting of road vehicles that has both accurate multimodal ground truth and authentic trajectories, and it provides a testbed for studying how models trained on real or synthetic imagery transfer across domains.
Real-world applications.
- Traffic Monitoring Systems at urban intersections, which the paper identifies as the most accident-prone locations in urban environments and as difficult to monitor because of complex layouts and dynamic traffic.
- Trajectory forecasting and motion prediction for road agents, used in autonomous driving and advanced driver-assistance pipelines.
- Road safety analysis and urban planning, where understanding complex urban traffic scenarios informs safer intersection design.
- Perception model training and evaluation for object detection, instance segmentation, and monocular depth estimation, since each frame ships with those annotations.
Industry relevance. The authors argue that surveillance-grade traffic monitoring often cannot rely on HD maps in real time, which limits the usefulness of HD-map-based datasets; a synthetic dataset with realistic trajectories and dense labels is presented as a practical alternative for building sensors, perception stacks, and forecasting components. The work was partially supported by the Italian Ministry of Enterprises and Made in Italy under agreements for innovation in the automotive sector, and partially financed by the European Union — Next Generation EU.
Future Directions
- Expand the dataset with more scenes, moving beyond the two intersections modeled after SinD's Tianjin and Chongqing scenarios.
- Increase realism by including different models for each vehicle type, addressing the mesh-quality issue the authors flag as the likely cause of weaker motorcycle detection.
- Add more challenging light conditions to better address the domain shift problem between synthetic and real imagery.
- Broaden and quantify the benchmark. Segmentation and depth results are presented only qualitatively, no numeric metrics are given for them, and the detection table omits trucks and buses even though the dataset covers those classes, leaving clear openings for more complete evaluation.
Target Audience
Readers who benefit most are computer vision researchers and graduate students working on trajectory forecasting, traffic monitoring, and urban scene understanding; engineers building autonomous driving or intelligent transportation perception systems who need annotated intersection data; and researchers studying synthetic-to-real domain shift, who will find the comparison against real and synthetic datasets such as KITTI, nuScenes, SinD, Argoverse 2, TUMTraf-I, SYNTHIA, SynTraC, Omni-MOT, and SHIFT directly relevant.
Authors’ abstract
Datasets are essential to train and evaluate computer vision models used for traffic analysis and to enhance road safety. Existing real datasets fit real-world scenarios, capturing authentic road object behaviors, however, they typically lack precise ground-truth annotations. In contrast, synthetic datasets play a crucial role, allowing for the annotation of a large number of frames without additional costs or extra time. However, a general drawback of synthetic datasets is the lack of realistic vehicle motion, since trajectories are generated using AI models or rule-based systems. In this work, we introduce R3ST (Realistic 3D Synthetic Trajectories), a synthetic dataset that overcomes this limitation by generating a synthetic 3D environment and integrating real-world trajectories derived from SinD, a bird's-eye-view dataset recorded from drone footage. The proposed dataset closes the gap between synthetic data and realistic trajectories, advancing the research in trajectory forecasting of road vehicles, offering both accurate multimodal ground-truth annotations and authentic human-driven vehicle trajectories.