Research
ROSA: Roundabout Optimized Speed Advisory with Multi-Agent Trajectory Prediction in Multimodal Traffic
Overview Research area: Multi-agent trajectory prediction and cooperative speed guidance for automated/connected driving in multimodal traffic at roundabouts (Intelligent Transportation Systems). Tech

- arXiv
- 2602.14780
- Published
- 2026-02-16
- Authors
- Anna-Lena Schlamp, Jeremias Gerner, Klaus Bogenberger, Werner Huber, Stefanie Schmidtner
AI summary
Overview
Research area: Multi-agent trajectory prediction and cooperative speed guidance for automated/connected driving in multimodal traffic at roundabouts (Intelligent Transportation Systems).
Technical level: Advanced. The paper combines Transformer-based multi-agent trajectory prediction with an autoregressive deployment scheme, occupancy-zone classification, and microscopic traffic simulation.
Scope: The paper presents ROSA, a system that predicts the joint future trajectories of vehicles and vulnerable road users (VRUs) at a roundabout and uses those predictions to issue real-time speed advisories to vehicles approaching and entering it, evaluated on real-world drone trajectory data and in SUMO simulation.
What This Paper Is About
Roundabouts are hard for automated driving because traffic is dense, interactive, and multimodal: vehicles, pedestrians, and cyclists share the same space, and automated vehicles cannot rely on eye contact or body language to resolve uncertainty. Existing roundabout coordination methods typically ignore VRUs entirely and assume fully automated, cooperative traffic, while existing trajectory prediction models rarely predict vehicles and VRUs jointly in roundabout settings. ROSA addresses both gaps by coupling an interaction-aware, multi-agent prediction model with a two-stage speed advisory that reacts proactively to predicted crosswalk and entry occupancy.
Key Contributions
-
A joint, interaction-aware prediction model for roundabouts. ROSA uses a Transformer architecture that predicts the trajectories of vehicles and VRUs together in a multi-agent (bird's-eye view) manner, trained on real-world data, trained for single-step prediction and deployed autoregressively over a five-second horizon to produce deterministic outputs suitable for real-time speed optimization.
-
An ablation on input modalities and a new evaluation of occupancy prediction. The authors quantify how agent-specific motion dynamics and route (exit intention) information affect prediction accuracy, and they evaluate how well predicted positions translate into correct crosswalk and entry occupancy classifications, which is the quantity the speed advisory actually consumes.
-
A proactive speed advisory for multimodal, mixed traffic. ROSA issues real-time speed advisories in response to predicted conflicts with both VRUs and circulating vehicles when approaching and entering a roundabout, designed for a VRU-prioritized setting and applicable to mixed fleets of automated and human-driven connected vehicles.
-
A simulation-based evaluation against perfect foresight. Using 6600 real-world scenarios in SUMO, the authors measure efficiency and safety effects of ROSA under prediction uncertainty (models with motion dynamics, and with motion dynamics plus exit intention) versus ground-truth occupancy.
Main Findings
-
Motion dynamics are decisive for accuracy. A baseline using only position and agent type degrades badly over time, reaching ADE 5.08 m and FDE 10.51 m at a five-second horizon. Adding velocity, tangential and lateral acceleration, and heading (sine/cosine) reduces errors to ADE 1.29 m / FDE 2.99 m at five seconds, which the paper states outperforms previous works on roundabout trajectory prediction for both vehicles and VRUs.
-
Route information helps further. Adding exit intention yields ADE 1.10 m / FDE 2.36 m at a five-second horizon, with the largest gains at longer horizons. The authors interpret this as evidence of the value of connected vehicle data.
-
Occupancy prediction is useful over the full five-second horizon. At one second, both models achieve precision of at least 0.97 and recall of at least 0.98; up to three seconds, precision stays at or above 0.85 and recall at or above 0.80. Beyond that, uncertainty grows: at five seconds for roundabout entries, the motion-dynamics-only model reaches precision 0.56 and recall 0.47, while the model that also uses exit intention reaches precision = recall = 0.75.
-
The dataset is highly imbalanced. Crosswalks are occupied in only 1.3% of cases, reflecting the low number of VRUs in the openDD data; entries are positive in 11.6% of samples. The authors note that accuracy and F1 may be inflated by this imbalance, so they emphasize precision and recall.
-
Only a minority of scenarios can be optimized. Approximately 16% of scenarios in the test dataset involve the ego vehicle encountering an occupied crosswalk and/or roundabout entry.
-
ROSA's headroom under perfect foresight. In the optimizable scenarios, perfect foresight gives reductions of about 17% in BEV energy consumption, 8% in fuel consumption and CO2 emissions, 5% in travel time, 95% in waiting time, and 93% in number of stops. Averaged over all 6600 scenarios, this amounts to roughly 1-3% reductions in energy and emissions, 0.8% in travel time, and 15% in both waiting time and stops.
-
ROSA keeps most of the benefit under prediction uncertainty. With predicted occupancy, energy consumption drops by around 10% (BEV), emissions by about 5%, travel time by about 3%, waiting time by about 66%, and stops by about 63% in optimizable scenarios. The paper describes this as roughly a one-third performance drop relative to perfect foresight.
-
False negatives, not false positives, drive the losses. Missed occupancies prevent necessary optimizations. In non-optimizable scenarios, false positives occur more often with the motion-dynamics-only model, producing increases of +0.31% BEV energy, +1.23% fuel and CO2, +0.43% travel time, and +1.26% waiting time and stops; with exit intention included, increases stay below 0.8%.
-
Exit intention gives a slight but consistent edge. The full model performs slightly better than the motion-dynamics-only variant in the speed advisory evaluation, but the difference is described as only slight.
-
Battery electric vehicles benefit more than conventional internal combustion engine vehicles in the evaluation.
Methodology in Plain English
The authors start from an existing Transformer-based, multi-agent prediction architecture that takes a bird's-eye view of all agents in a roundabout area rather than an ego-centric view. An attention mask restricts each agent's embedding at a given time step to attend to its own history and to all agents at the same time step, cutting the number of attention connections from (N × s)² to N × s + s.
They use the openDD dataset, specifically the urban roundabout rdb1, which is drone-recorded at 30 Hz and includes prioritized VRU crossings at all entries and exits. The data is downsampled to 1 Hz, keeping class label, position, velocity, tangential and lateral acceleration, heading angle, and exit label. Exit intention is derived from each agent's final position relative to the roundabout center; agents that stay inside the roundabout or are pedestrians/cyclists are labeled −1. The data is split 80% training, 10% validation, 10% test, with balanced VRU representation across splits.
Rather than predicting only positions, the model predicts a tuple of position, velocity, tangential acceleration, lateral acceleration, and sine/cosine of heading. Training uses a weighted composite loss: MSE for position and speed, Smooth L1 for accelerations, and MSE plus a geometric regularization term for orientation. The model is trained for single-step prediction with a three-second history, then deployed autoregressively: predicted states are fed back as input to produce the next step, generating five-second trajectories with deterministic (single) outputs.
To evaluate whether the predictions are useful for speed guidance, the authors define three crosswalk zones and three entry zones and treat occupancy as a binary classification problem: an occupied zone is a positive case. This yields 19,944 binary classifications per prediction step and zone type across the test set.
The speed advisory itself is modeled on the logic of GLOSA (Green Light Optimized Speed Advisory), but for roundabouts. A central unit or the vehicle receives past and current trajectory data, predicts occupancy over the horizon, and computes the time-to-arrival at the crosswalk. If the crosswalk is predicted clear, the vehicle keeps its speed; if occupied, an optimal speed is computed as v_opt = 2d/t − v. The entry is then handled in a second stage using the crosswalk-optimal speed and the remaining distance. The algorithm targets arrival at t+1 to maximize the chance of conflict-free passage, respects a maximum deceleration of 2 m/s², and re-runs at one-second intervals.
Evaluation uses the SUMO microscopic traffic simulator. Real-world trajectories from the test set define realistic occupancy states, producing 6600 scenarios. A connected, fully automated ego vehicle approaches from a fixed distance of 250 meters; it is initially controlled by SUMO's default dynamics, decelerating from 50 km/h around 100 meters before the entry. ROSA triggers roughly 46 meters before the VRU crosswalk and 54 meters before the entry, with the two conflict zones about 8 meters apart. The ICE vehicle uses SUMO's standardized EURO4 model, and a BEV is evaluated on the same trajectories; results are compared with and without advisories, and against perfect foresight.
Why This Matters
Impact on research. The paper targets a specific gap: multi-agent models are underrepresented in trajectory prediction, joint vehicle–VRU prediction at roundabouts is underexplored, and existing roundabout models with extensive evaluations usually focus on a single agent type. It also introduces route/exit intention as an input and shows it reduces error, and it evaluates prediction quality through the downstream metric that matters for control (occupancy), not only through displacement error. The source code is released at github.com/urbanAIthi/ROSA.
Real-world applications:
- Connected speed advisory in vehicles, issuing proactive recommendations when approaching a roundabout with prioritized crossings, using trajectories from infrastructure sensors or from connected vehicles acting as Floating Car Observers via V2X.
- Infrastructure-side traffic management, where a central unit predicts occupancy and broadcasts advisories or occupancy information to equipped vehicles.
- Roundabout design and operations assessment, using the simulation setup to estimate efficiency and emission effects of VRU-prioritized crossing design under real-world trajectory data.
- Fleet-level energy and emissions planning, since the evaluation distinguishes BEV energy consumption from ICE fuel consumption and CO2 emissions, and reports waiting time and stop counts.
Industry relevance. The approach is explicitly designed for mixed traffic rather than requiring full automated-vehicle penetration, which the authors contrast with prior coordination work that requires a full AV penetration rate. That makes it relevant to automakers and suppliers building driver-assistance and automated driving features, to municipalities operating connected infrastructure, and to operators evaluating energy and emissions outcomes. The paper notes its data-driven architecture uses no prior assumptions such as graph construction or map inputs, so it can scale to all motorized road users.
Future Directions
-
Multi-vehicle cooperative optimization. The authors suggest extending ROSA to optimize the speed of multiple vehicles cooperatively, where finding the optimum in a high-dimensional decision space is the key challenge, for example with Reinforcement Learning.
-
Broader traffic conditions. The current focus was analyzing ROSA's potential on real-world trajectories; the authors propose complementing this with evaluation of varying traffic conditions through additional real-world datasets or simulation.
-
Perception, V2X, and compliance. ROSA is proposed as an example function for studying the role of different perception approaches (infrastructure compared to FCOs), V2X-related uncertainty and latency, and human compliance in coordinating multimodal and mixed traffic.
-
Real-world validation and data gaps. The authors call for real-world experiments to validate the findings and support practical deployment, and note that stable performance in broader contexts requires more trajectory prediction datasets with sufficient VRU representation.
Target Audience
Researchers and practitioners in trajectory prediction, automated and connected driving, intelligent transportation systems, and traffic engineering, particularly those working on multimodal interaction with vulnerable road users. It is also relevant to engineers building speed advisory or vehicle coordination functions, to simulation and evaluation specialists interested in occupancy-based metrics, and to readers studying multi-agent systems and Transformer-based motion forecasting applied to real-world traffic. The paper is advanced in technical content; readers benefit from familiarity with displacement-error metrics, Transformer models, and microscopic traffic simulation.
Authors’ abstract
We present ROSA -- Roundabout Optimized Speed Advisory -- a system that combines multi-agent trajectory prediction with coordinated speed guidance for multimodal, mixed traffic at roundabouts. Using a Transformer-based model, ROSA jointly predicts the future trajectories of vehicles and Vulnerable Road Users (VRUs) at roundabouts. Trained for single-step prediction and deployed autoregressively, it generates deterministic outputs, enabling actionable speed advisories. Incorporating motion dynamics, the model achieves high accuracy (ADE: 1.29m, FDE: 2.99m at a five-second prediction horizon), surpassing prior work. Adding route intention further improves performance (ADE: 1.10m, FDE: 2.36m), demonstrating the value of connected vehicle data. Based on predicted conflicts with VRUs and circulating vehicles, ROSA provides real-time, proactive speed advisories for approaching and entering the roundabout. Despite prediction uncertainty, ROSA significantly improves vehicle efficiency and safety, with positive effects even on perceived safety from a VRU perspective. The source code of this work is available under: github.com/urbanAIthi/ROSA.