Research
ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
Overview Research area: Robotics / Agentic Embodied AI — specifically the intersection of the Model Context Protocol (MCP), large language models (LLMs), and ROS/ROS 2 robot log data analysis. Technic

- arXiv
- 2511.03497
- Published
- 2025-11-05
- Authors
- Lei Fu, Sahar Salimpour, Leonardo Militano, Harry Edelman, Jorge Peña Queralta, Giovanni Toffetti
AI summary
Overview
Research area: Robotics / Agentic Embodied AI — specifically the intersection of the Model Context Protocol (MCP), large language models (LLMs), and ROS/ROS 2 robot log data analysis.
Technical level: Intermediate. Readers will benefit from basic familiarity with ROS/ROS 2 concepts (bags, topics, TF trees, LiDAR, odometry) and with LLM tool-calling, but the paper is written to be readable for those new to MCP.
One-sentence scope: The paper introduces rosbags-mcp, an MCP server plus a benchmarking UI (MCP Lab) that lets LLMs and VLMs analyze, visualize, and process ROS and ROS 2 bag files through natural language, and it benchmarks the tool-calling ability of eight models on ten robotics analysis tasks.
What This Paper Is About
Robot logs recorded as ROSbags contain large volumes of time-synchronized sensor data, but LLMs cannot natively parse the binary format, often exceed context limits, and lack access to the structured metadata and sensor semantics needed for accurate answers. The authors build an MCP server that gives conversational AI models a set of robotics-aware tools for reading, filtering, analyzing, and plotting ROSbag data, so that a user can interrogate a robot dataset in plain language without expert robotics or programming knowledge. Alongside the server they build a testing interface and use it to measure how well different proprietary and open-source models actually call those tools.
Key Contributions
-
rosbags-mcp, an MCP server for ROS and ROS 2 bag analysis. The authors state that to the best of their knowledge this is the first paper to propose an MCP interface for analyzing and processing ROS native data. It supports native analysis of trajectories, laser scan data, transforms, and time series data, plus an interface to standard ROS 2 CLI tools (ros2 bag list,ros2 bag info) and bag filtering by topic subset or time trim. -
A robotics-domain-grounded toolset. Compared with existing open-source tooling, the authors introduce a wider and more comprehensive toolset organized in three categories: Core Data Access and Management, Domain-Specific Analysis, and Visualization and Plotting. The initial release focuses on mobile robotics.
-
MCP Lab, a lightweight benchmarking UI. MCP Lab is a modular testing environment for MCP servers that supports multiple LLM providers (including remote connections), standardizes access to ROSbag repositories, manages tool invocation, records structured metrics, and provides built-in visualization of plots generated by the MCP server.
-
The first benchmark of tool-calling capabilities for LLMs and VLMs in a robotics data context. The authors frame this as distinctive because their tools are more complex than MCP servers in other domains that act mainly as translators.
Main Findings
-
A large divide in tool-calling capability. Across ten tasks, Kimi K2 and Claude Sonnet 4 were the most consistent, completing all tasks (the conclusion states 100% task completion rates). GPT-5 mini followed closely with 9 successful completions (90% success rate). Qwen3 (32B) and GPT OSS (20B) each completed 7 tasks (70% success rate), struggling on searching for specific data and on frame transformations.
-
Smaller models showed notable limitations. Sonnet 3.7, GPT-4o Mini, and Llama 4 Scout frequently failed to complete multi-step analytical tasks, often producing incomplete results or requiring extensive guidance on tool calling.
-
Frontier models handle multi-tool composition better. The best-performing models showed particular strength on complex tasks requiring combined tool calls. For example, on the trajectory-plotting task, Sonnet 3.7 and GPT-4o Mini answered correctly but additionally invoked
analyze_trajectory, and GPT OSS invoked bothanalyze_trajectoryandanalyze_lidar_scan. -
Response time varies sharply by model. Kimi K2 and Claude 4 had the most stable and predictable response times, with tight distributions around low median values and minimal variance. Smaller models showed substantially higher response time variability and broader distributions, which the authors attribute to suboptimal decision-making rather than efficient tool use.
-
Tool-use efficiency differs. Kimi K2, Claude 4, and Claude 3.7 had the most efficient tool-usage patterns, consistently requiring fewer tools to complete the same tasks, whereas GPT models and other open-source models frequently invoked more tools than necessary.
-
Several factors drive tool-calling success rates. The authors identify: schema clarity (ambiguous descriptions lead to malformed requests); argument complexity (smaller models frequently fail to construct valid parameter sets); context preservation across sequential tool calls (models often forget previously established parameters such as folder paths); and the total number of available tools, which inversely correlates with success rates, particularly when tools share similar functionality or overlapping parameter spaces.
-
Example outputs demonstrate the analysis depth. In showcased conversations, the system reported a closest approach of 0.217 m (21.7 cm) to a target position at coordinates (1.793, –1.933) at timestamp 1755341898.5013318; three moments in the first 30 seconds where angular velocity exceeded 0.4 rad/s (at 6.19s, 16.61s, 25.78s, all at 0.656 rad/s); and a traveled X range from -2.074 to 2.168 meters (span 4.243 m) and Y range from -2.017 to 2.000 meters (span 4.017 m).
-
Positioning against prior tooling. The authors compare their work to Bagel, a "ChatGPT for physical data" platform that supports ROS 1, ROS 2, PX4, and Ardupilot log formats, stating that they provide more specific tooling in addition to an in-depth benchmarking.
Methodology in Plain English
The authors implemented specialized MCP servers in Python using the MCP library, with FastAPI as the underlying web framework for request validation, documentation, and error handling. Each analysis capability is defined as a dedicated function with an explicit input and output schema; these schemas validate user input and give MCP clients the metadata needed to generate interactive documentation. Servers communicate with clients over JSON-RPC, which keeps the interaction language-agnostic and platform-independent.
The tools are grouped into three categories. Core Data Access and Management covers setting a bag path (B1), listing bags (B2), retrieving bag metadata such as topics, counts, and duration (B3), getting a message at a specific timestamp (M1), getting messages in a time range (M2), searching messages with conditions such as regex or equals (M3), and creating a filtered copy of a bag by topic, time, or rate (FB). Domain-Specific Analysis covers trajectory metrics such as distance, speed, and waypoints (AT), LiDAR scan analysis for obstacles, gaps, and statistics (AS), ROS log parsing and filtering by level or node (AL), TF tree retrieval (GT), and extraction of a camera image at a given time as base64 JPEG (GI). Visualization and Plotting covers time series plots (P1), 2D trajectory plots in XY (P2), and LiDAR scans as polar plots (P3). The metric function get_messages_in_range, for example, takes topic, start_time, end_time as required arguments and max_messages (default 100) and bag_path as optional ones.
For evaluation, all experiments ran inside MCP Lab. Eight models received identical task prompts with access to the same ROSbag data and analysis utilities. A task counted as successful if the model called the correct tools and produced valid results without manual intervention; partial success was recorded if the reasoning was correct but required human guidance or additional tool calls beyond the baseline tools. The paper describes ten tasks, from simple bag listing to multi-tool workflows such as finding the maximum commanded linear velocity between two times and reporting when and where it occurred.
Why This Matters
Impact on research. The paper sits at the intersection of two fast-moving verticals — agentic AI and physical/embodied AI — where the authors note the literature "remains scarce." It contributes a reusable, open-source, permissively licensed implementation and, more unusually, a quantitative benchmark of tool calling in a robotics data context rather than a generic one. It also makes a design argument: MCP standardizes tool integration, but model selection and interface design still determine whether an agent works in practice.
Real-world applications:
- Robot debugging and fault investigation: engineers can query logs in natural language to find when a sensor saw an obstacle, when velocity commands exceeded a threshold, or where a path deviated.
- Performance evaluation and regression testing: trajectory metrics, velocity plots, and filtered bag extraction support systematic comparison of robot runs.
- Access for non-experts: operators and analysts without robotics or programming expertise can interrogate datasets through conversation, which is the paper's explicit accessibility goal.
- Selective data handling: bag filtering by topic, time, or rate lets teams trim large recordings before downstream review or cloud processing, echoing the event-driven capture model the authors cite from Heex Technologies.
Industry relevance. The benchmark directly informs deployment choices: if smaller or cheaper models fail multi-step analytical workflows and require more tool calls, teams cannot treat model selection as interchangeable. The finding that the total number of available tools inversely correlates with success rates, especially when tools overlap, is an actionable constraint on MCP server design. The open-source release also targets community-driven development of natural-language interfaces for robotics.
Future Directions
- Systematic ablation studies to dissect successful tool design patterns and establish best practices for MCP-based robotic systems.
- Schema optimization strategies, given the observed link between schema clarity and invocation accuracy.
- Context-aware tool chaining mechanisms, addressing the finding that models forget previously established parameters like folder paths across sequential calls, plus adaptive tool exposure based on model capabilities.
- Real-time robot control and multi-robot coordination, which the authors describe as compelling opportunities for advancing agentic embodied AI applications.
Target Audience
Robotics engineers and ROS/ROS 2 developers who want conversational access to bag data; agentic AI and MCP server developers looking for a concrete robotics-domain tool design and a benchmark methodology; LLM practitioners evaluating model selection for tool-calling pipelines; and researchers working on embodied AI who need a grounded, reproducible reference implementation — the code is available at https://github.com/binabik-ai/mcp-rosbags. Teams choosing between proprietary and open-source models for robotic data analysis will find the per-model comparison table most immediately useful.
Authors’ abstract
Agentic AI systems and Physical or Embodied AI systems have been two key research verticals at the forefront of Artificial Intelligence and Robotics, with Model Context Protocol (MCP) increasingly becoming a key component and enabler of agentic applications. However, the literature at the intersection of these verticals, i.e., Agentic Embodied AI, remains scarce. This paper introduces an MCP server for analyzing ROS and ROS 2 bags, allowing for analyzing, visualizing and processing robot data with natural language through LLMs and VLMs. We describe specific tooling built with robotics domain knowledge, with our initial release focused on mobile robotics and supporting natively the analysis of trajectories, laser scan data, transforms, or time series data. This is in addition to providing an interface to standard ROS 2 CLI tools ("ros2 bag list" or "ros2 bag info"), as well as the ability to filter bags with a subset of topics or trimmed in time. Coupled with the MCP server, we provide a lightweight UI that allows the benchmarking of the tooling with different LLMs, both proprietary (Anthropic, OpenAI) and open-source (through Groq). Our experimental results include the analysis of tool calling capabilities of eight different state-of-the-art LLM/VLM models, both proprietary and open-source, large and small. Our experiments indicate that there is a large divide in tool calling capabilities, with Kimi K2 and Claude Sonnet 4 demonstrating clearly superior performance. We also conclude that there are multiple factors affecting the success rates, from the tool description schema to the number of arguments, as well as the number of tools available to the models. The code is available with a permissive license at https://github.com/binabik-ai/mcp-rosbags.