Skip to content
AI.info

Research

Event-Grounding Graph: Unified Spatio-Temporal Scene Graph from Robotic Observations

Overview Research area: Robotics, specifically semantic scene representation and spatio-temporal memory for autonomous robots (3D scene graphs, robot memory, vision-language models). Technical level:

Event-Grounding Graph: Unified Spatio-Temporal Scene Graph from Robotic Observations
arXiv
2510.18697
Published
2025-10-21
Authors
Phuoc Nguyen, Francesco Verdoja, Ville Kyrki

AI summary

Overview

Research area: Robotics, specifically semantic scene representation and spatio-temporal memory for autonomous robots (3D scene graphs, robot memory, vision-language models).

Technical level: Intermediate. The paper combines 3D scene graph formalisms, graph pruning algorithms, and off-the-shelf vision-language and language models, but the core idea is explained with intuitive examples.

Scope: This paper introduces EGG (Event-Grounding Graph), a scene graph framework built from real robotic observations in which dynamic events are grounded to the spatial elements (objects, rooms) that they change, enabling an LLM to answer free-form spatio-temporal queries about a robot's environment and its history.

What This Paper Is About

Existing robotic scene representations either capture rich spatial and semantic structure without any notion of dynamic events (3D scene graphs such as Hydra), or they record events and interactions as free-form video captions without grounding them to the specific objects involved (video-caption memory systems such as ReMEmbR). The paper's goal is to close that gap: connect the semantic attributes of interactions (the symbolic knowledge of "washing a mug") with the spatial features of the scene (the specific blue mug), so a robot can recall, reason about, and explain what happened in a scene over time. The authors build this representation from real robot data and test whether an LLM agent can use it to answer questions in natural language.

Key Contributions

  1. EGG, a unified spatio-temporal scene graph framework that combines spatial features of the environment (rooms, objects with time-invariant and time-variant attributes) with dynamic event nodes and timestamped event edges that ground each event to the objects it affects.
  2. A graph manipulation and pruning method that exploits EGG's spatio-temporal queryability to extract a task-relevant compressed subgraph from a natural language query, using time pruning, location pruning, object pruning, event pruning, and history expansion.
  3. Experimental evaluation on real robotic data collected with a Hello Robot Stretch 2 platform, showing EGG retrieves relevant information and responds to human queries, and outperforming a purely spatial baseline (Hydra) and a spatio-temporal memory baseline (ReMEmbR) across all reported modalities.
  4. An open-source release of the EGG framework implementation and the evaluation dataset at https://github.com/aalto-intelligent-robotics/EGG.

Main Findings

  • EGG outperforms both baselines across all modalities. Averaged over 5 trials, EGG scored 0.70 on the text-semantic score (S^text), 0.78 F1^binary, 0.76 F1^node, and 0.78 time accuracy (A^time). Hydra scored 0.52, 0.00, 0.68, and 0.10 respectively. ReMEmbR scored 0.54, 0.49, and 0.38, and was not evaluated on the node modality because it has no access to a scene graph.

  • Purely spatial representations fail on event-related queries. Hydra, implemented by removing event nodes and event edges from EGG, obtained 0.00 F1^binary and 0.10 A^time; the authors report that the LLM tends to rely only on object positions, which is ineffective for queries requiring human-interaction detail (e.g., describing the most frequently used mug for coffee-making at the coffee machine).

  • Caption-only memory systems fail on instance grounding and usage history. ReMEmbR scored lower on node queries (0.38) and struggled with questions such as "How many different mugs were seen in the events?" and questions requiring the usage history of a specific object, because its events are not grounded to spatial features.

  • Graph pruning improves both accuracy and token efficiency. The pruned EGG reached 0.83 overall accuracy (A^all) with 998 tokens and 64.72% graph compression, versus 0.77 A^all, 1576 tokens, and 0.00% compression for EGG without pruning. The paper reports roughly 65% graph compression and a 36% reduction in token usage from pruning.

  • Event edges add value, but a reduced version remains usable. Removing event edges after pruning (EGG w.o edges) gave 0.79 A^all, 0.81 S^text, 0.72 F1^binary, 0.82 F1^node, 0.80 S^time, 922 tokens, and 68.02% compression — close to the full model, though the paper notes it misses details for understanding specific interactions.

  • Caption quality directly limits representation quality. With automatic video captioning (VideoRefer) instead of ground-truth event captions, EGG scored 0.76 A^all, 0.70 S^text, 0.78 F1^binary, 0.76 F1^node, and 0.78 S^time, with 985 tokens and 66.55% compression. The paper states automatic captioning reduces overall accuracy by 9% compared to ground-truth event labeling.

  • Most errors come from the LLM, not the graph. In the failure-mode analysis with pruning and ground-truth captions, the primary error sources were semantic context (the LLM misinterpreting the question or event context) and hallucinated nodes and edges (incorrect connections between nodes or misidentification of objects and events related to the query). The authors identified only 1 instance where pruning removed a vital node related to the query.

Methodology in Plain English

The authors define a scene as a set of tracked spatial elements (objects, rooms), each with attributes split into time-invariant ones (name, semantic class, caption, and for rooms a manually labeled name and position) and time-variant ones (position, state of operation). An event is defined as a sequence of interactions within a time interval that change the attributes of at least one tracked element. Each event is stored as a node defined spatially by the mean of the robot's camera positions during the observation interval and by its time interval, plus a summary caption; edges connect that event to each involved object, with a per-object description of the role it played (for example, "washing a mug" linked to the specific mug).

To build EGG, the robot's RGB-D stream, frame timestamps, and pose are used. Object positions come from projecting segmented point clouds from depth images; for simplicity, only each object's initial and final position is recorded per event. GPT4o is used to generate concise appearance captions for objects to distinguish instances of the same class, and VideoRefer is used to generate event summary captions and per-object role descriptions.

For question answering, the full graph is serialized as JSON. A multi-stage strategy prompts an LLM agent to approximate the relevant information types: time interval, locations, objects, and events. In the first stage the agent selects a time interval and rooms, and the graph is pruned accordingly. In the second stage the agent is shown object names, captions, and event summaries from that reduced graph and picks the relevant objects and events; if both sets are non-empty, they are merged so that only objects connected to the relevant events remain (the paper's example: for "Where is the mug that I was drinking coffee with in the coffee room yesterday?", the merged set contains only mugs used for drinking coffee in the coffee room yesterday). Finally, the history of the relevant objects is expanded across the whole graph, the result is pruned back to the relevant time interval, re-serialized as JSON, and passed to the LLM agent to produce the answer.

Why This Matters

Impact on research: The paper formalizes the grounding of dynamic events into spatial scene elements, an extension of the grounding problem for 3D scene graphs, and provides an open-source implementation plus a purpose-collected evaluation dataset. It gives a concrete comparison showing that spatial-only and caption-only representations each fail on complementary query types, motivating unified representations.

Real-world applications (as discussed or supported by the paper):

  • Robots answering everyday human questions about the environment and its history, such as "Where is my phone?" by recalling where it was last used.
  • Semantic object search, where knowledge of past interactions informs more efficient search decisions.
  • Learning mobile manipulation skills from human demonstrations, using recorded interactions as data.
  • Assessing the states of routine tasks, tracking personal items by usage history, and offering insights about unusual occurrences.

Industry relevance: Service and assistive robotics, household robots, and warehouse or workplace robots that operate in environments shared with people would need memory systems that remember not only what is where but who did what to it. The framework relies on off-the-shelf models (GPT4o, VideoRefer) and standard robotic sensing, which lowers the barrier to adoption, though the current pipeline is offline rather than real-time.

Future Directions

  • More robust semantic search over the graph. The retrieved subgraph is not guaranteed to be optimal; the heuristic can include irrelevant objects connected to pertinent events, which matters more in larger, more complex environments with many more events.
  • Hierarchical structure in the time domain. EGG currently has no connectivity between recorded events, forcing the LLM to infer implicit links from timestamps or ordering; summaries over multiple events are suggested as an improvement.
  • Online, real-time construction. Building EGG is currently an offline process, which the authors state is not yet suitable for real-time robotic applications.
  • Inferring unseen interactions and transitions, and improving video captioning models, since automatic captioning measurably lowered downstream retrieval accuracy and a method is still required to automatically detect human-activity moments and select the objects involved (done manually in these experiments).

Target Audience

Robotics researchers and engineers working on scene representation, semantic mapping, and robot memory; researchers applying LLMs and VLMs to embodied and spatio-temporal reasoning; and practitioners building assistive or service robots that must recall and explain past interactions with a shared environment.

Authors’ abstract

A fundamental aspect for building intelligent autonomous robots that can assist humans in their daily lives is the construction of rich environmental representations. While advances in semantic scene representations have enriched robotic scene understanding, current approaches lack a connection between spatial features and dynamic events; e.g., connecting the blue mug to the event washing a mug. In this work, we introduce the event-grounding graph (EGG), a framework grounding event interactions to spatial features of a scene. This representation allows robots to perceive, reason, and respond to complex spatio-temporal queries. Experiments using real robotic data demonstrate EGG's capability to retrieve relevant information and respond accurately to human inquiries concerning the environment and events within. Furthermore, the EGG framework's source code and evaluation dataset are released as open-source at: https://github.com/aalto-intelligent-robotics/EGG.

Read the original paper