Skip to content
AI.info

Research

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

Overview Research area: Computer vision, specifically egocentric and exocentric human action anticipation, multimodal fusion, and vision-language modeling for instructional/industrial activity underst

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
arXiv
2512.02846
Published
2025-12-02
Authors
Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

AI summary

Overview

  • Research area: Computer vision, specifically egocentric and exocentric human action anticipation, multimodal fusion, and vision-language modeling for instructional/industrial activity understanding.
  • Technical level: Intermediate. The method is built from familiar components (DINOv2 features, cross-attention, a text encoder), but prior background in action anticipation benchmarks and attention mechanisms helps.
  • Scope: The paper asks whether a single frame enriched with depth and a text-based action history can replace temporal video aggregation for predicting actions one second ahead, tested on IKEA-ASM, Meccano, and Assembly101.

What This Paper Is About

Most action anticipation systems process many video frames and aggregate them over time to guess what a person will do next. This paper asks whether that temporal machinery is necessary: given a single frame plus the right extra information, can a model predict the upcoming action just as well? The authors build a method called AAG (Action Anticipation at a Glimpse) that combines one RGB frame, a depth map, and a textual record of previous actions, then test how close this "glimpse" approach gets to full video-based methods across three assembly datasets.

Key Contributions

  1. A single-frame anticipation study: The authors systematically investigate single-frame action anticipation as a competitive alternative to video-based approaches, using recent progress in multimodal learning, vision models, and language models.
  2. The AAG method: AAG fuses RGB with depth cues from one frame through cross-attention, and fuses the result with a text-encoded action history through self-attention. Code is released on GitHub (https://github.com/ManuBenavent/AAG).
  3. Three strategies for encoding action history: (a) prompting a vision-language model (VLM) with a single frame to describe past actions, (b) generating descriptions from action labels produced by a single-frame action recognizer, and (c) encoding each past action class separately and concatenating the embeddings.
  4. Benchmarking against video-based state of the art: Comparisons with AVT, RULSTM, TempAgg, and VLMAH on IKEA-ASM, Meccano, and Assembly101, including a computational cost and parameter comparison.

Main Findings

  • Single-frame anticipation is competitive, not universally better. Under realistic (predicted) action history, AAG reaches 44.66/82.87 Top-1/Top-5 accuracy on IKEA-ASM, 26.43/54.95 on Meccano, and 13.56/32.13 with 8.90 class-mean Recall@5 on Assembly101. With ground-truth history it reaches 61.83/89.64, 31.86/66.80, and 26.82/53.05 with 24.86 recall respectively.
  • Action history is the single largest contributor. On IKEA-ASM, adding a ground-truth action history to RGB alone lifts Top-1 accuracy from 34.37 to 60.14, and an RGB+Depth+action-history single-frame model reaches 61.83. Ground-truth action history alone (63.55 Top-1 on IKEA-ASM) outperforms combinations with visual input in some settings, indicating that the history itself carries highly informative temporal cues.
  • Depth helps mainly in third-person views. Adding depth improves IKEA-ASM by +4.45% (34.37 to 38.82 Top-1 with no action history), an exocentric setting where depth captures object localization. In Meccano (egocentric) and Assembly101 (top-down/close-up) depth is less informative or misleading, because those views capture near-planar workspaces with limited depth variation.
  • Video still wins on complex, variable tasks. On Assembly101, video-based methods lead: TempAgg reaches 11.43/34.16 Top-1/Top-5 and VLMAH reaches 9.17/27.63 with 25.13 class-mean Recall@5, versus AAG's 13.56/32.13 and 8.90. The authors attribute this to Assembly101's multiple valid execution orders and timing variation, which require explicit temporal modeling.
  • AAG beats RGB-only video baselines on some benchmarks. AAG's RGB-only single-frame setup scores 34.37/83.43 on IKEA-ASM versus AVT at 27.10/69.70, TempAgg at 26.90/70.14, and RULSTM at 26.37/70.05. On Meccano, AAG RGB reaches 27.21/50.02 against AVT's 27.43/53.38 and RULSTM's 24.08/58.23.
  • AAG is far cheaper at inference. AAG processes 1 frame with 202M total and 24M trainable parameters, compared with AVT (404M/392M, 10 frames), TempAgg (135M/123M, 37 frames), RULSTM (79M/67M, 14 frames), and VLMAH (52M/45M, 8 frames). VLMAH additionally depends on external action recognizers, while AAG uses self-supervised, task-agnostic feature extractors.
  • The optimal history length depends on the dataset. IKEA-ASM performs best with 7 or 3 past actions, while Meccano and Assembly101 perform best with 3. Longer histories (5, 7, 10 actions) introduce noise on the more complex datasets; for example Assembly101 drops from 26.82/53.05 at 3 actions to 24.35/48.97 at 7.
  • VLMs are weaker at inferring past actions from one frame. On IKEA-ASM, Llama-3.2 Vision reaches 38.66/86.47 with no context, 37.54/86.48 with dataset context, and 38.66/83.31 with action context; GPT-4o with action context reaches 38.66/85.23 and DeepSeek-VL2 reaches 38.18/82.99. Label-based histories are far better: 61.83/89.64 (ground-truth, concatenation) and 44.66/82.87 (predicted, concatenation).
  • Concatenating per-action embeddings is the best history encoding. Ground-truth histories scored 57.18/90.28 with a single description, 61.83/89.64 with concatenation, and 52.18/87.43 with a transformer aggregator. Concatenation preserves each action's distinct semantics while capturing order.
  • Predicted history quality gates performance. Meccano performs worse than its visual-only baseline under realistic history, which the authors attribute to the lower accuracy of the action recognizer in that domain. The single-frame action recognizer reaches 66.43% on IKEA-ASM, 31.21% on Meccano, and 34.19% on Assembly101 without depth, and 66.84%, 31.14%, and 32.77% with depth.
  • Depth frames are estimated, not always native. Because depth is missing or noisy in some datasets, Depth Anything V2 generates depth, which is color-mapped into an RGB representation so standard RGB feature extractors can be used.

Methodology in Plain English

The authors start by taking one frame captured a set amount of time before the action to be predicted (the paper evaluates δ = 1 second into the future). They feed that frame to DINOv2, a self-supervised vision transformer, to get a compact image representation — no fine-tuning required.

To add spatial understanding, they also generate a depth map of the same frame with Depth Anything V2 and recolor it into an RGB-style image so the same kind of extractor can encode it. The RGB and depth embeddings are then combined with a cross-attention transformer, letting each modality sharpen the other.

The second ingredient is context about what has already happened. The authors try three ways to get it: ask a vision-language model to guess past actions from the single frame; take the action labels predicted by a single-frame action recognizer and turn them into text; or simply encode each past action class on its own and concatenate the resulting embeddings. The text is encoded with DistilBERT, and an extra transformer fuses the visual and textual representations before a linear classifier and softmax produce the prediction. Training uses cross-entropy loss.

For comparison, the same architecture can be swapped to accept aggregated video frames via a temporal transformer with 3 layers and 8 attention heads, so the authors can isolate how much temporal aggregation actually contributes under identical conditions. They train with AdamW (weight decay 0.01, learning rate 5e-5) for up to 100 epochs with early stopping patience of 10 and an improvement threshold of 0.001, batch size 32, on Nvidia RTX 4090 GPUs. The transformer encoders use 2 layers, 4 attention heads, and an embedding dimension of 768.

Why This Matters

The work tests a widely assumed requirement: that accurate action anticipation demands long video sequences and heavy temporal aggregation. If a single frame plus depth and action history gets competitive results, then anticipation systems can be dramatically cheaper, simpler, and easier to deploy in settings where processing every frame is impractical.

Real-world applications named in the paper:

  • Autonomous driving: predicting the behavior of distracted pedestrians or erratic drivers to improve safety.
  • Industrial settings: forecasting upcoming actions to help prevent defects and enhance workplace safety.
  • Human-robot collaboration: anticipating human actions before they occur to support fluency and success in collaborative scenarios.
  • Instructional assembly assistance: the datasets studied are furniture and toy assembly tasks (IKEA-ASM, Meccano, Assembly101), the kind of step-by-step guidance a single-frame system could support with minimal camera infrastructure.

Industry relevance: the efficiency profile is the strongest argument. AAG processes one frame with 24M trainable parameters, versus 10 frames and 392M trainable parameters for AVT, and it pre-encodes action classes so no text encoder is needed at inference. The paper also notes that reusing the same features for both recognition and anticipation yields competitive action recognizers without meaningfully raising cost.

Future Directions

  • Integrating richer dataset-specific visual cues. The authors suggest that fine-grained cues related to hands, objects, and their interactions would help close the gap in close-up, egocentric settings where depth is unhelpful.
  • Adding a memory mechanism for past actions. The conclusion names a memory mechanism as a likely requirement for further improving single-frame anticipation.
  • Improving the underlying action recognizer. The paper shows predicted-history quality governs results, especially on Meccano, and explores replacing the single-frame recognizer with more advanced video-based architectures in the appendix, leaving the tradeoff between history quality and cost as an open question.
  • Characterizing when temporal modeling is genuinely necessary. The authors frame dataset structure — procedural predictability and variability — as the deciding factor, raising the question of how to predict which regime a new domain belongs to before committing to an architecture.

Target Audience

Researchers and practitioners in video action understanding who want to know how much temporal modeling is actually required, and engineers designing efficient anticipation pipelines for industrial, robotics, or assisted-assembly settings where a single camera stream and limited compute are the norm. It also suits readers interested in how far vision-language models can substitute for observed history, since the paper provides a direct comparison of three VLMs against label-based history on the same benchmark.

Authors’ abstract

Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single moment from a scene, when given sufficient context. Can a model achieve this competence? The short answer is yes, although its effectiveness depends on the complexity of the task. In this work, we investigate to what extent video aggregation can be replaced with alternative modalities. To this end, based on recent advances in visual feature extraction and language-based reasoning, we introduce AAG, a method for Action Anticipation at a Glimpse. AAG combines RGB features with depth cues from a single frame for enhanced spatial reasoning, and incorporates prior action information to provide long-term context. This context is obtained either through textual summaries from Vision-Language Models, or from predictions generated by a single-frame action recognizer. Our results demonstrate that multimodal single-frame action anticipation using AAG can perform competitively compared to both temporally aggregated video baselines and state-of-the-art methods across three instructional activity datasets: IKEA-ASM, Meccano, and Assembly101.

Read the original paper