Research
Procedural Mistake Detection via Action Effect Modeling
Overview Research area: Computer vision, egocentric procedural video understanding, and mistake detection in task-oriented video. Technical level: Advanced. The paper assumes familiarity with one-clas

- arXiv
- 2512.03474
- Published
- 2025-12-03
- Authors
- Wenliang Guo, Yujiang Pu, Yu Kong
AI summary
Overview
- Research area: Computer vision, egocentric procedural video understanding, and mistake detection in task-oriented video.
- Technical level: Advanced. The paper assumes familiarity with one-class classification, contrastive learning, vision-language models, scene graphs, and action segmentation backbones.
- Scope: The paper proposes Action Effect Modeling (AEM), a framework that detects mistakes in procedural tasks by modeling not only how an action is performed but also what outcome it produces, evaluated on the EgoPER and CaptainCook4D benchmarks under the one-class classification setting.
What This Paper Is About
Most existing mistake detection systems judge whether an action was performed correctly by looking only at the execution process (motion patterns, action sequences, prototypes, or task graphs). The authors argue that many errors are invisible in the motion itself and only become apparent in the result, such as a mixture spilled onto a table or a cucumber cut into an unexpected shape. The goal is to detect both execution errors and outcome errors by explicitly modeling the action effect alongside the action execution.
Key Contributions
- A probabilistic reformulation of mistake detection. The authors cast mistake detection as a marginalization over latent action effects, decomposing the task into three subproblems: effect frame sampling, effect-aware learning, and mistake classification.
- Action Effect Modeling (AEM). A module that learns effect-aware action representations by capturing object states and spatial relationships from complementary visual (grounded object features) and symbolic (scene graph) cues, aligned in a shared latent space.
- A prompt-based detector. Instead of a binary or frame-level classifier, the detector aligns each action segment with a task-specific textual prompt contrastively, capturing temporal execution dynamics in context.
- State-of-the-art results under the OCC setting. The framework reports the best performance on EgoPER and CaptainCook4D, with ablations isolating the contribution of each component.
Main Findings
- EgoPER benchmark results: The method achieves an average AUC of 73.8 and EDA of 66.7 across the five recipe tasks (Quesadilla 80.8 AUC / 68.1 EDA, Oatmeal 77.0 / 68.6, Pinwheel 69.9 / 61.2, Coffee 70.3 / 66.4, Tea 71.1 / 69.4). The paper reports this as an average improvement of 5.3% on AUC and 2.3% on EDA over prior methods.
- CaptainCook4D benchmark results: The method reaches 68.1 Precision, 62.5 AUC, and 71.9 EDA, surpassing AMNAR by 2.7% in Precision and 2.3% in AUC. The authors note that EDA is occasionally lower than AMNAR on both datasets, which they attribute to AMNAR's use of a dynamic programming block that constructs task graphs to mitigate noise in action-segment labels.
- Effect modeling matters: Removing AEM entirely gives 67.6 AUC and 65.6 EDA on EgoPER. Adding all multimodal effect supervision (visual and textual, for both state and relation) raises this to 73.8 AUC and 66.7 EDA.
- Relation effects beat state effects: Explicit supervision for spatial relations (69.4 AUC / 66.3 EDA) outperforms state supervision (68.4 / 66.1), a 1.0% AUC improvement that the authors attribute to spatial relationships being more consistent and discriminative than subtle object state changes.
- Visual supervision beats textual supervision: Using visual features yields 71.7 AUC and 66.4 EDA, while textual features yield a smaller gain (69.9 / 66.0). The authors suggest scene-graph-derived text can be abstracted and noisy for fine-grained egocentric cues.
- Cross-modal alignment is necessary: Without aligning visual and textual supervision, performance drops to 66.8 AUC and 64.7 EDA, lower than the model with no effect modeling at all. Aligning in the relation space (72.6 AUC) gives a stronger gain than the state space (69.9 AUC), and aligning both reaches 73.8 AUC / 66.7 EDA.
- Informed frame selection helps: Selecting the last frame of each segment raises AUC to 70.6 (EDA 65.7), while the proposed semantic-relevance-plus-visual-clarity sampling reaches 73.8 AUC / 66.7 EDA, versus 67.6 / 65.6 with no effect frame.
- Dynamic fusion contributes: Adding the multi-scale dynamic fusion module yields a 2% improvement in AUC and a 3.3% gain in EDA. Without it, the model still surpasses AMNAR by 3.3% AUC.
- Action segmentation also improves: Using ActionFormer as a shared backbone, the method achieves 58.5 IoU, 69.7 Edit, 58.5 F1@0.5, and 73.5 Acc, versus AMNAR at 56.3, 69.4, 57.3, and 75.3, and EgoPED at 44.6, 61.3, 47.5, and 68.5.
- Open-source scene graphs are competitive: Replacing GPT-4o with Qwen3-VL (30B) for scene graph generation gives 73.3 AUC and 66.6 EDA overall, comparable to GPT-4o's 73.8 and 66.7.
- Qualitative coverage of both error types: Visualizations show the model detecting mistakes that appear in final outcomes, and also execution errors during the process even when the visual outcome looks correct.
Methodology in Plain English
The system processes a procedural video that has already been split into non-overlapping action segments, each labeled with a start time, end time, action label, and a binary mistake label. Training uses only correct actions (the one-class classification setting), while testing includes both correct and erroneous segments.
Three things happen in sequence. First, the model picks an "effect frame" from within each action segment that best shows the result of the action, scoring candidates on how well they match a GPT-4o-generated description of the expected post-action state and on image sharpness (measured with a Laplacian operator). Second, from that frame it extracts two kinds of knowledge. A visual branch uses Grounding DINO to find objects and encodes their appearance and positions. A textual branch asks GPT-4o to build a scene graph with object, relation, and attribute nodes, which is encoded by a text encoder and processed by a graph neural network, then split into a state subgraph and a relation subgraph. Third, a learnable "effect token" attached to the segment features is trained to match these visual and textual signals through mean-squared-error losses and contrastive losses, distilling outcome knowledge into a compact representation. Crucially, the expensive external models are only used during data pre-processing and training; at inference the model relies only on the learned token.
For the detection stage, the enriched segment features are pooled into an action embedding and compared contrastively with a textual prompt built from a template such as "An image showing [ACTION] for [TASK]" plus a learnable prefix. At inference, the mistake probability is one minus the similarity between the action embedding and its prompt, thresholded to produce a label. All components, including the segmentation backbone and the dynamic fusion module that aggregates multi-scale temporal features, are trained end-to-end with a combined objective.
Why This Matters
The paper shifts mistake detection away from a purely execution-centric view toward a joint view of execution and outcome, which the authors argue is necessary because correct-looking motions can produce flawed results. This reframing could influence how future procedural video models, anomaly detection systems, and task-assistance tools are designed.
Real-world applications suggested or implied by the work:
- Cooking assistance: Detecting spilled mixtures or irregularly cut ingredients in egocentric kitchen video, as illustrated with the CaptainCook4D examples.
- Assembly and manufacturing: Verifying that a step produced the intended physical result, not just that the motion resembled a reference.
- Medical procedures: Assessing whether a procedural step produced the correct outcome, where subtle execution deviations can have serious consequences.
- Life assistive systems and skill learning: Providing feedback to users learning a task, which the authors cite as motivation for building intelligent support systems.
Industry relevance: The design choice to run effect-frame sampling and multimodal knowledge extraction only during pre-processing means inference relies on a lightweight learned token, which matters for deployment on edge or AR/VR hardware. The finding that Qwen3-VL (30B) can replace the closed-source GPT-4o for scene graph generation also lowers the cost barrier for industrial pipelines.
Future Directions
- Spatial-temporal effect modeling: The authors propose extending the framework beyond a single effect frame to long-range procedural reasoning.
- Interpretability: Generating human-understandable explanations of detected mistakes using large language models.
- Integrating task graphs: The authors note that AMNAR's dynamic programming block for task graph construction lies outside their scope and leave incorporating it for future exploration, which may close the remaining EDA gap.
- Broader video understanding tasks: The segmentation results suggest action effects could serve as useful cues for action recognition and action analysis more generally.
Target Audience
This paper is best suited to researchers and graduate students working on video understanding, egocentric vision, procedural learning, and anomaly or mistake detection, as well as practitioners building task-assistance or quality-verification systems who need to know the state of the art under the one-class classification setting. Readers without a background in contrastive learning and multimodal supervision will find the method section dense, though the core intuition about execution versus outcome is accessible.
Authors’ abstract
Mistake detection in procedural tasks is essential for building intelligent systems that support learning and task execution. Existing approaches primarily analyze how an action is performed, while overlooking what it produces, i.e., the \textbf{action effect}. Yet many errors manifest not in the execution itself but in the resulting outcome, such as an unintended object state or incorrect spatial arrangement. To address this gap, we propose Action Effect Modeling (AEM), a unified framework that jointly captures action execution and its outcomes through a probabilistic formulation. AEM first identifies the outcome of an action by selecting the most informative effect frame based on semantic relevance and visual quality. It then extracts complementary cues from visual grounding and symbolic scene graphs, aligning them in a shared latent space to form robust effect-aware representations. To detect mistakes, we further design a prompt-based detector that incorporates task-specific prompts and aligns each action segment with its intended execution semantics. Our approach achieves state-of-the-art performance on the EgoPER and CaptainCook4D benchmarks under the challenging one-class classification (OCC) setting. These results demonstrate that modeling both execution and outcome yields more reliable mistake detection, and highlight the potential of effect-aware representations to benefit a broader range of downstream applications.