Skip to content
AI.info

Research

ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos

Overview Research area: Computer vision and structured sequence prediction, specifically procedural planning in instructional videos (predicting the sequence of actions that takes a viewer from a star

ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
arXiv
2603.04265
Published
2026-03-04
Authors
Luigi Seminara, Davide Moltisanti, Antonino Furnari

AI summary

Overview

Research area: Computer vision and structured sequence prediction, specifically procedural planning in instructional videos (predicting the sequence of actions that takes a viewer from a start visual state to a goal visual state).

Technical level: Intermediate. The paper assumes familiarity with hidden-state sequence models, Viterbi decoding, and standard procedure-planning benchmarks, but it explains its probabilistic formulation step by step.

Scope: The paper introduces ViterbiPlanNet, a lightweight planner that embeds a Procedural Knowledge Graph (PKG) directly into training through a Differentiable Viterbi Layer (DVL), and it re-benchmarks prior methods under a unified, statistically rigorous evaluation protocol.

What This Paper Is About

Given only a start frame and a goal frame from an instructional video, procedure planning asks for the ordered list of actions that connects them. Current state-of-the-art planners (diffusion models, LLM-based planners, transformer architectures) learn procedural structure implicitly and therefore need large models and large amounts of data. The authors instead encode procedural knowledge explicitly as a graph of actions and transition probabilities, and make the classical Viterbi decoding algorithm differentiable so gradients can flow from the plan back into a small visual network trained end to end.

Key Contributions

  1. ViterbiPlanNet and the Differentiable Viterbi Layer (DVL). A framework that integrates a Procedural Knowledge Graph end to end by replacing the non-differentiable max and argmax operations of Viterbi decoding with smooth relaxations, so the model learns emission probabilities from visual input rather than memorizing procedural rules. The DVL adds no trainable parameters; it is parametrized by fixed transition probabilities estimated from action co-occurrence statistics.

  2. A standardized, open-sourced evaluation benchmark. A unified pipeline with consistent data splits and metric implementations, with each model trained using five random seeds. Results are reported as means with 90% confidence intervals computed by bootstrapping, and differences are marked statistically significant only when the confidence interval excludes zero.

  3. A cross-horizon testing protocol. Models trained at a longer planning horizon are evaluated at shorter horizons to test consistency and robustness.

  4. Demonstration of parameter and sample efficiency. The planner operates with roughly 5–7M parameters and is shown to require fewer training sequences than a comparable transformer-based baseline.

Main Findings

  • Structure-aware training drives the gains, not post-hoc decoding. On CrossTask with T=3, configurations 6–8 (DVL used during training) reach about 6% absolute improvement in Success Rate over the base model (configuration 1), whose SR is 32.47 ± 0.32. The full model (configuration 8, DVL at training and standard Viterbi Decoding at inference) reaches SR 38.45 ± 0.32, mAcc 63.07 ± 0.17, mIoU 83.89 ± 0.16, an improvement of 5.98 ± 0.47 SR, 2.44 ± 0.29 mAcc and 1.44 ± 0.16 mIoU over configuration 1.

  • Adding decoding at inference to a baseline barely helps. Configurations 2 (32.99 ± 0.28 SR), 3 (32.09 ± 0.26) and 4 (30.77 ± 0.19) show that bolting standard Viterbi Decoding, the DVL, or both onto a base model at inference time alone does not bring substantial improvement.

  • Emissions alone are not usable as plans. Using DVL-trained emissions directly for prediction (configuration 5) drops to SR 20.05 ± 0.63, because emissions are distributions over states rather than actions; the same emissions perform well once decoded (configurations 6 and 7).

  • DVL is backward-compatible with classical Viterbi decoding. Replacing DVL with standard VD at inference does not change performance substantially (configuration 6: SR 38.09 ± 0.39 versus configuration 7: 37.66 ± 0.45), and stacking VD on top of the DVL soft plan gives comparable results (configuration 8: 38.45 ± 0.32).

  • All compared methods benefit from the PKG, but ViterbiPlanNet benefits most. On CrossTask with T=3, enabling the PKG raises KEPP by 2.56 ± 5.93 SR, PlanLLM by 1.97 ± 1.81, SCHEMA by 3.31 ± 0.84, and ViterbiPlanNet by 5.98 ± 0.47.

  • Best Success Rate across all reported settings. On CrossTask T=3, ViterbiPlanNet reaches SR 38.45 ± 0.32 against SCHEMA's 37.24 ± 0.60 (+1.21 ± 0.69); on COIN T=3, SR 33.99 ± 0.23 (+0.55 ± 0.27 over the best competitor); on NIV T=3, SR 32.37 ± 0.96 (+2.37 ± 1.63). At T=4 the corresponding numbers are CrossTask 24.64 ± 0.30 (+0.46 ± 0.61), COIN 23.92 ± 0.29 (+0.73 ± 0.44), NIV 27.54 ± 0.70 (+3.15 ± 1.93). On CrossTask and COIN the mAcc and mIoU differences are small or statistically inconclusive, while NIV shows positive and statistically significant improvements on all metrics.

  • LLM and VLM in-context baselines underperform trained planners. Vision-only prompting with Qwen2.5-VL-32B is particularly weak (SR 11.48 on CrossTask T=3), symbolic reasoning helps (Qwen2.5-32B: 25.14; Qwen3-30B: 23.37), adding the PKG to the prompt does not improve results (Qwen3-30B + PKG: 23.31), and even Gemini 2.5 Pro (29.18 on CrossTask T=3) trails training-based methods. The simple PKG beam-search baseline (22.38 ± 0.26) outperforms most LLM/VLM baselines.

  • Competitive with a much larger diffusion model. Against MTID on CrossTask, ViterbiPlanNet attains mIoU 76.92 at T=3 versus MTID's 69.17, and 80.67 versus 67.67 at T=4, with comparable SR (39.75 vs 40.45 at T=3; 24.19 vs 24.76 at T=4) and comparable mAcc (67.39 vs 67.19; 61.12 vs 60.69), despite MTID having 1,085.20M parameters against ViterbiPlanNet's 5.49M and 5.52M.

  • Strong parameter efficiency. The method runs with roughly 5–7M parameters (5.57M on CrossTask T=3, 6.67M on COIN T=3, 5.48M on NIV T=3), which the abstract describes as an order of magnitude fewer than diffusion- and LLM-based planners; the paper body describes this as two to three orders of magnitude fewer than language models (about 30B–100B), MTID (1.08B) and PlanLLM (about 385M).

  • Better sample efficiency. With a frozen visual encoder and the same PKG, ViterbiPlanNet outperforms SCHEMA (a similar-parameter model, roughly 6M) when trained on progressively larger fractions of the training data, and the gap narrows as more data becomes available, which the authors attribute to memorization.

  • Robust to shorter unseen horizons. Under the cross-horizon protocol (training at horizon 6, testing shorter), ViterbiPlanNet reaches SR 27.77 ± 0.43 at [6→3] and 18.45 ± 0.39 at [6→4], against SCHEMA's 16.12 ± 1.24 and 9.69 ± 0.90 and Gemini 2.5 Pro's 20.97 and 10.46. The value for the [6→5] setting of ViterbiPlanNet is not present in the provided content (the table is truncated), so it cannot be reported here.

Methodology in Plain English

The problem is framed probabilistically. Latent actions generate visual states, and each action depends only on the previous action (the Markov property). The probability of a plan given the start and goal states factorizes into two parts: transition probabilities, which say how likely one action is to follow another, and emission probabilities, which say how compatible an action is with the visual evidence.

Transitions come from a Procedural Knowledge Graph built once per dataset from co-occurrence statistics in the training data, so they are pre-computed and fixed. Emissions are the only thing the neural network has to learn: a visual encoder extracts features from the start and goal clips using a frozen backbone plus a learnable projection, and a transformer encoder with an MLP and sigmoid output predicts an emission matrix over actions and time steps.

The critical piece is decoding. Standard Viterbi decoding uses max and argmax, which block gradients. The authors use log-sum-exp and softmax relaxations to define differentiable versions, "S-max" and "S-argmax". State scores are updated with the smooth maximum over predecessor scores, and a soft backpointer distribution is computed in parallel. During the backward pass, these soft backpointers are recursively composed into a soft plan, a time-indexed distribution over actions that smoothly approximates the discrete Viterbi solution. Because the layer holds no trainable parameters, the only thing optimized through it is the emission network.

Training uses three equally weighted losses: a planning loss (mean squared error between the soft plan and the one-hot ground-truth plan), a visual-semantic alignment loss, and a task classification loss. At inference, standard Viterbi decoding is applied to the soft plan unless stated otherwise.

For benchmarking, the authors re-trained and re-evaluated prior methods using official implementations, ran each model with five random seeds, and reported 90% confidence intervals via bootstrapping. They note they did not re-train LLM-based models or MTID because those exceed their computational budget.

Why This Matters

Impact on research. The paper argues that the field's reliance on implicit procedural learning is a fundamental bottleneck, and it shows that a small, structure-aware model can beat much larger ones. It also addresses a long-noted problem: prior work reported inconsistencies in training and testing protocols, metric implementations, feature extraction schemes, data loaders and parameter counts. The open-sourced unified benchmark, multi-seed runs and confidence intervals provide a fairer basis for comparison and could change how results in this area are reported.

Real-world applications:

  • Wearable AI assistants that guide a user through a daily activity step by step from a start and goal observation.
  • Assistive and instructional systems that generate checklists or step orderings for cooking, repair and assembly tasks.
  • Robotics and embodied agents that need to sequence actions under valid preconditions rather than produce impossible orderings.
  • Augmented-reality training tools that verify whether a sequence of steps is procedurally valid.

Industry relevance. The method's parameter count (roughly 5–7M) and its sample efficiency lower the cost of training and deployment relative to diffusion or LLM-based planners with hundreds of millions to tens of billions of parameters. The explicit graph also makes the plan structure inspectable and constraints enforceable, which suits production systems where invalid action sequences are unacceptable.

Future Directions

  • Extending the graph beyond fixed co-occurrence statistics. The DVL uses a fixed transition matrix estimated from training data; learning transition parameters jointly with emissions, as in the differentiable dynamic programming work the authors build on, is a natural extension. The supplementary material reportedly examines dependence on PKG quality and results using a single PKG across datasets.
  • Scaling the cross-horizon protocol. The [6→5] result is not included in the provided content; completing these tests and probing longer horizons (results for T ∈ {5,6} are said to be in the supplementary material) would clarify how far the robustness extends.
  • Incorporating intermediate visual observations. The main model uses only start and goal states; the supplementary material reports experiments using intermediate observations, and integrating them into the main framework is an open path.
  • Improving LLM integration with structured knowledge. Adding the PKG to prompts did not improve Qwen3-30B results, leaving open how to make foundation models actually exploit explicit procedural structure.

Target Audience

Researchers and graduate students working on procedure planning, instructional video understanding, and structured prediction with differentiable dynamic programming. It is also relevant to practitioners who need efficient, verifiable planners rather than large generative models, and to anyone benchmarking sequence-planning methods who would benefit from the unified evaluation protocol and multi-seed statistical reporting.

Authors’ abstract

Procedural planning aims to predict a sequence of actions that transforms an initial visual state into a desired goal, a fundamental ability for intelligent agents operating in complex environments. Existing approaches typically rely on large-scale models that learn procedural structures implicitly, resulting in limited sample-efficiency and high computational cost. In this work we introduce ViterbiPlanNet, a principled framework that explicitly integrates procedural knowledge into the learning process through a Differentiable Viterbi Layer (DVL). The DVL embeds a Procedural Knowledge Graph (PKG) directly with the Viterbi decoding algorithm, replacing non-differentiable operations with smooth relaxations that enable end-to-end optimization. This design allows the model to learn through graph-based decoding. Experiments on CrossTask, COIN, and NIV demonstrate that ViterbiPlanNet achieves state-of-the-art performance with an order of magnitude fewer parameters than diffusion- and LLM-based planners. Extensive ablations show that performance gains arise from our differentiable structure-aware training rather than post-hoc refinement, resulting in improved sample efficiency and robustness to shorter unseen horizons. We also address testing inconsistencies establishing a unified testing protocol with consistent splits and evaluation metrics. With this new protocol, we run experiments multiple times and report results using bootstrapping to assess statistical significance.

Read the original paper