Research
On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning
On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Authors: Sacha Morin (Université de Montréal; Mila – Quebec AI Institute), Moonsub Byeon (Samsung Electronics

- arXiv
- 2602.02762
- Published
- 2026-02-02
- Authors
- Sacha Morin, Moonsub Byeon, Alexia Jolicoeur-Martineau, Sébastien Lachapelle
AI summary
On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation LearningAuthors: Sacha Morin (Université de Montréal; Mila – Quebec AI Institute), Moonsub Byeon (Samsung Electronics, Suwon), Alexia Jolicoeur-Martineau (Samsung AI Lab, Montreal), Sébastien Lachapelle (Samsung AI Lab, Montreal) arXiv: 2602.02762v2 [cs.LG], 02 Jul 2026 | Keywords: Machine Learning, ICML | License: CC BY 4.0
Overview
- Research area: Semi-supervised imitation learning (SSIL), inverse dynamics models (IDMs), offline policy learning from action-free video, and latent action learning.
- Technical level: Intermediate, with Advanced sections. The unification and consistency arguments rely on KL-divergence and MDP notation; the complexity argument draws on statistical learning theory (VC dimension, Rademacher complexity) and neural network simplicity bias.
- Scope: One sentence — the paper asks why inverse-dynamics-based methods learn policies more sample-efficiently than behavior cloning, unifies two popular IDM-based methods under one optimal policy, and proposes an improved latent-action algorithm (LAPO+) evaluated on Procgen, Push-T and LIBERO.
What This Paper Is About
Behavior cloning (BC) requires action-labeled expert demonstrations, which are expensive to collect. SSIL instead combines a small action-labeled dataset with a much larger dataset of action-free videos, typically by training an inverse dynamics model that predicts what action produced a transition from state s to next state s'. Prior work observed empirically that IDM-based methods generalize better than BC from the same number of labeled samples, but only offered partial explanations. This paper formalizes that advantage, argues it comes from the ground-truth IDM being both less complex and less stochastic than the expert policy, and uses those insights to build a better algorithm.
Key Contributions
- A unification result. The authors show that, at optimality, VM-IDM and IDM labeling recover the same policy when the unlabeled dataset is infinite and model capacity is sufficient. They name this common policy the IDM-based policy (Section 3).
- Two claimed causes of IDM sample efficiency. IDM learning is argued to be more sample-efficient than BC when (i) the ground-truth IDM lies in a lower-complexity hypothesis class relative to the expert policy, and/or (ii) the ground-truth IDM is less stochastic than the expert policy. These claims are supported with statistical learning theory and novel experiments (Section 4), including a study of IDM-based policies using recent architectures for unified video-action prediction (UVA).
- An extensive empirical comparison. A novel comparison across the 16 Procgen environments, Push-T and LIBERO, discussing how IDM properties correlate with the improved performance of IDM-based policies over BC (Section 5).
- An improved algorithm. An improved version of the LAPO algorithm for latent action policy learning, called LAPO+, demonstrated superior on the Procgen benchmark (Section 5.1). The authors further show that sampling a recent UVA architecture as a VM-IDM can improve policy success (Section 5.2).
Main Findings
- VM-IDM and IDM labeling are the same policy in the limit. With infinitely many unlabeled pairs
(s, s'), expressive enough hypothesis classes, and global optimization, both methods learnπ̂_{v*,ĥ}, the IDM-based policy. The VM-IDM policy is written asπ̂_{v̂,ĥ}(a|s) = ∫ ĥ(a|s,s') v̂(s'|s) ds', and the IDM-labels objective is shown to reduce to a cross-entropy toward that same distribution. - The IDM-based policy is consistent. With infinitely many action-labeled samples and sufficient capacity, IDM learning recovers the ground-truth IDM
h*(a|s,s') := p_{π*}(a|s,s'). Sinceĥ(a|s,s') v*(s'|s) = p_{π*}(a,s'|s), integrating overs'givesp_{π*}(a|s) = π*(a|s), soπ̂_{v*,h*} = π*. BC, VM-IDM and IDM labeling are all consistent estimators of the expert. - A formal inequality links IDM error to policy error. The paper proves (Appendix B) that
E_{p_{π*}(s)} D_KL(π* || π̂_{v*,ĥ}) ≤ E_{p_{π*}(s,s')} D_KL(h* || ĥ). Combined with the claim that IDM learning has lower KL error than BC at the same labeled dataset sizeN_L, this suggests IDM-based policies should outperform BC in SSIL. - In mazes, the ground-truth IDM is less complex than the expert policy. On mazelab-generated mazes (sizes 10x10, 20x20, 50x50), low-capacity VM-IDM reaches perfect test accuracy given enough samples (indicating
h*is in the low-capacity class), while low-capacity BC never reaches perfect accuracy (indicatingπ*is not). The authors give an explicit form: because the agent never runs into a wall, the action is recoverable froms' - s, soh*(a|s,s') = softmax_a(V(s,s'))for someV ∈ ℝ^{4×4}— that is, a linear classifier. For image states, they showh*can be expressed by a single-layer convolutional network. - The IDM advantage grows with environment complexity. In the low-data regime, high-capacity VM-IDM outperforms high-capacity BC, and this gap increases for more complex mazes. The trend is also present in image states but less strongly than for position states.
- Capacity control explains part of the effect. Because
h*is simple, a smaller architecture reduces variance without adding bias, which is the classical bias-variance trade-off formalized by VC dimension and Rademacher complexity bounds. - Simplicity bias explains the rest. Even with the same 5-layer network, high-capacity VM-IDM beats high-capacity BC despite the IDM having twice as many inputs (a larger hypothesis class). The authors attribute this to the implicit simplicity bias of neural networks in the overparametrized interpolation regime.
- Goal diversity affects the expert more than the IDM. In a 10x10 maze where the goal location changes per trajectory, BC without goal conditioning cannot reach perfect test accuracy even with all possible transitions, while VM-IDM without goal conditioning does reach perfect accuracy — showing the goal
gis unnecessary for predicting the action from(s, s'). BC with goal conditioning eventually reaches perfect accuracy but is outperformed by VM-IDM in the low-data regime. - Stochasticity of the expert. Claim 2 states the ground-truth IDM is often less stochastic than the expert policy, contributing to IDM sample efficiency. The provided text shows only the caption of Figure 3 (state visitation distributions for different experts, comparing average reward of BC and IDM labeling, averaged over 10 seeds); the specific numerical results of that experiment are in the truncated portion and are not reported here.
- A stochastic maze variant. In an appendix setting where the agent remains static with some probability regardless of the action, the ground-truth IDM is not linear, but the authors report reaching similar conclusions.
- LAPO+ outperforms LAPO on Procgen. The paper's proposed algorithm is presented in Table 1 as an alternative to the three-stage LAPO pipeline; the exact performance numbers are in the truncated Section 5.1 and are not reported here.
Methodology in Plain English
The authors mix theory and controlled experiments.
- Theory. They set up a finite-horizon MDP, define state-visitation and transition-visitation distributions, and write the learning objectives of BC, VM learning, IDM learning, IDM labeling and LAPO in a common notation. They then use standard KL-divergence arguments to show what each objective converges to when data is infinite and capacity is sufficient. This produces the equivalence result, the consistency result, and the inequality bounding policy error by IDM error.
- Controlled maze experiments. To test whether the ground-truth IDM is genuinely simpler than the expert, they use mazelab to generate mazes and a deterministic expert to solve them. They then vary four things: maze size (10x10, 20x20, 50x50), the fraction of states present in the labeled dataset, the state representation (agent
(x, y)position versus maze images), and model capacity (linear classifier and 5-layer MLP for positions; 1-layer CNN and 5-layer CNN for images). In this section VM-IDM uses the ground-truth video modelv*, echoing the infinite-unlabeled-data regime. Results are averaged over 5 seeds. - Goal-diversity experiments. They keep the 10x10 maze but let the goal location change between trajectories, observing the goal
gwith each transition. They compare BC with and without goal conditioning, and VM-IDM with a goal-conditioned ground-truth VM paired with an IDM that is either goal-conditioned or not. Averaged over 5 seeds. - Stochastic-environment and expert-stochasticity experiments. These probe the second claimed cause of sample efficiency, with results averaged over 10 seeds.
- Benchmark evaluation. The theoretical framework is then tested on the 16 Procgen environments, Push-T and LIBERO, and used to motivate two practical changes: a modified LAPO (LAPO+), where the latent IDM decoding step replaces the latent policy decoding step and IDM labeling follows, and sampling a UVA architecture as the VM in a VM-IDM policy.
Why This Matters
The paper provides a predictive framework rather than a single empirical win: it tells practitioners when to expect IDM-based methods to beat BC. Because the IDM drops the goal/return information that the policy must encode, and because its induced distribution is often less noisy, it can be learned reliably from far fewer labeled transitions. That matters for any domain where labeled actions are the bottleneck and raw video is abundant.
Applications:
- Robotics manipulation. Pretrained vision models, object tracking, optical flow and inverse kinematics can supply an IDM, while large-scale robot or human video supplies the VM — the paper cites robot manipulation and generated robot videos as prior successes.
- Autonomous driving. IDM labeling variants have been applied here, where huge volumes of driving video exist but precisely labeled control actions are scarcer.
- Computer-use agents. The paper notes applications to computer-use agents, where screen recordings are plentiful and ground-truth action labels are not.
- Game-playing agents from gameplay video. VPT on Minecraft is the canonical example, using contractor data for
D_Land unlabeled online gameplay forD_U.
Industry relevance: The work is co-authored by researchers at Samsung AI Lab (Montreal) and Samsung Electronics (Suwon), reflecting direct industrial interest in learning policies from cheap video rather than expensive labeled demonstrations. The LAPO+ and UVA results point to concrete algorithmic and architectural choices for teams building video-driven agents. Code for all experiments is stated to be publicly available.
Future Directions
- Quantifying the two causes per environment. The paper offers complexity and stochasticity as an explanatory framework, but it does not provide a direct, measurable diagnostic that predicts the BC-versus-IDM gap on a new benchmark; developing such a metric is a natural next step.
- Extending beyond analysis to more environments. The framework is tested on the 16 Procgen environments, Push-T and LIBERO; whether the same correlation between IDM properties and performance holds in higher-dimensional, real-robot video domains is left open.
- Ground-truth stochastic IDMs. The linear form for
h*breaks in the stochastic maze variant, and no simple analytic form is expected in most realistic settings. Understanding how far the low-complexity argument extends in those regimes is unresolved. - Better UVA-based VM-IDMs. Section 5.2 shows that sampling a unified video-action prediction architecture as a VM-IDM can improve policy success; how to design and train such architectures to exploit the IDM advantage most effectively remains an active question.
Target Audience
Researchers and engineers working on imitation learning, offline reinforcement learning, learning from observations, and video-based robot learning. It is especially useful for those who already use IDM labeling or VM-IDM pipelines and want to understand the conditions under which those pipelines should beat behavior cloning. The paper is also relevant to theorists interested in applying statistical learning theory and neural network simplicity bias to practical policy-learning questions, and to practitioners at robotics, autonomous driving and agent-development companies deciding how to allocate annotation budget between action labels and raw video.
Authors’ abstract
Semi-supervised imitation learning (SSIL) consists in learning a policy from a small dataset of action-labeled trajectories and a much larger dataset of action-free trajectories. Some SSIL methods learn an inverse dynamics model (IDM) to predict the action from the current state and the next state. An IDM can act as a policy when paired with a video model (VM-IDM) or as a label generator to perform behavior cloning on action-free data (IDM labeling). In this work, we first show that VM-IDM and IDM labeling learn the same policy in a limit case, which we call the IDM-based policy. We then argue that the previously observed advantage of IDM-based policies over behavior cloning is due to the superior sample efficiency of IDM learning, which we attribute to two causes: (i) the ground-truth IDM tends to be contained in a lower complexity hypothesis class relative to the expert policy, and (ii) the ground-truth IDM is often less stochastic than the expert policy. We argue these claims based on insights from statistical learning theory and novel experiments, including a study of IDM-based policies using recent architectures for unified video-action prediction (UVA). Motivated by these insights, we finally propose an improved version of the existing LAPO algorithm for latent action policy learning. We experiment on the Procgen, Push-T and LIBERO benchmarks.