Research
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
Overview Research area: Offline imitation learning in machine learning, specifically the comparison between behavior cloning (BC) and predictive inverse dynamics models (PIDMs). Technical level: Advan

- arXiv
- 2601.21718
- Published
- 2026-01-29
- Authors
- Lukas Schäfer, Pallavi Choudhury, Abdelhak Lemkhenter, Chris Lovett, Somjit Nath, Luis França, Matheus Ribeiro Furtado de Mendonça, Alex Lamb, Riashat Islam, Siddhartha Sen, John Langford, Katja Hofmann, Sergio Valcarcel Macua
AI summary
Overview
Research area: Offline imitation learning in machine learning, specifically the comparison between behavior cloning (BC) and predictive inverse dynamics models (PIDMs).
Technical level: Advanced. The paper uses expected prediction error decompositions, covariate shift and density-ratio analysis, Fisher information, and biased Cramér-Rao lower bounds; the empirical sections are accessible but the theory is not.
Scope: A theoretical and empirical analysis of when and why conditioning an inverse dynamics model on a predicted future state yields lower prediction error and better sample efficiency than standard behavior cloning.
What This Paper Is About
Behavior cloning (BC) is the standard offline imitation learning method, but it often needs large numbers of expert demonstrations, which are costly to collect. PIDMs—architectures that combine a future-state predictor with an inverse dynamics model—have been reported to outperform BC, yet the paper states that the reasons behind these benefits remain unclear. This work derives conditions under which PIDMs achieve lower prediction error and higher sample efficiency than BC, then tests those predictions in 2D navigation and a 3D video game.
Key Contributions
-
A generalization-gap result for optimal predictors (Theorem 1). The paper shows that for optimal BC and IDM predictors, the predicted error gap equals the expected conditional variance of the optimal IDM predictor over future states, and is therefore always nonnegative:
Δ_{p*} = E_{s_t ~ d}[ Var_{s_{t+k} ~ p*}( f*_idm(s_t, s_{t+k}) ) ] ≥ 0. -
A decomposition of the gap for non-optimal estimators under covariate shift (Corollary 1). Training the IDM with future states from the true distribution
p*but testing with future states from an approximate predictorp̂induces a covariate shift. The realized gap isΔ̂_p̂ = Δ_{p*} + δ + β + γ, separating the variance-reduction benefit from differences in estimator variance (δ), bias (β), and irreducible noise weighted by the density ratiow(s_t, s_{t+k}) = p̂(s_{t+k}|s_t) / p*(s_{t+k}|s_t)(γ). -
Conditions for sample-efficiency gains (Theorem 2, Lemma 1, Corollary 2). Using Fisher information and the biased Cramér-Rao lower bound, the paper derives a ratio
η = n/mof samples needed by BC (n) versus IDM (m) to reach a target MSEε, shows the Fisher ratioF_{ξ*} ≥ F_{μ*}under a local shared-parameterization assumption, and concludesη ≳ 1when the intrinsic bias terms and their derivatives match for BC and IDM. -
Empirical validation in 2D and 3D. On four 2D navigation tasks with human demonstrations, BC requires up to five times more demonstrations than PIDM (three times on average) to reach comparable performance. In a 3D video game with high-dimensional visual inputs, stochastic transitions, and real-time inference, BC requires over 66% more samples than PIDM to reach an 80% success rate.
Main Findings
-
Conditioning on the future state cannot hurt the optimal predictor. Theorem 1 establishes that the optimal IDM predictor's expected prediction error is at most that of the optimal BC predictor, with the gap equal to the expected variance of the IDM predictor with respect to future states.
-
Prediction introduces a tradeoff. Because the state predictor is approximate, the IDM is tested on a different future-state distribution than it was trained on. This covariate shift adds bias and variance that can shrink, though not necessarily eliminate, the gain over BC.
-
Covariate shift can sometimes help. The paper notes that when
0 < w < 1, the approximate predictor places more mass on future states with lower irreducible variance (γ > 0), and it may reduce estimator bias (β > 0) or variance (δ > 0). -
Additional data sources widen the gap. Action-free demonstrations of the same task can train a more accurate state predictor, reducing the covariate-shift term so that
δandβapproach zero. Fork = 1, expert or even non-expert demonstrations from other tasks in the same environment can reduce the variance of the IDM estimator, makingδ > 0. The paper cites Xie et al. (2025) as having empirically reported gains from such additional data. -
Sample efficiency in 2D navigation. Reported efficiency ratios
η_PIDM: at 80% of maximum performance, 4.0 (Four room), 2.0 (Zigzag), 5.0 (Maze), 2.0 (Multiroom), average 3.25; at 90%, 4.0 / 2.0 / 4.0 / 4.0, average 3.5; at 95%, 4.0 / 1.33 / 5.0 / 2.0, average 3.08. Maximum reached goal ratio for BC was 0.98 / 0.97 / 0.95 / 0.99 and for PIDM 0.99 / 0.98 / 0.99 / 0.98. -
Less action diversity increases PIDM's advantage. Datasets collected with a deterministic A* planner, which have lower state-prediction error, further increased PIDM's advantage. The paper also reports that PIDM significantly outperforms a retrieval-based BC variant (Appendix E), separating the PIDM decomposition from the retrieval-based state predictor.
-
Where the gains come from. State-wise prediction error gaps are largest in states surrounding the goals, where human actions are more diverse. Visualizations of the learned IDM indicate that the policy only attends to future states in states with large
Δ(s). -
Correlation between state-prediction error and performance. The paper reports a correlation between state prediction error and rollout performance for PIDM across all datasets and tasks.
-
3D world result. In the "Tour" task, BC needs 66% more samples than PIDM to reach 80% success rate, even with a simple state predictor.
Methodology in Plain English
The authors first treat imitation learning as a regression problem. BC predicts an action from the current state. PIDM instead predicts a plausible future state from the current state, then predicts the action from the current state and that predicted future state. The paper writes the expert policy as an integral over future states of an expert IDM policy, then compares the expected prediction error of BC and IDM in two ways: at test time (generalization) and during training (sample efficiency).
For generalization, the analysis shows the optimal IDM's error is below the optimal BC's error by exactly the amount of action uncertainty explained by the future state. For non-optimal estimators, the analysis adds terms for the mismatch between the true future-state distribution and the approximate predictor's distribution, using a density ratio to quantify that covariate shift. For sample efficiency, the analysis uses Fisher information and the Cramér-Rao lower bound to relate the number of samples each method needs to hit a target mean squared error.
Empirically, the authors run four 2D navigation tasks in which an agent must reach a sequence of goals, using low-dimensional states (agent position, goal positions, whether each goal was reached) so that gains from representation learning are excluded. They collect 50 human demonstrations per task, with Gaussian noise N(0, 0.2) added to actions, and vary dataset size across (1, 2, 5, 10, 20, 30, 40, 50) demonstrations with 20 random seeds, four checkpoints (5000, 10,000, 50,000, 100,000 optimization steps) and 50 rollouts each. They also test datasets from a deterministic A* planner. For the 3D setting, they use third-person video frames from the Dojo practice level of the game Bleeding Edge, encode them with a pretrained ViT-B/16 Theia encoder, and use continuous actions in [−1,1]^4 at 30 FPS. In 2D, PIDM uses k = 1 and MLPs; in 3D, PIDM is trained with k ∈ {1, 6, 11, 16, 21, 26} and evaluated with k = 1. State predictors are deliberately simple: nearest-neighbor retrieval in 2D, and in 3D a prediction conditioned on time step and a randomly sampled training demonstration.
Why This Matters
Impact on research. The paper supplies a theoretical account of an empirical observation that prior PIDM work (Du et al. 2023; Tian et al. 2025; Tot et al. 2025; Xie et al. 2025; Pai et al. 2026) had reported but not explained. It isolates the benefit of the action decomposition from the representational benefits of multi-step inverse dynamics noted in earlier work (Lamb et al. 2023; Koul et al. 2024; Levine et al. 2024; Islam et al. 2023), and it complements analysis of BC as implicit trajectory modeling (Foster et al. 2024). It also gives a formal reason why additional data sources help: they shrink the covariate-shift terms in the error gap.
Real-world applications:
- Robotics, where collecting expert demonstrations is often costly, time-consuming, or infeasible (the paper cites Schaal 1999 and Fang et al. 2019).
- Autonomous driving, another domain the paper lists as a target for offline imitation learning (Pan et al. 2020).
- Gaming and character control, where the paper demonstrates a 3D video game task with real-time inference requirements (Pearce and Zhu 2022; Pearce et al. 2023; Schäfer et al. 2025).
- Any offline control setting where reward functions are unavailable and environment interaction is not permitted during training.
Industry relevance. The paper is authored primarily at Microsoft, with collaborators at McGill University/Mila and Tsinghua University, and code is released at https://github.com/microsoft/understanding_pidm_for_imitation_learning. The finding that a simple state predictor paired with an IDM yields large sample-efficiency gains—and that BC needed 66% more samples in a real-time, partially observable, visually rich deployment with random latency and visual artifacts—is directly relevant to teams that pay for demonstration collection or face real-time inference constraints.
Future Directions
- Characterizing when the covariate shift helps rather than hurts. Corollary 1 indicates
γ > 0,β > 0, orδ > 0are possible, but the paper does not give practical conditions for identifying these cases in advance. - Extending the analysis beyond the assumptions used. The theory covers scalar parameters, single-point predictors under mean-squared loss, and i.i.d. sampling; the paper defers vector parameters, other losses, and the non-i.i.d. case to appendices and leaves their practical consequences open.
- Better state predictors. Since the error gap depends on how closely
p̂matchesp*, improving the state predictor—for example with action-free demonstrations—is a direct lever the paper identifies for closing the gap. - Choosing the prediction horizon. PIDM is trained with
k ∈ {1, 6, 11, 16, 21, 26}in 3D and evaluated atk = 1; the paper points to an appendix discussion of the choice ofk, leaving principled horizon selection as an open question. - Testing whether the small-data gains generalize further. The theory's sample-efficiency conditions are derived for sufficiently large
nandm, while the paper's own evidence is that gains hold in the small-data regime—an extension the paper explicitly flags as empirical rather than proven.
Target Audience
Researchers and practitioners in offline imitation learning, robot learning, and game AI who need to decide between behavior cloning and predictive inverse dynamics architectures, and who want a principled explanation rather than a purely empirical comparison. The theory sections assume familiarity with expected prediction error, covariate shift, Fisher information, and the Cramér-Rao bound, so the paper is best suited to readers with a graduate-level machine learning background. Engineers focused only on deployment may find the 3D video game experiment, its sample-efficiency curves, and the simple state-predictor design the most immediately useful parts.
Authors’ abstract
Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent works have introduced a class of architectures named predictive inverse dynamics models (PIDMs) that combine a future-state predictor with an inverse dynamics model. While PIDMs often outperform BC, the reasons behind their benefits remain unclear. In this paper, we provide a theoretical explanation: PIDMs introduce a tradeoff. Conditioning the IDM on the predicted future state can significantly reduce variance, but the prediction itself introduces additional bias and variance. We establish conditions for PIDMs to achieve higher sample efficiency and lower prediction error than BC, with the gap widening when additional data sources are available. We validate the theoretical insights empirically in 2D navigation tasks, where BC requires up to five times (three times on average) more demonstrations than PIDM to reach comparable performance. Results are also illustrated in a complex 3D environment in a modern video game with high-dimensional visual inputs and stochastic transitions, where BC requires over 66\% more samples than PIDM.