Skip to content
AI.info

Research

Structured Imitation Learning of Interactive Policies through Inverse Games

Overview Research area: Robotics, imitation learning, multi-agent interaction, and game theory. Technical level: Advanced. Scope: The paper introduces a structured imitation learning framework that co

Structured Imitation Learning of Interactive Policies through Inverse Games
arXiv
2511.12848
Published
2025-11-17
Authors
Max M. Sun, Todd Murphey

AI summary

Overview

Research area: Robotics, imitation learning, multi-agent interaction, and game theory.
Technical level: Advanced.
Scope: The paper introduces a structured imitation learning framework that combines single-agent generative policy learning with an inverse game to learn interactive policies, demonstrated on a synthetic 5-agent social navigation task.

What This Paper Is About

Imitation learning of interactive policies is difficult because robots must coordinate with humans in shared spaces without explicit communication, and multi-agent behavior is more complex than single-agent behavior. The problem is that each agent’s actions influence all others, and demonstrations reflect both individual intents, such as reaching a goal, and collective intents, such as avoiding collisions. The goal is to learn interactive policies from multi-agent demonstrations by separating individual behavior learning from inter-agent dependency learning.

Key Contributions

  1. A structured two-stage imitation learning framework: first learn non-interactive policies from multi-agent demonstrations using standard single-agent imitation learning, then learn inter-agent dependencies by solving an inverse game problem.
  2. A game-theoretic formulation in which interactive policies are the Nash equilibrium of a game whose joint cost function is learned as a neural network, preserving the expressiveness of generative behavioral models.
  3. Preliminary results in a synthetic 5-agent social navigation benchmark showing that the interactive policy significantly improves the non-interactive policy and performs comparably to the ground-truth iLQGames policy using only 50 demonstrations.
  4. An implementation that uses a conditional variational autoencoder (CVAE) for the non-interactive policy, a multi-layer perceptron (MLP) for the joint cost function, JAX and Flax for implementation, and iLQGames as the ground-truth simulator.

Main Findings

  • Improved safety without efficiency loss: The proposed interactive policy significantly outperforms the corresponding non-interactive policy, improving the safety performance measured by minimum distance without compromising navigation efficiency measured by runtime cost.
  • Comparable to ground truth: The proposed interactive policy performs comparably to the ground-truth iLQGames policy and closely imitates its behavior from only 50 demonstrations.
  • Data efficiency through structure: Leveraging explicit structure for modeling multi-agent interaction significantly improves data efficiency in the 5-agent social navigation benchmark.
  • Evaluation protocol: The experiment uses 100 randomized trials with 5 agents, one of them being the robot during tests. The non-robot iLQGames agents operate under the assumption that the robot is an iLQGames agent with a presumed runtime cost function.
  • Quantitative and qualitative evidence: Qualitative results from a representative navigation trial are shown in Fig. 3, and quantitative results including median, quartiles, and distribution of metrics are shown in Fig. 4.

Methodology in Plain English

The authors collect multi-agent demonstrations and first treat each agent’s state-action sequence as a standard single-agent imitation learning problem. They learn a non-interactive policy for each agent using a generative model, implemented as a CVAE, so the policy captures individual behavior without considering other agents.

Next, they model inter-agent dependencies as a game. Each agent minimizes an individual objective that combines a collective intent, such as avoiding collisions, and an individual intent, represented by a KL-divergence term that keeps the agent close to its non-interactive policy. The authors assume the expert interactive policies form a Nash equilibrium of this game.

Because the joint cost function of the game is unknown, they model it as an MLP and solve an inverse game problem. The Nash equilibrium calculation is differentiable, so the joint cost function can be learned by maximizing the likelihood of the demonstrated actions through backpropagation. This is the multi-agent equivalent of inverse optimal control or inverse reinforcement learning.

The ground-truth data is generated with the iLQGames dynamic game solver. Each agent is modeled as a circular disk under Dubins car dynamics, and the individual runtime cost includes a navigation goal, a preferred longitudinal velocity, and a straight-line reference trajectory. The method is implemented in JAX and Flax.

Why This Matters

This work addresses an open challenge in imitation learning: learning interactive policies that coordinate with other agents from limited multi-agent demonstrations. It suggests that combining generative single-agent imitation learning with an explicit game-theoretic structure can reduce the amount of data needed while improving safety in interactive tasks. The framework is compatible with any generative model-based single-agent imitation learning method because the game-theoretic optimization problem can be solved using arbitrary non-interactive policies.

Real-world applications include:

  • Social navigation and collision avoidance in shared human spaces.
  • Autonomous driving and multi-agent traffic coordination.
  • Cooperative manipulation where multiple agents interact with the same object.
  • Robotic characters expressing emotional behaviors during interaction.

Industry relevance spans robotics, autonomous systems, human-robot interaction, and collaborative robots. The approach is especially relevant for safety-critical settings where robots must anticipate and coordinate with humans without explicit communication and where only limited demonstrations are available.

Future Directions

  • Integrate the method with a wider range of imitation learning methods.
  • Expand the task beyond navigation to domains such as cooperative manipulation.
  • Investigate the game formula further to improve the computational efficiency of the inverse game process, enabling rapid learning and adaptation of interactive policies with online observations.
  • Validate the approach on real-world human demonstrations and more complex multi-agent scenarios, since the reported results are preliminary and use a synthetic 5-agent social navigation benchmark.

Target Audience

This paper benefits robotics and AI researchers working on imitation learning, multi-agent systems, game theory, human-robot interaction, and autonomous navigation. It is also relevant to practitioners building interactive robots that must coordinate with humans in shared spaces. The content is advanced because it relies on game-theoretic concepts such as Nash equilibrium, inverse games, maximum likelihood estimation, and generative models like CVAEs.

Authors’ abstract

Generative model-based imitation learning methods have recently achieved strong results in learning high-complexity motor skills from human demonstrations. However, imitation learning of interactive policies that coordinate with humans in shared spaces without explicit communication remains challenging, due to the significantly higher behavioral complexity in multi-agent interactions compared to non-interactive tasks. In this work, we introduce a structured imitation learning framework for interactive policies by combining generative single-agent policy learning with a flexible yet expressive game-theoretic structure. Our method explicitly separates learning into two steps: first, we learn individual behavioral patterns from multi-agent demonstrations using standard imitation learning; then, we structurally learn inter-agent dependencies by solving an inverse game problem. Preliminary results in a synthetic 5-agent social navigation task show that our method significantly improves non-interactive policies and performs comparably to the ground truth interactive policy using only 50 demonstrations. These results highlight the potential of structured imitation learning in interactive settings.

Read the original paper