Skip to content
AI.info

Research

Agent Priors-guided Policy Learning

Agent Priors-guided Policy Learning Overview Research area: Robot learning from demonstration, specifically imitation learning with diffusion policies, compositional skill generalization, and LLM-agen

Agent Priors-guided Policy Learning
arXiv
2609.35690
Published
2026-09-28
Authors
Puming Jiang, Tianrun Hu, Haozhe Du, Yibo Li, Zhiwei Xue, Xinhu Li, Harold Soh

AI summary

Agent Priors-guided Policy Learning

Overview

Research area: Robot learning from demonstration, specifically imitation learning with diffusion policies, compositional skill generalization, and LLM-agent-driven design of robot learning pipelines.

Technical level: Advanced. The paper assumes familiarity with imitation learning, diffusion policies, task-and-motion planning abstractions, and language-model agents.

Scope: The paper proposes APPL, a system in which a language-model construction agent designs multiple structural priors per skill, trains and verifies one diffusion policy per prior, and exposes those priors as the interface a runtime agent uses to select and compose frozen policies.

What This Paper Is About

Robots that learn from a few demonstrations must generalize in two coupled ways: each individual skill must still work when objects move or when it starts from a state another skill left behind (skill generalization), and learned skills must be recombined into new task sequences (compositional generalization). The paper argues that information is lost between these two levels, because the task level usually sees a skill only through a name, instruction, or symbolic operator that says nothing about the physical conditions under which the learned policy actually works. The goal is to make the structural assumption used to train a policy also serve as the description that tells a runtime agent when and how to reuse it.

Key Contributions

  1. A dual role for structural priors. The paper formalizes a structural prior as having a training realization that shapes where a policy generalizes (coverage) and a description realization that forms part of the runtime interface (selection). The stated intent is improved interface faithfulness (A_i ⊆ W(π_i)) and informativeness (A_i ≈ W(π_i)).

  2. The APPL construction pipeline. A construction agent segments complete, unsegmented demonstrations into reusable skills with overlapping handoffs, proposes several structural priors per skill from different families, implements and trains one conditional diffusion policy (Chi et al., 2023) per prior against a fixed policy interface, and verifies each policy before freezing the library.

  3. A four-field runtime interface and a runtime agent. Each frozen policy is documented as I_i = (d_i, h_i, s_i, v_i), covering the prior description, handoff conditions, observed training support, and verification evidence, plus typed call arguments. A separate runtime agent reads these interfaces each episode to pick a policy, set its arguments, duration, and stop conditions, and chain calls toward a new goal, without modifying the library.

  4. Empirical evaluation of both roles. Experiments on six MetaWorld tasks and five long-horizon ManiSkill tasks test prior design for skill generalization, and prior descriptions plus runtime selection for composition, including ablations that hide interface information while keeping the same trained policies.

Main Findings

  • Agent-designed priors raise out-of-distribution skill generalization on MetaWorld. On the six adapted MetaWorld tasks (pick-place-wall, assembly, drawer-open, door-open, peg-insert-side, stick-push; 20 demonstrations each), the best-of-four APPL candidate (q_4) reaches 89.58% OOD success at N = 2 and 93.13% at N = 20, versus 28.96% and 46.67% for the vanilla Diffusion Policy (B0) and 37.92% and 56.25% for a fixed relational prior (B1). Averaged over the four values of N, mean OOD is 92.40% for q_4, 85.21% for q_3, 67.92% for q_1, 47.29% for B1, and 37.86% for B0.

  • The margin over the fixed relational prior is large at every data budget. The reported q_4 margins over B1 are 51.67, 50.83, 41.04, and 36.88 percentage points for N = 2, 5, 10, and 20. IID success is similar across methods at larger N, which the authors read as gains coming from generalization rather than fitting.

  • Selection and feedback each add performance. The first proposal alone beats B1 by 18.1–22.5 percentage points; proposing three priors and selecting by development OOD success adds 10.2–24.0 points; the feedback round adds 2.9–12.1 more. A4 was selected in 15 of 24 conditions, with test success rising in 10, tying in 13, and dropping in one. The three reported variants share 144 trained models.

  • Design revisions are concrete and can be targeted. At drawer/N = 2, A1 expressed the hand–handle relation and actions in the cabinet frame (75.00%); because demonstrated planar actions are close to 4(p_handle − p_hand), A4 subtracts that term so the network predicts only a residual, reaching 100.00%. At assembly/N = 5, development feedback showed crossed layouts failing, so A4 attenuates goal features and the goal-aligned frame until the ring is lifted, rising from 41.25% to 93.75%. A revision can also hurt: door/N = 20 falls from 100.00% under q_3 to 96.25% under q_4, and the paper notes that q_4 − q_3 does not isolate the effect of feedback.

  • APPL composes better than full-task policies under shifted objects. On five long-horizon ManiSkill tasks with twelve demonstrations each, APPL succeeds on 50.0% of motion-OOD layouts, versus 10.0% for both full-task Diffusion Policy (DP) and SinglePrior, with paired McNemar p < 0.001 for both comparisons.

  • Task-level OOD and unseen compositions. APPL reaches 92.5% success on the eight-per-task task-level OOD cases (resuming near a demonstrated stage or stopping at a requested intermediate sub-goal) and solves 8 of 16 previously unseen compositions.

  • The same library fails without informed selection. Hiding every field of the interface while keeping the verified four-policy-per-skill library drops motion-OOD from 50.0% to 20.0%, task-level from 92.5% to 65.0%, and composition from 8/16 to 2/16. Replacing the runtime agent with a fixed rule that follows the demonstrated skill order gives 22.5%, 52.5%, and 1/16. Hiding only prior, handoff, and support descriptions while keeping verification scores and exit criteria (w/o prior information) yields 45.0%, 72.5%, and 3/16.

  • A language-model interface alone is weaker. Agent+VLA, which replaces the library with one vision-language-action policy (Black et al., 2025) called by the same runtime agent using a skill name as instruction, reaches 32.5% motion-OOD, 70.0% task-level, and 6/16 compositions.

  • Verification helps overall but can misrank under shift. Verification improves selection overall, yet scores from demonstrated entry states misranked policies in covered peg assembly: its report discouraged a policy useful under runtime handoff rules and promoted another, and success fell from 8/8 to 1/8, offsetting gains on the other four tasks.

  • Remaining failures are recovery-related. Failures commonly follow a dropped object, a drawer pushed closed, or a skill invoked outside its trained entry conditions; the demonstrations contain no recovery trajectories and the library has no skills for re-approaching lost objects.

Methodology in Plain English

The method splits work into an offline construction phase and an online execution phase, both driven by language models.

Construction (offline, library built once). The construction agent receives N complete, unsegmented demonstrations, their task goals, and observation/control conventions. It proposes skill boundaries on every trajectory and groups segments performing the same operation into skill datasets. Adjacent segments are deliberately overlapped around handoffs, so a successor policy's training data include the tail of its predecessor and the head of its successor; this reuses existing transitions and adds no new demonstrations. For each skill, the agent proposes several structural priors from different families rather than committing to one design. Each prior is implemented against a fixed policy interface: inputs are built from causal observation history, demonstrated actions are re-encoded in the prior's coordinates, predicted actions are decoded back to native commands, and an optional auxiliary loss may be added. For example, an object-frame prior transforms end-effector targets by the object's rotation and position at the start of an action chunk, encouraging object-relative rather than absolute motion. Each prior yields a separate conditional diffusion policy trained with a common recipe, and a shared inverse-kinematics module converts decoded end-effector targets into joint commands.

Verification. Each candidate policy is initialized from the entry state of its corresponding segment in every demonstration, executed for 1.5 times the segment length, and checked against the demonstrated exit condition at the final state. Only demonstrated skill-entry states are used. The agent receives these results, writes a verification report separating observed outcomes from inferences, and on that basis proposes one additional prior per skill, which is trained and verified the same way.

Interfaces. Every trained policy gets an interface of four fields: a prior description stating which observations and relations the policy depends on, which variations it should tolerate, and known limitations; a handoff description with entry and exit conditions and the predecessor/successor states represented by the overlapping segments; a support description summarizing demonstrated training support such as ranges of entry states; and verification evidence. Typed arguments such as the object to manipulate and its destination are also declared.

Execution (online, every episode, no weight updates). The library, interfaces, and the runtime agent's model, prompt, and tools are frozen. At each decision point the runtime agent receives the goal, current state and goal predicates, recent history, an episode notebook, and the available interfaces; it can inspect any policy's full interface before calling it. A call specifies typed arguments, an execution duration, and stop conditions over observed quantities that may be absolute or relative to the call start (for example, gripper width below 3 cm and the grasped object lifted more than 4 cm). An executor runs the policy on fresh observations and checks stop conditions after every step; it also interrupts immediately if the overall goal is satisfied, returning the updated state and goal predicates. The agent can then continue, switch to a different policy for the same skill, invoke a different skill, or terminate. It cannot write control code or modify the frozen library, and its memory resets between episodes.

Evaluation design. Experiment 1 isolates coverage: six MetaWorld tasks with 20 demonstrations each, test states split into IID (within demonstrated factor combinations), C (recombined factors), and E (both factors outside demonstrated intervals), with OOD weighting C and E equally. All policies get the same state channels, the same diffusion backbone, a 20,000-update budget, the same controller, and the same evaluation, but candidates may change representations, encoders, action coordinates, and auxiliary losses. In each of 24 task–N conditions the agent submits three candidates (A1–A3) before any performance feedback, then designs a fourth (A4) from their development results and trains it from scratch; all designs and selections were frozen before testing. Experiment 2 tests coverage and selection together: five long-horizon ManiSkill tasks with twelve scripted-planner demonstrations each, eight frozen motion-OOD layouts and eight task-level OOD cases per task, and 16 composition cases. Both experiments are simulation-only; no real-robot evaluation is reported.

Why This Matters

The paper's distinctive claim is that the inductive bias chosen during training is itself useful information at deployment time. Rather than discarding the design rationale when a policy is frozen, APPL keeps it as the interface over which an agent reasons, which turns "which policy structure should I use?" from a one-time design decision into a per-skill hypothesis that can be proposed, implemented, verified, and reused. For robot learning research, this offers a concrete mechanism for coupling the two generalization problems that are usually studied separately, and it provides a measurable framing (faithfulness and informativeness of an applicability region) for what an interface should provide.

Real-world applications:

  • Warehouse and home manipulation, where objects are frequently placed in new positions and a limited demonstration budget must cover many goals.
  • Multi-step assembly or kitting, where sub-operations are reordered or skipped and each step must start from whatever state the previous one left behind.
  • Household tasks that begin mid-way or stop at a requested intermediate stage, such as returning a container without first emptying it.
  • Deployments where robots are taught from a handful of expert trajectories but must adapt to new task variants without retraining.

Industry relevance: The construction pipeline is a way to amortize a small number of expert demonstrations into a documented, reusable library, and the verification step provides a lightweight sanity check before freezing policies. The explicit interface documentation also targets a practical integration problem: agentic systems that call policies as tools currently know a tool only by its name, which the paper identifies as a common source of transition failures.

Future Directions

  • Deployment-aligned verification. The paper calls for verification that reflects deployment states and handoff rules rather than demonstrated entry states, since scores from demonstrated starts can misrank policies under shift (as in covered peg assembly) and OOD trials on real robots are often unavailable.
  • Expanding skill support. The authors propose targeted data collection or augmentation to widen the states each skill covers, motivated by failures that occur when a skill is invoked outside its trained entry conditions.
  • Recovery and re-approach behavior. The current library has no skills for re-approaching lost objects and the demonstrations contain no recovery trajectories, so dropped objects and drawers pushed closed remain unresolved failure modes.
  • Integration with complementary tools. The paper suggests integrating motion planners alongside the learned policy library, and more broadly retaining design knowledge so downstream agents can reason about how learned behaviors can be reused, combined, and extended.

Target Audience

This paper is most valuable to robotics and robot-learning researchers working on imitation learning, few-demonstration policy generalization, and skill composition, particularly those using diffusion policies or task-and-motion-planning abstractions. It also speaks to researchers building language-model agents that orchestrate robot skills as tools, and to practitioners in industrial or service robotics who need a small set of demonstrations to cover many deployment conditions. Readers without a background in imitation learning or policy representations will find the problem framing accessible, but the method and results assume familiarity with diffusion policies and manipulation benchmark suites.

Authors’ abstract

Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.

Read the original paper