Skip to content
AI.info

Research

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions

Overview Research area: Robotics and embodied AI, specifically humanoid "loco-manipulation" — the combination of walking/motion control and object manipulation — using hierarchical world models, Visio

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions
arXiv
2601.17507
Published
2026-01-24
Authors
Yutong Shen, Hangxu Liu, Kailin Pei, Ruizhe Xia, Tongtong Feng

AI summary

Overview

Research area: Robotics and embodied AI, specifically humanoid "loco-manipulation" — the combination of walking/motion control and object manipulation — using hierarchical world models, Vision-Language Model (VLM) task planning, and transfer of pre-trained expert skills.

Technical level: Advanced. The paper assumes familiarity with model-based reinforcement learning (TD-MPC2, DreamerV3), latent dynamics models, Model Predictive Control (MPC), temporal-difference learning, contraction-mapping convergence arguments, and the "symbol grounding" debate in robotics.

Scope: The paper introduces MetaWorld, a three-layer framework that lets a VLM turn natural-language instructions into weighted combinations of pre-validated expert policies, which a TD-MPC2-based latent dynamics controller then executes, and it evaluates this on HumanoidBench tasks (arXiv:2601.17507v1 [cs.RO], 24 Jan 2026).

What This Paper Is About

Humanoid robots struggle with an "abstraction gap": large VLMs can understand what a task means ("open the door") but produce plans that violate a robot's kinematics and dynamics, while low-level controllers know how to move but cannot generalize to composed, long-horizon instructions. MetaWorld tries to close this gap by keeping the VLM as a pure semantic interface — it never emits joint actions, only weights over a library of physically feasible expert skills — and letting a hierarchical world model handle dynamic selection, fusion, and physical execution. The goal is efficient online adaptation without the sample inefficiency of end-to-end RL or the brittleness of pure imitation learning.

Key Contributions

  1. A hierarchical, modular world model architecture that decomposes task representation across a semantic planning layer and a physical execution layer, formalized as π(a_t | s_t, 𝓣) = π_phys(a_t | s_t, π_sem(𝓣)), allowing independent optimization of semantic understanding and physical control.

  2. A dynamic expert selection and motion prior fusion mechanism that reuses a pre-trained multi-expert policy library, selecting and weighting experts online via a state-conditioned probability distribution p(i | s_t), fused with VLM semantic weights through a coefficient α.

  3. A VLM-as-semantic-interface paradigm that converts the symbol grounding problem into a linear combination of expert policies: the VLM outputs an expert weight vector w = f_VLM(𝓣, 𝓔), which is normalized with a softmax-style expression, so that the resulting plan π_sem(𝓣) = Σ_i w_i π_exp^i is physically feasible by construction.

  4. A theoretical sample-complexity claim, showing that under contraction-mapping conditions the sample complexity is reduced from 𝒪(|𝓢||𝓐|/[(1−γ)²ε²]) to 𝒪(1/[(1−γ)²ε²]+K) compared with traditional methods.

Main Findings

  • Average return improvement of 135.6%: Across the four reported tasks (Stand, Walk, Run, Door), MetaWorld reaches an average return of 966.1 versus 410.0 for TD-MPC2 and 398.0 for DreamerV3. The improvement percentages in the paper's table are computed relative to TD-MPC2.

  • Stand: MetaWorld reports a return of 793.4 ± 13.5, against 699.3 ± 62.7 for both DreamerV3 and TD-MPC2 (a 5.8% improvement). Convergence ("Conv. (M)") drops to 0.8 versus 1.9 for TD-MPC2 and 5.5 for DreamerV3, a 57.9% improvement.

  • Walk: MetaWorld reports 701.2 ± 7.6 versus 644.2 ± 162.3 for TD-MPC2 and 428.2 ± 14.5 for DreamerV3 (8.8% improvement); convergence 1.4 versus 1.8 and 6.0 (22.2% improvement).

  • Run: The largest reported gain — 1689.9 ± 13.6 versus 66.1 ± 4.7 for TD-MPC2 and 298.5 ± 84.5 for DreamerV3, an improvement of 2456.3%; convergence 1.9 versus 2.0 and 6.0 (5.0% improvement). The paper attributes this to high-quality motion priors from the imitation-learned expert library supplying the VLM planner with a physically feasible action space.

  • Door (manipulation): 680.0 ± 50.0 versus 179.8 ± 52.9 for TD-MPC2 and 165.8 ± 50.2 for DreamerV3, a 278.3% improvement; convergence 1.7 versus 2.0 and 9.0 (15.0% improvement). The paper credits VLM decomposition of the "open door" instruction into sub-actions plus dynamic expert composition.

  • Convergence speed: The average convergence figure is 1.5 (M) for MetaWorld versus 1.9 for TD-MPC2 and 6.6 for DreamerV3, described as a 24.9% improvement.

  • Ablation — removing the VLM: Causes a 72.7% performance collapse on the Door task, which the authors interpret as a semantic planning failure — the robot cannot decompose "open the door" into skills like "approach handle – rotate – push/pull."

  • Ablation — removing dynamic expert selection (fixing α = 1.0): Only a 15.4% performance drop, which the authors read as robustness: without online adaptation the pre-trained expert base still supplies basic motion priors.

  • Ablation — removing expert guidance (λ = 0): A 52.9% performance loss, which the authors attribute to the synergy between imitation learning and model-based RL; expert policies supply feasible action bounds for online fine-tuning.

  • Baseline note: Other skills in the base expert policy repository (the paper names reach as an example) are inherited from TD-MPC2 and therefore excluded from evaluation.

  • Task-set inconsistency in the reported content: Section 4 states that the locomotion tasks selected are walk, stand, and reach, while the results table and Figure 3 caption report Stand, Walk, Run plus Door. The reported numbers correspond to Run, not Reach.

Methodology in Plain English

MetaWorld splits robot control into three tiers.

  1. Semantic layer (VLM). A Vision-Language Model reads the task description and outputs a weight vector over a library of pre-trained expert policies, rather than raw actions. Prompts are engineered so the response can be parsed and normalized into weights. Because every expert is physically feasible, any mixture of them is also feasible — this is how the framework sidesteps the classic symbol grounding problem.

  2. Skill transfer layer (dynamic selection and fusion). The framework also computes a state-aware selection distribution over experts, based on an encoding of the current state and a feature representation of each expert. The VLM's semantic weights and these state-aware probabilities are blended using α ∈ [0,1]; the paper sets α = 0.7, giving semantic planning more weight while still allowing short-term adaptation. The blended weights produce a reference expert action, a_ref, which serves as a high-quality initial guess for the controller.

  3. Physical layer (latent control). Execution uses TD-MPC2: observations are encoded into latent states, a dynamics model predicts how those latents evolve, and an MPC problem is solved over a short future horizon (H = 3) to maximize predicted reward plus a terminal value estimate. Training loss combines the temporal-difference loss with a term penalizing deviation of the chosen action from the expert reference action, weighted by λ = 0.05.

Other implementation details: VLM temperature τ = 0.3, discount factor γ = 0.99, learning rate 0.001, batch size 256. Locomotion tasks are learned via imitation learning from the AMASS expert dataset with trajectory-tracking rewards. Baselines are TD-MPC2 and DreamerV3, and success criteria follow HumanoidBench definitions.

Why This Matters

Impact on research: The paper argues that complex tasks can be achieved through semantic parsing and expert composition rather than end-to-end training, offering a way to use VLMs in robotics without inheriting their physical infeasibility. It positions hierarchical separation of semantic and physical layers as a route around both the sample-efficiency bottleneck of RL and the robustness limits of imitation learning.

Real-world applications:

  • Humanoid robots in unstructured human environments, such as opening doors or handling objects while walking.
  • General-purpose service or assistive robots that must accept natural-language instructions and convert them into safe, physically valid motion.
  • Industrial or warehouse settings where a limited library of validated skills must be recombined on the fly for varied task sequences.
  • Sim-to-real deployment pipelines where reusing pre-validated expert policies reduces the risk of infeasible generated plans.

Industry relevance: The framework's design — a frozen library of vetted low-level skills plus a language front-end — maps directly onto how robotics companies typically ship systems, where safety-critical controllers are validated separately from high-level planning. The reported convergence improvements (1.5 M versus 1.9 M and 6.6 M) and large reported task gains on Run and Door suggest potential reductions in training and fine-tuning cost.

Future Directions

  1. Dynamic reward shaping for imitation learning. The paper states the framework currently depends on simple trajectory-matching rewards that cannot dynamically weigh fine-grained aspects such as joint-level motion errors.

  2. Mixture-of-Experts (MoE) semantic-aware routing. Expert selection currently uses static weighted fusion rather than an intelligent routing mechanism, which the authors say limits precise skill composition and semantically conditioned switching.

  3. Few-shot transfer generalization. The system is described as lacking few-shot generalization, with constrained adaptability and scalability for novel complex task combinations.

  4. Gradient interference in multi-skill coordination. The paper notes this problem is not explicitly addressed, leaving open how multiple skills should be trained together without interfering with one another.

Target Audience

Researchers and engineers working on embodied AI, humanoid control, and robot learning — particularly those interested in world models, model-based RL, and integrating VLMs with low-level control. It is also relevant to practitioners who need a modular, validation-friendly architecture for combining language-level planning with pre-tested motor skills, and to students with a background in deep RL who want a concrete example of hierarchical skill transfer on a standard benchmark (HumanoidBench).

Authors’ abstract

Humanoid robot loco-manipulation remains constrained by the semantic-physical gap. Current methods face three limitations: Low sample efficiency in reinforcement learning, poor generalization in imitation learning, and physical inconsistency in VLMs. We propose MetaWorld, a hierarchical world model that integrates semantic planning and physical control via expert policy transfer. The framework decouples tasks into a VLM-driven semantic layer and a latent dynamics model operating in a compact state space. Our dynamic expert selection and motion prior fusion mechanism leverages a pre-trained multi-expert policy library as transferable knowledge, enabling efficient online adaptation via a two-stage framework. VLMs serve as semantic interfaces, mapping instructions to executable skills and bypassing symbol grounding. Experiments on Humanoid-Bench show MetaWorld outperforms world model-based RL in task completion and motion coherence. Our code will be found at https://anonymous.4open.science/r/metaworld-2BF4/

Read the original paper