Skip to content
AI.info

Research

World Action Learning via Interaction-Centric Spectral Latent Guidance

World Action Learning via Interaction-Centric Spectral Latent Guidance (WING) Overview Research area: Robotics — robot manipulation policy learning, vision-language-action (VLA) models, world-action m

World Action Learning via Interaction-Centric Spectral Latent Guidance
arXiv
2610.03607
Published
2026-10-02
Authors
Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin, Fangqi Zhu, Xiaoyi Pang, Quanxin Shou, Zhengyang Yan, Haodong Wang, Song Guo

AI summary

World Action Learning via Interaction-Centric Spectral Latent Guidance (WING)

Overview

Research area: Robotics — robot manipulation policy learning, vision-language-action (VLA) models, world-action models (WAMs), and latent action learning from egocentric human video.

Technical level: Advanced. The paper combines latent action modeling, teacher–student distillation, discrete cosine transform (DCT) spectral analysis, and VLA/WAM policy training. A reader would benefit from background in representation learning, imitation learning, and frequency-domain signal analysis.

Scope: The paper proposes a framework, WING, that extracts interaction-centric latent actions from large-scale egocentric human videos, isolates their slowly varying (low-frequency) temporal structure as guidance, and uses that guidance to condition robot action generation — evaluated on LIBERO, RoboTwin 2.0, RoboCasa–GR1, and four real-world bimanual tasks.

Published as arXiv:2610.03607v1 [cs.RO], 02 Oct 2026, under CC BY 4.0, from The Hong Kong University of Science and Technology, Hong Kong SAR, China. Project page: https://mikuz12.github.io/wing/

What This Paper Is About

Training general-purpose robot policies needs large amounts of real-world interaction data, but collecting it by teleoperation is expensive and hard to scale, whereas egocentric human videos contain abundant first-person human–object interaction that could serve as a substitute source of experience. The problem is that existing latent action models infer actions by reconstructing consecutive frames, so they capture any visually predictive variation — including task-irrelevant ego-camera motion — rather than the interaction dynamics that are actually transferable, and human and robot executions of the same task differ substantially in temporal dynamics. WING's goal is to extract interaction-centric latent actions from egocentric video and transfer them to robot control as structured guidance.

Key Contributions

  1. Characterization of nuisance structure in egocentric latent actions. The authors show that observer-induced variations become entangled with interaction-relevant dynamics in egocentric latent actions, and further show that low-frequency latent dynamics exhibit stronger correspondence across human and robot interactions than full or high-frequency trajectories.

  2. The WING framework. Motivated by those findings, WING combines interaction-centric latent action learning (WING-LAM) with spectral latent guidance, suppressing observer-induced variation while exploiting temporal structures that are more consistently shared across embodiments.

  3. Extensive validation. The authors validate WING through representation analyses and closed-loop experiments in simulation and on a real-world dual-arm platform, reporting that it outperforms representative robot policies and alternative egocentric pretraining strategies.

  4. A demonstrated pretraining recipe. A controlled comparison under the same 200 h egocentric pretraining budget shows that predicting spectral latent guidance to condition action generation outperforms video-only pretraining, hand-pose supervision, direct latent-action prediction, and a non-debiased latent action model.

Main Findings

  • Observer motion dominates reconstruction-based latent actions, but WING-LAM suppresses it. In the MLP-probe analysis of Figure 3(a), both the WING-LAM teacher and student retain strong interaction-motion information while encoding substantially less camera-motion information than DreamDojo; CD-LAM, which is explicitly designed to reduce action-irrelevant variation, still retained more camera-motion information under the authors' evaluation protocol.

  • Lower latent sensitivity under camera perturbation, without losing action semantics. Under progressively stronger camera-motion perturbations (Figure 3(b)), the WING-LAM student shows the lowest latent sensitivity. It also achieves the highest average action-category classification accuracy using frozen latent representations across the LARYBench egocentric datasets (Figure 3(c)).

  • Shared cross-embodiment structure lives in low frequencies. In Figure 3(d), low-frequency components of latent action trajectories produce a substantially larger separation between semantically matched and mismatched human–robot pairs than either the full trajectory or the high-frequency components.

  • State-of-the-art simulation success rates. WING reaches average success rates of 99.20% on LIBERO (Table 2(a): Spatial 99.00, Object 100.00, Goal 98.80, Long 99.00), 93.80% on RoboTwin 2.0 (Table 2(b): Clean 94.56, Randomized 93.04), and 57.7% on RoboCasa–GR1 (Table 1). In each case, the authors report these as outperforming the strongest baseline — for RoboCasa–GR1, the compared methods range from GR00T-N1.6 at 47.6 to LDA-1B at 55.4, with FastWAM (reproduced) at 51.9. On LIBERO the next-best average is LingBot-VA at 98.50; on RoboTwin 2.0 it is LingBot-VA at 92.20.

  • Strong real-world performance. WING achieves the highest average success rate of 75.0% under the standard real-world setting and consistently outperforms FastWAM and GR00T-N1.7 in average success rate and progress score across both standard and generalization settings. Under generalization settings, WING reaches comparable average success to π0.5 while attaining a slightly higher average progress score.

  • Both components contribute independently. In the ablation without ego-pretraining of the WAM (Table 3), removing WING-LAM or the DCT module each reduces average performance across LIBERO and real-world evaluations. Combining both gives the best LIBERO suite scores (99.6 / 100.0 / 99.2 / 96.2) and real-world base 72.5 / generalization 45.0, versus 65.0 / 36.0 with neither component.

  • Four retained DCT coefficients work best. Sweeping K ∈ {2, 4, 8}, where K = 8 corresponds to the full spectrum, K = 4 consistently performs best across all suites. K = 2 is overly restrictive and discards useful temporal dynamics, while the full spectrum adds high-frequency components without improving downstream control.

  • Spectral latent guidance beats other egocentric pretraining strategies. Under the same 200 h budget, video-only pretraining (Vid-P.) consistently underperforms; replacing WING-LAM with a non-debiased DreamDojo LAM degrades performance in most settings; and direct latent-action prediction (LA-P., following LAPA) also performs worse. The authors conclude ego-derived latent actions work best as auxiliary guidance rather than as direct action targets.

Methodology in Plain English

The system has two parts that are trained in sequence.

Part 1 — Learning interaction-centric latent actions (WING-LAM). The insight is that egocentric observer motion and hand–object interaction produce different spatial motion signatures: observer motion changes the whole image coherently, while interaction motion is localized. The authors use an offline point tracker to find corresponding points between two video frames, use training-time region masks to split background tracks from hand–object interaction tracks, and fit a global image warp from the background tracks as the observer-motion target. After compensating for that warp, whatever displacement remains for each tracked point is the interaction-motion target. A two-branch teacher — one branch for camera, one for interaction — is trained on these privileged motion cues. Because these cues are not available at deployment, the teacher's interaction latent is distilled into an RGB-only student that takes just a pair of frames. The distillation loss has two terms: matching the teacher's interaction latent on both the original and a camera-perturbed frame pair, and forcing the student's outputs on those two views to agree, so the representation is consistent under camera-like warps.

Part 2 — Spectral latent guidance. WING-LAM encodes H successive transitions into a latent trajectory matrix, then applies an orthonormal DCT along the temporal dimension and keeps only the lowest K components as the guidance target. This truncation affects the guidance only, not the robot actions. Because the target requires future observations, a lightweight predictor estimates it from the current observation, language instruction, and proprioception. The predicted coefficients are projected into the action model's hidden dimension, LayerNorm-ed, tagged with a learned frequency-mode embedding, and injected into the action model's hidden sequence via gated residual attention, where a learned scalar per layer controls how strongly the guidance influences the network.

Training. Stage one pretrains on curated egocentric data using Wan2.2, training a video expert for future-frame generation and the guidance predictor for low-frequency spectral guidance. Stage two transfers the pretrained visual backbone and guidance predictor to the robot policy and jointly trains them with the action expert on robot demonstrations; the guidance predictor is supervised on robot data while the WING-LAM encoder stays frozen. Future frames are used only to construct training targets.

Why This Matters

Impact on research. The paper reframes the egocentric-to-robot transfer problem: rather than modeling all observed dynamics, robot learning from human experience should identify the interaction structures that remain consistent across observers and embodiments. It provides both a diagnostic analysis (observer variation is retained in reconstruction bottlenecks, characterized with a linear LAM proposition in Appendix A) and a practical recipe, arguing that ego-derived latents work better as auxiliary guidance than as direct action prediction targets. This is directly relevant to the growing line of work on VLA models, world-action models, and latent action models that rely on costly robot data.

Real-world applications:

  • Generalist manipulation policies that generalize across task suites and scene randomizations, as measured on LIBERO, RoboTwin 2.0, and RoboCasa–GR1.
  • Bimanual and humanoid robot control, since the evaluation spans single-arm, bimanual, and humanoid manipulation benchmarks (RoboCasa–GR1) and four real-world bimanual tasks.
  • Learning from wearable/ego-camera footage, where the method's key claim is that scalable human interaction data can substitute for some robot demonstrations.
  • Robustness under distribution shift, since WING is evaluated under both standard and diverse generalization settings in the real world and under randomized scenes in RoboTwin 2.0.

Industry relevance. The paper targets the central cost bottleneck in robot learning — teleoperation data. A pretraining pathway that extracts usable control guidance from human video could reduce the amount of robot demonstration data required; the authors note that WING retains stronger downstream performance as robot supervision becomes more limited (Appendix E.3.3 and E.3.4).

Future Directions

  • More self-contained, depth-aware motion modeling. Constructing the WING-LAM teacher currently relies on external motion cues and an image-space camera–interaction decomposition in which the dominant observer motion is approximated by a single global affine warp — an approximation that can break down under strong parallax, large depth variation, or pronounced 3D camera motion, potentially injecting noise into the interaction supervision.

  • More efficient guidance mechanisms. Latent-action pretraining and spectral guidance add training cost and pipeline complexity; the authors flag efficiency as open work.

  • Determining how much spectral bandwidth is optimal in general. K = 4 wins on the tested suites, but whether the ideal number of retained DCT coefficients generalizes to other embodiments, tasks, and sampling rates is left open.

  • Extending the transfer claim beyond the evaluated platforms. The evidence covers three simulation benchmarks and four real-world bimanual tasks; whether the same interaction-centric spectral guidance transfers to broader embodiments and richer manipulation skill sets is not established.

Target Audience

Researchers and engineers working on robot manipulation policies, VLA/world-action models, and imitation learning, especially those interested in learning from human video rather than teleoperation. It is also relevant to practitioners in humanoid and bimanual robotics who care about cross-embodiment transfer, and to representation-learning researchers interested in disentangling nuisance variation from task-relevant dynamics. Readers without a background in latent action models, robot policy learning, or spectral analysis will find the methodology demanding, though the core arguments about observer motion and low-frequency structure are stated in accessible terms.

Authors’ abstract

Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos offer abundant human interaction experience with task-relevant semantics for robotic manipulation, but direct transfer is challenging for two reasons: latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion, and human and robot behaviors often exhibit different temporal dynamics. We propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then exploits the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures, identifying shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and using them to guide action generation. WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1, and also performs strongly across four real-world manipulation tasks under diverse generalization settings. These results show that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. Project page: https://mikuz12.github.io/wing/

Read the original paper