Skip to content
AI.info

Research

Anisotropic Representations Improve Planning in JEPA World Models

Overview Research area: Robotics / visual model-based reinforcement learning — specifically latent (JEPA-style) world models for goal-conditioned visual planning. Technical level: Advanced. The paper

Anisotropic Representations Improve Planning in JEPA World Models
arXiv
2609.37441
Published
2026-09-29
Authors
Mingu Kang, Yoori Oh, Sookyung Kim, Joonseok Lee

AI summary

Overview

  • Research area: Robotics / visual model-based reinforcement learning — specifically latent (JEPA-style) world models for goal-conditioned visual planning.
  • Technical level: Advanced. The paper pairs a formal geometry analysis (matrix-covariance arguments, regret bounds) with a representation-regularization change and benchmarked visual control experiments.
  • Scope: The paper argues that isotropic Gaussian regularization in JEPA world models induces a planning-misaligned latent geometry, and proposes AnisoWM with Λ Reg — a learnable, constrained anisotropic Gaussian target — to fix it across four visual control environments.

What This Paper Is About

JEPA-based world models such as LeWorldModel (LeWM) plan by rolling out candidate action sequences in a learned latent space and picking the one whose predicted terminal representation is closest (in Euclidean distance) to a goal representation. Because the encoder defines that distance, it also defines how the planner weights terminal errors — yet LeWM's regularizer, SIGReg, targets an isotropic Gaussian purely for representation quality, not for planning alignment. The paper's goal is to show that this mismatch produces positive planning regret even when prediction is accurate and representations do not collapse, and to learn the target variance allocation instead of fixing it.

Key Contributions

  1. A failure-mode result for isotropic regularization. The authors show that the expected prediction–SIGReg objective can select a planning-misaligned latent geometry even with noncollapsed representations and exact conditional-mean prediction (Proposition 1, Theorem 1).
  2. AnisoWM with Λ Reg. A method that replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under a fixed-trace and condition-number constraint, with the prediction objective, predictor architecture, and Euclidean planner left unchanged.
  3. Theory of prediction-driven geometry selection. The analysis characterizes how the learnable target reshapes the state-space metric, gives the target that exactly aligns with a task metric, and shows conditions under which planning regret is driven to zero (Proposition 2, Eqs. 13–15).
  4. Empirical validation across four visual control environments. AnisoWM improves planning success over LeWM in all four, and its latent costs agree better with recorded task outcomes.

Main Findings

  • Isotropic covariance induces inverse-covariance weighting. Lemma 1 shows that if the encoded covariance is isotropic (A Σ Aᵀ = sI), then the induced metric is M = AᵀA = sΣ⁻¹, so latent Euclidean cost equals s(x − x_g)ᵀΣ⁻¹(x − x_g). Directions with lower training variance receive greater weight, and this agrees with a task cost ΔxᵀQΔx for all residuals only if Q is proportional to Σ⁻¹.
  • Joint training actually selects this geometry. Proposition 1 states that for sufficiently small process-noise scale η, every global minimizer of the isotropic-target objective is invertible (despite singular encoders being admissible), the predictor recovers the exact encoded conditional mean, and K_{B,η} → qI_D with M_{B,η} = AᵀA → qΣ⁻¹, while the prediction loss (η/2)tr(K_{B,η}R) → 0. Accurate prediction therefore does not rescue planning alignment.
  • Positive limiting regret on a finite horizon. Theorem 1 constructs a fully actuated linear Gaussian control family (any state dimension D ≥ 2, any fixed finite horizon H ≥ 1, any action bound ū > 0, any nonscalar Σ ≻ 0) where every exact Euclidean latent-planning minimizer has regret converging to Δ_H > 0 as η ↓ 0, even though training prediction loss and fixed-H encoded rollout error vanish.
  • The gap appears with a nonlinear encoder too. In a two-dimensional controlled system with an MLP encoder and action-conditioned predictor, decreasing process-noise scale by 30× substantially decreased held-out conditional-mean prediction error, while mean normalized physical planning regret stayed near 0.16 over ten seeds — closely matching the 0.1607 inverse-covariance reference.
  • Planning success improved in all four environments. With a single shared bound κ = 2: 93% vs 87% in TwoRoom, 89% vs 86% in Reacher, 97% vs 96% in PushT, and 79% vs 74% in Cube (AnisoWM vs LeWM).
  • Better agreement between latent cost and task outcome. Under the planning cost J_pred, ordering agreement improved in all four environments (TwoRoom 0.575 → 0.646; Reacher 0.676 → 0.723; PushT 0.594 → 0.621; Cube 0.537 → 0.553). Under J_enc (representation only), gains were +0.097 in TwoRoom, +0.103 in Reacher, +0.021 in Cube, and −0.005 in PushT.
  • Local cost geometry gets closer to physical distance. Spearman's rank correlation between latent cost and task cost rose from 0.15 to 0.91 in Cube and from 0.65 to 0.92 in PushT.
  • Learned spectra differ across environments. At the shared κ = 2, the number of coordinates above the mean target variance ranges from 95 in PushT to 124 in Reacher, and in TwoRoom the allocation keeps evolving after the condition-number bound is reached — so κ constrains but does not determine the spectrum.
  • More anisotropy is not monotonically better. The sweep over κ shows non-monotone sensitivity, with the best observed anisotropic setting differing by environment; each κ > 1 point is a single training run, so the authors treat the sweep as a sensitivity diagnostic rather than a tuned comparison (numerical values are in Appendix F.5).

Methodology in Plain English

The authors start from an existing recipe: an encoder turns observations into latent vectors, a predictor rolls out candidate actions in that space, and a regularizer (SIGReg, from LeJEPA) pushes the latent features toward an isotropic Gaussian so they do not collapse. Planning then picks the action sequence whose predicted endpoint is nearest the goal latent, using plain Euclidean distance.

Their first step is analytical. In a linear Gaussian model of state, action, and next state, they reduce the joint training objective to a function of one Gram matrix and ask what metric the training ends up preferring. The answer: as process noise vanishes, isotropic regularization dominates and forces the encoder toward a whitening solution, so latent distance behaves like inverse-state-covariance weighting rather than the task's own cost. They then build an explicit control problem where that weighting picks a different action than the true task cost, which yields a strictly positive regret that does not shrink with better prediction. A small two-dimensional MLP experiment checks that the effect survives a nonlinear encoder.

Their fix keeps everything else intact and changes only the training target. Instead of a fixed isotropic Gaussian, they parameterize a diagonal target covariance Λ whose trace is fixed at D and whose condition number is bounded by κ, and they standardize features by Λ^{-1/2} only inside the regularizer. The allocation across latent directions is learned jointly with the encoder and predictor, initialized at Λ = I_D (so κ = 1 exactly recovers the isotropic baseline). They then prove that this broadens the set of achievable metrics to qΣ^{-1/2}LΣ^{-1/2}, that predictive training selects L by minimizing tr(LR) over the feasible set, and that for a Euclidean task metric the aligned choice is the trace-normalized state covariance — which, in their finite-horizon construction with W = I_D and κ = cond(Σ), drives the regret to zero.

Finally they evaluate on TwoRoom, Reacher, PushT, and Cube, reusing LeWM's datasets, architecture, and visual goal-planning protocol, with a 192-dimensional representation, regularization weight λ = 0.09, three training seeds per environment, and CEM planning in the original latent coordinates. They also measure how well latent costs rank recorded action sequences against the outcomes those sequences actually reached, and visualize low-cost neighborhoods around goals.

Why This Matters

  • Impact on research: The paper reframes representation regularization as a planning decision, not only a representation-quality decision. It shows that non-collapse and low prediction error are insufficient guarantees for a usable planning cost, and it offers a minimal, theoretically characterized modification (learn the target covariance) that leaves the predictor and planner untouched. This connects the JEPA/latent world-model literature to the broader line of work on control-relevant geometry (bisimulation, DeepMDP, metric-learning for planning).
  • Real-world applications:
    • Vision-based robot manipulation, where a goal image or goal state must be reached and where errors along different state dimensions matter unequally.
    • Navigation and mobile robotics in structured indoor environments, matching the TwoRoom-style setting.
    • Learned controllers for articulated arms and contact-rich manipulation tasks, matching the Reacher and PushT-style settings.
    • Model-predictive control pipelines that already score candidate action sequences in a learned latent space and want better action ranking without a new planner.
  • Industry relevance: Because the method discards Λ after training and adds no planning-time metric, it can be dropped into existing latent world-model stacks as a training-time regularizer change, incurring no extra inference or planning cost — attractive for deployment where planning overhead matters.

Future Directions

  • Choosing κ in practice. The paper uses one shared bound κ = 2 and reports non-monotone sensitivity with single-run points per κ > 1; how to select κ per environment or per task is left open (numerical sweep values are in Appendix F.5, not in the main text).
  • Exact task alignment and its feasibility. Equation 14 gives the target that exactly matches a task metric Q, but alignment also requires that geometry to lie in the feasible set and that predictive training actually select it — the paper notes the condition rather than solving it generally.
  • Generalization beyond the Gaussian model. The analysis rests on linear encoders, linear predictors, and a Gaussian training distribution; extending the selection rule to nonlinear encoders, non-Gaussian data, and higher-dimensional distribution-dependent spectrum selection (Appendix D) remains to be done.
  • Separating representation geometry from predictor rollout error. The J_enc versus J_pred comparison shows the two contribute differently across environments (PushT's representation-only ordering is nearly unchanged), which raises the question of when fixing geometry alone suffices and when predictor quality must also improve.

Target Audience

Researchers and graduate students working on latent world models, JEPA-style self-supervised representations, and model-based control; practitioners building visual goal-conditioned planners who are willing to engage with the matrix-geometry analysis; and readers interested in how representation regularization choices propagate into downstream decision quality. Familiarity with Gaussian distributions, covariance matrices, and planning costs will help, since much of the argument is expressed in those terms.

Note: author affiliations listed in the paper are Seoul National University and Ewha Womans University. Project website: https://rkdrn79.github.io/AnisoWM-page/

Authors’ abstract

Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with $Λ$Reg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/

Read the original paper