Skip to content
AI.info

Research

D-JEPA: A Decision-Aligned Latent World Model

Overview Research area: Robotics and world-model-based control (latent/JEPA world models, goal-conditioned planning, robot manipulation, autonomous driving). Technical level: Advanced. One-sentence sc

D-JEPA: A Decision-Aligned Latent World Model
arXiv
2609.24749
Published
2026-09-21
Authors
Shuaijun Liu, Chengyu Wu, Qifu Wen, Feiyang You, Chenglong Zhang, Shuyang Hao, Xi Lin, Ningxin Su

AI summary

Overview

Research area: Robotics and world-model-based control (latent/JEPA world models, goal-conditioned planning, robot manipulation, autonomous driving).

Technical level: Advanced.

One-sentence scope: The paper diagnoses a mismatch between predicted latent goal distance and executed success among the few candidate futures that actually compete for execution, and introduces D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes and realizes them in JEPA-compatible future representations.

What This Paper Is About

Latent world models predict what will happen if an agent takes a given action, and a goal-conditioned planner typically executes the candidate future that looks closest to the goal. The authors show that accurate prediction does not guarantee that this latent distance reflects which candidate will actually succeed: a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative, and this error is concentrated among the handful of candidates competing for execution (the "decision-local prediction gap").

The goal of D-JEPA is to learn, from executed candidate outcomes, which relative distinctions among candidate futures matter at decision time, and then to encode that learned decision structure back into the model's future representations so it can be used through the model's native latent-distance planning interface.

Key Contributions

  1. Decision-local diagnosis. The authors identify and quantify the mismatch between predicted goal distance and executed outcomes near action selection, and formulate candidate-outcome supervision specifically for this regime.
  2. Relational decision alignment. They introduce a bounded, permutation-equivariant operator that learns decision-relevant relations among pretrained candidate futures, complemented by restricted predictor adaptation and an ordinal interface across heterogeneous predictive geometries.
  3. Decision-aligned future representations. They realize the learned decision structure in JEPA-compatible future geometry (bounded temporal transport and exact ordinal realization), so that aligned action choices are recoverable through native latent distance.
  4. Broad validation. They validate decision alignment across simulated latent control, pretrained action-producing models, driving trajectory selection, and physical robots, with additional geometric and visual shift evaluations.

Main Findings

  • The decision-local gap is measurable. On matched PushT decisions, within-start Spearman correlation between predicted and realized latent costs falls from 0.90 to 0.11 for LeWM and 0.80 to 0.13 for TD-JEPA as the full candidate set narrows to the preferred four. Across a 96-start audit, each model prefers a failing candidate on 8 starts despite an available successful alternative, and mean success–failure pair inversion rates within mixed-outcome top-four shortlists reach 49.0% for LeWM and 38.2% for TD-JEPA (top-four means use 34 LeWM and 29 TD-JEPA starts). Among each model's four lowest-cost candidates, within-start Spearman correlation drops below 0.13 for both models.
  • Improved action selection on core tasks. On the PushT confirmation cohort (n = 256), D-JEPA reaches 87.89% success, versus 85.16% for JEPA-WM, 83.59% for LeWM, 82.03% for DINO-WM and 76.95% for TD-JEPA. On Reacher (n = 128) it reaches 93.75%, versus 86.72% for LeWM and 68.75% for TD-JEPA. On Granular (n = 64) it reports 39.06 (main) and 18.75 (strict) with Chamfer distance 0.358; the largest Granular gain is at the strict threshold.
  • Gains over matched adaptation controls. PEGrad-inspired control 82.03%, Var-JEPA-inspired 80.08%, conflict-safe variational adaptation 81.25%, expected-value planning 78.91%, CVaR planning 79.69% on PushT, and four-geometry fusion 85.94%, all below D-JEPA's 87.89%.
  • Transfer to action-producing foundations and physical robots. Across four RoboTwin tasks (128 starts each, 512 total), D-JEPA raises average native VLA success from 61.72% to 76.76% (+15.04 points), with per-task gains of +12.50 (Grab roller), +14.84 (Bread to skillet), +15.63 (Object to cabinet) and +17.19 (Block handover). On two physical PiPER tasks with a shared V-JEPA 2-AC planner (50 paired trials each, 100 total), success rises from 64.0% to 81.0% (+17.0 points): Real PushT 56.0 to 74.0 (+18.0), Two-cube stack 72.0 to 88.0 (+16.0). Gains/losses are 12/3 on PushT and 10/2 on stacking.
  • Driving trajectory selection. Across seven focused scenes (32 shared trajectories per scene, three source logs, equal-weight mean), mean PDMS rises from 57.36 to 95.34 (+37.98) under a candidate-outcome Oracle ceiling of 100.00. Scene-level gains range from +12.87 (S7) to +80.00 (S1).
  • Robustness to shape and appearance shifts. On three unseen PushObj geometries (100 held-out starts each), D-JEPA reaches a mean of 53.67% versus 44.67% for calibrated rank fusion and 28.33% for native predictive selection (+9.00 points over fusion). On fresh PushT starts under seven appearance conditions (50 fresh identities per condition), D-JEPA means 73.14% versus 62.86% for calibrated rank fusion and 56.57% for native selection (+10.29 points over fusion), with the largest appearance gain on object colour.
  • Ablations isolate the relational operator. In the mechanism ablation (256 PushT confirmation starts), TD-JEPA with native distance scores 76.95%, predictive plasticity 79.69% (+2.74), relational alignment 87.11% (+10.16) and calibrated composition 87.89% (+10.94). In the 128-start predictive-source ablation, dual LeWM + TD-JEPA evidence with set relation reaches 93.75%, against 84.38% for LeWM descriptors alone and 83.59% for temporal descriptors alone (+9.37).
  • Multiple geometries help. Under expanded predictive evidence, DINO-WM with native distance reports 95.31 (−3.13), JEPA-WM native distance 98.44 (reference), four-geometry convex fusion 99.22 (+0.78) and multi-geometry alignment 99.22 (+0.78). The same relational operator matches calibrated four-source fusion.
  • Decisions can be read out natively. After exact ordinal realization, native goal distance recovers 100% of relational action choices and candidate ranks. Reported median end-to-end latency per start is 35.03 ms for relational readout, 52.16 ms for calibrated composition and 34.02 ms for ordinal realization.
  • Performance grows with candidate availability. D-JEPA improves as candidate availability increases, which the authors attribute to learning to exploit relations among a richer set of plausible futures.

Methodology in Plain English

The authors begin from the interface that existing predictive models already expose: a set of candidate action sequences, each producing a predicted future and a predicted distance to a goal. Instead of trusting that distance alone, they train a small network that looks at the whole candidate set at once.

Each candidate is described in two complementary ways: a normalized "goal-relative" descriptor (the direction and offset of the predicted terminal latent from the goal embedding) and an ordinal coordinate, which is simply the candidate's rank by the pretrained model's own native cost, rescaled to between 0 and 1. Ranks are scale-free, so evidence from different predictive models can be combined even though their latent spaces differ. In the dual-model configuration the token for each candidate concatenates LeWM and TD-JEPA descriptors and ranks, and the base score is a weighted combination of the two ranks.

A shared encoder and two permutation-equivariant Transformer layers process the complete candidate set, so the model reasons about candidates in relation to one another rather than in isolation. The output produces a bounded, zero-initialized correction (rank-eight, bounded by ε = 0.2) added to the base score. Because the correction is bounded, preferences are preserved across base-score gaps above 2ε and selection can only change within 2ε of the base minimum. Training maximizes decision mass on successful candidates, sharpens local success/failure ordering near the current decision boundary, and penalizes unnecessary global reordering; a calibrated margin gate then chooses between the relational winner and the base winner. Labels come from training-time counterfactual execution; at deployment the model sees only predictive evidence.

Two complementary mechanisms extend the approach. First, predictive plasticity adapts only the final TD-JEPA predictor block and projection, so the predictor geometry itself can respond to decision supervision while adaptation stays localized; its native-distance winner supplies a complementary proposal that calibrated composition can accept when its score advantage exceeds a calibrated threshold. Second, an ordinal interface lets the same relational operator ingest ranks from additional models (JEPA-WM and DINO-WM in the four-model instance).

Finally, representation lifting writes the learned decision structure back into future latents. For same-dimensional models, a small time-conditioned network predicts a signed coefficient bounded by ρ = 0.1 and applies a bounded temporal transport to same-action steps. For PushT, exact ordinal realization sets each candidate's terminal goal-relative direction to the strict rank of its final gated score, so the native mean-squared goal distance becomes (π_i/(K+1))² and the aligned choice can be recovered simply by picking the nearest realized future.

Why This Matters

Impact on research. The paper reframes a standard assumption in latent world modeling: that a globally accurate predictive geometry is sufficient for good control. By quantifying where predictive accuracy and decision quality diverge, and

Authors’ abstract

Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

Read the original paper