Research
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training Overview Research area: Robotics — action-conditioned robot world models and video generation for robot learning and poli
- arXiv
- 2610.12468
- Published
- 2026-10-08
- Authors
- Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu, Mingchao Sun, Hongyu Pan, Mu Xu, Lue Fan, Zhaoxiang Zhang
AI summary
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingOverview
Research area: Robotics — action-conditioned robot world models and video generation for robot learning and policy evaluation.
Technical level: Advanced. The paper assumes familiarity with diffusion/flow-matching video generators (DiT, VACE conditioning), reinforcement learning post-training (GDPO, DiffusionNFT), camera calibration geometry, and embodied video benchmarks.
Scope: The paper introduces a multi-view, cross-embodiment robot world model, trained with offline geometric calibration and counterfactual reward-guided post-training, that predicts future videos which follow specified robot actions and produce physically plausible object interactions.
What This Paper Is About
Robot world models must predict how a scene evolves under a given action sequence, but training them on existing robot datasets runs into two problems: imprecise camera calibration misaligns rendered action conditions with recorded videos, and data dominated by successful demonstrations biases the model toward predicting success even when an action fails. DreamTrue addresses both by refining camera geometry from existing RGB recordings alone, and by post-training the generator on counterfactual (perturbed) action trajectories using a reward model learned from human annotations of generation defects.
Key Contributions
- DreamTrue, a multi-view, cross-embodiment robot world model that improves action following and interaction plausibility without collecting additional real-robot training data.
- An offline geometric calibration method that refines camera intrinsics, extrinsics, distortion, and arm mounting offsets from existing recordings, without dedicated calibration sequences or depth sensors. The authors release the resulting calibration for three datasets, covering 153,666 retained episodes and over 1,660 hours.
- A reward-guided post-training approach that learns from human annotations of robot (L1), object (L2), and interaction (L3) defects, enabling feedback under counterfactual actions where no paired ground-truth future video exists. Annotations for 44.9K videos, including 30.4K defect annotations, are released.
- A cross-embodiment action representation that converts embodiment-specific actions into a shared image-space condition (rendered robot RGB, depth, amodal masks, Plücker camera rays, and gripper openings encoded into RGB channels), so one model can be trained across datasets and robot arm types.
Main Findings
- State-of-the-art action following on AgiBot: The full model achieves the highest nDTW of 0.8772 under recorded actions and 0.8831 under counterfactual actions among evaluated methods.
- Large reduction in interaction defects: Reward-guided post-training reduces the human-assessed interaction defect rate from 48.12% to 6.25%, and object defects from 31.88% to 3.12%, while nDTW remains comparable — interaction quality improves without hurting action following.
- Best visual fidelity: The full model reports PSNR 23.06, SSIM 0.936, and LPIPS 0.094 on AgiBot, outperforming EnerVerse-AC (19.17 / 0.863 / 0.158), Genie Envisioner (20.72 / 0.890 / 0.112), DreamDojo (19.27 / 0.812 / 0.163), and GE-Sim 2.0 (19.91 / 0.874 / 0.130).
- First place in the AgiBot World Challenge 2026 world model track: The model records the highest overall EWMScore (0.829), action following (0.9651), and visual quality (0.6246) among participating teams, ahead of PAIWorld (0.8245) and Loop (0.8241).
- Cross-dataset and cross-embodiment generalization: A single shared checkpoint predicts across four datasets. On DROID it improves over Ctrl-World in PSNR (22.00 to 22.95) and LPIPS (0.162 to 0.073). It also generalizes without fine-tuning to an unseen robot embodiment (WidowX250) and unseen real-world scenes.
- More faithful policy outcome evaluation: Across 24 task–policy–setting combinations, the model's predicted success rates fit y = 1.032x − 0.009 versus the baseline's y = 0.919x + 0.087, cutting success-rate MAE from 12.92 to 7.08 percentage points, with Spearman ρ = 0.937 and near-zero aggregate bias (+0.42 pp). Baseline rollouts hallucinate completions on failed grasps, floating above the diagonal.
- Calibration measurably improves alignment: Area-weighted IoU between robot renderings and SAM3 masks improves by 0.228 mean IoU on 130,182 AgiBot episodes and 0.212 on 63,061 DROID episodes, with improvements on 97.1% and 91.6% of episodes respectively. Using RGB only, the method outperforms PointWorld on 75.5% of 38,356 DROID episodes. Applying refined calibration at both training and inference raises PSNR by 1.28 dB on AgiBot and 1.30 dB on DROID, and reduces synchronous position error to 4.60 px and 8.06 px.
- The reward model tracks human judgment: On 4,893 held-out manipulation clips, the embodied video reward model reaches 86.33% average accuracy and 74.21% macro-F1. Per dimension: L1 accuracy 96.69% / macro-F1 69.37 / MAE 0.041; L2 79.24% / 79.00 / 0.243; L3 83.08% / 74.27 / 0.198.
- Human evaluation reliability: Defect rates are based on majority agreement among three independent annotators, with Fleiss' κ = 0.7817 for any-defect labels.
Methodology in Plain English
The framework has three parts.
Offline geometric calibration. For each recording, the pipeline renders the robot's known URDF geometry using recorded robot states and initial camera parameters, then matches rendered pixels to observed pixels (using RoMaV2 for correspondences and SAM3 masks fine-tuned on robot segmentation to restrict matching to arm regions). Two additional correspondence sets are used: static background points tracked across time, and points shared between synchronized camera views. All three sets are optimized jointly — reprojection error for robot-observation matches, and a ray-coplanarity constraint for temporal and cross-view matches — to refine intrinsics, extrinsics, distortion, and arm mounting offsets. No dedicated calibration sequences and no depth sensors are needed. The refined parameters then render the robot along the action trajectory to produce spatially aligned action conditions.
Stage I — supervised training. Embodiment-specific actions are converted into a shared image-space representation: rendered robot RGB, depth maps, amodal masks, and dense Plücker camera rays, plus gripper openings linearly mapped to background RGB intensities in [0, 255]. These are injected into a video DiT (initialized from Wan2.1-VACE-14B) through a VACE conditioning branch, with a 3D convolutional geometry encoder handling masks and Plücker maps. The model processes 101-frame clips with three synchronized views at 240×320 resolution, conditioned on the first RGB frame of each view and the task instruction, and predicts the remaining 100 frames. Training uses multi-view video latents Tiled along the width dimension and a flow-matching objective against paired future videos.
Stage II — counterfactual post-training. To expose the model to failure modes, recorded trajectories are perturbed: an SE(3) perturbation is applied to the final end-effector pose, the path is interpolated from the fixed initial pose to that endpoint, and inverse kinematics converts it into a feasible action sequence — with the initial scene and instruction unchanged. Because no ground-truth future exists for these counterfactual actions, the authors instead train a reward model on human annotations of defects across three dimensions (L1 embodiment, L2 object, L3 interaction). The reward is the negative predicted defect probability, R^k(x̂) = −p_φ^k(x̂). Starting from Stage I, the generator samples groups of future videos per action condition, scores them with the frozen reward model, normalizes rewards within each group following GDPO, and optimizes with DiffusionNFT. Predictions under recorded actions also receive a PSNR reward against paired ground-truth futures with a weight of 0.5, while the L1/L2/L3 channels are weighted 1 each. Only the DiT and VACE LoRA branches are updated (approximately 0.7B trainable parameters, LoRA rank 128, learning rate 5×10⁻⁶); the base generator, geometry encoder, and reward model stay frozen.
Training data. DreamTrue is jointly trained on AgiBotWorld-Beta, DROID, RoboMIND 2.0, and RoboTwin 2.0 — after filtering, 2,232 h of multi-view trajectories across five robot arm types in real and simulated environments. The reward model is fine-tuned from Qwen3.5-9B and trained on a corpus that also includes single-view Bridge dataset predictions (Bridge is used only for reward-model training and validation, not for the video generator). Predictions in that corpus come from DreamTrue and seven external models: DreamDojo, Ctrl-World, EnerVerse-AC, Genie Envisioner, GE-Sim 2.0, Cosmos Predict 2.5, and IRASim.
Why This Matters
Impact on research. The paper shows that the two dominant failure modes of robot world models — misaligned action conditioning and success-biased prediction — can be substantially reduced without new real-robot data collection. The released calibration (153,666 episodes, 1,660.31 h across AgiBotWorld-Beta 1,225.60 h, DROID 239.73 h, and RoboMIND 2.0 194.98 h) and the released defect annotation corpus (44.9K videos, 30.4K defect labels) are reusable assets for the broader embodied-video community. The finding that world-model rollouts align with ground-truth policy success rates (Spearman ρ = 0.937, near-zero bias) strengthens the case for using world models as policy evaluators rather than only as data generators.
Real-world applications:
- Evaluating robot policies in simulated rollouts before deploying them on physical hardware, reducing the cost of real-world trial and error.
- Training and benchmarking manipulation policies that need to handle failed grasps, slipping objects, and unsuccessful contact, which successful-demonstration data underrepresents.
- Cross-embodiment transfer — a single checkpoint generalizes to an unseen robot (WidowX250) and to unseen real-world scenes (self-collected Piper recordings), which matters for labs and companies running heterogeneous robot fleets.
- Cleaning up and reusing existing open robot datasets by correcting calibration errors offline, rather than discarding recordings with bad geometry.
Industry relevance. The work is a collaboration between NLPR/CASIA and Amap, Alibaba Group, and its first-place result in the AgiBot World Challenge 2026 world model track signals industrial interest in world models as an evaluation layer for robot policies. The counterfactual reward approach offers a practical alternative to collecting expensive failure trajectories on real hardware.
Future Directions
- Extend beyond image space. The authors note that the action representation, generator, and VLM reward model all operate in image space, so under occlusion or limited views the framework may generate or reward interactions that look plausible but are physically incorrect. Incorporating 3D or physical-state reasoning is a natural next step.
- Broaden counterfactual construction. Counterfactuals here are generated by perturbing the final end-effector pose and interpolating from a fixed initial pose with inverse kinematics, retaining only kinematically feasible conditions. Other perturbation families (mid-trajectory deviations, contact-force perturbations, multi-stage interventions) remain unexplored.
- Scale reward-model coverage. The reward model is trained on annotations from a specific set of source datasets and eight generators. How well it discriminates defects from generators outside that distribution — and whether the reward can be gamed — is an open question.
- Tighten the loop with policy learning. The paper demonstrates that rollouts support policy outcome evaluation; whether policies trained inside these rollouts transfer reliably to hardware, and how the world model should be updated as the policy improves, is not addressed.
Target Audience
Robotics and embodied-AI researchers working on world models, action-conditioned video generation, or policy evaluation; practitioners building simulation and evaluation infrastructure for manipulation; and engineers interested in dataset calibration, reward-model-driven post-training, or reinforcement-learning fine-tuning of video diffusion models. Readers without a background in diffusion transformers, camera geometry, or RL post-training will find the high-level framing useful but the methodology sections dense.
Authors’ abstract
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.