Research
MoSE3: Learning World-Space SE(3) at Every Pixel
Overview Research area: Computer vision, specifically monocular 3D motion estimation from video — dense 6-DoF rigid motion (SE(3)) prediction per pixel, 3D point tracking, and articulated/deformable o

- arXiv
- 2610.03716
- Published
- 2026-10-02
- Authors
- Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang
AI summary
Overview
Research area: Computer vision, specifically monocular 3D motion estimation from video — dense 6-DoF rigid motion (SE(3)) prediction per pixel, 3D point tracking, and articulated/deformable object understanding.
Technical level: Advanced. The paper assumes familiarity with SE(3) rigid transforms, Procrustes/Horn alignment, transformer backbones, and geometric foundation models such as π³.
Scope: The paper introduces MoSE3, a feed-forward model that predicts per-pixel world-space SE(3) motion from monocular RGB video via 3D point tracks and learned rigidity embeddings, plus the Art-Kubric synthetic dataset and generation pipeline used to train it.
What This Paper Is About
Existing motion representations from video — optical flow, 2D point tracking, 3D point tracking — are all translation curves per pixel: they record where pixels go, but not how the underlying body rotates, and not which pixels move together as one rigid part. The paper's goal is to predict, in a single feed-forward pass from monocular RGB video, a full 6-DoF rigid transform at every pixel in a shared world coordinate frame, so that rotation, translation, and grouping of pixels into rigid bodies are all captured at once. Because SE(3) labels are scarce and rotations are awkward to regress directly, the authors decompose the problem into two easier, more widely supervised intermediates and recover SE(3) analytically from them.
Key Contributions
-
MoSE3, a feed-forward dense SE(3) predictor. Described as the first feed-forward model for dense SE(3) motion from monocular RGB video, requiring no category priors, no depth sensors, and no per-video optimization.
-
A decomposition of SE(3) into 3D point tracks and rigidity embeddings. The two intermediates are connected by a differentiable closed-form weighted Horn fit, which makes the full prediction end-to-end trainable and lets the model exploit more widely available tracking supervision alongside scarce SE(3) annotations.
-
Art-Kubric, a large-scale synthetic dataset and generation pipeline with dense SE(3) and rigidity labels for articulated scenes, built from PartNet-Mobility, Articraft, and Infinigen-Articulated assets.
-
State-of-the-art results for feed-forward SE(3) estimation at pixel, part, and object levels, and for 3D point tracking, across multiple benchmarks.
Main Findings
-
State-of-the-art dense SE(3) estimation. MoSE3 reports the best per-pixel and object/part-level SE(3) results on both the rigid HO3D benchmark and the articulated iTACO benchmark across all reported metrics. With rigidity-embedding clustering, MoSE3 achieves per-pixel RRE 18.28° and ADD 2.16 cm on HO3D (AUC_r 0.569, AUC_ADD 0.785) and RRE 12.30°, ADD 11.37 cm on iTACO (AUC_r 0.741, AUC_ADD 0.594). At object/part level it reports RRE 17.96°, ADD 2.06 cm, AUC_ADD 0.802, IoU 0.709 on HO3D, and RRE 10.76°, ADD 10.49 cm, AUC_ADD 0.625, IoU 0.784 on iTACO.
-
**Gains over adapted baselines even under the same protocol. The strongest listed baseline on per-pixel HO3D is Track4World-Pi3x at RRE 33.21°, ADD 3.37 cm, AUC_r 0.371, AUC_ADD 0.669; on iTACO it is RRE 28.07°, ADD 24.82 cm, AUC_r 0.447, AUC_ADD 0.312. MoSE3 using the same k-NN protocol as the baselines (rather than its learned rigid clustering) still reports RRE 24.09°/ADD 2.61 cm on HO3D and RRE 20.78°/ADD 14.47 cm on iTACO, outperforming every baseline on every metric.
-
Rigidity embeddings improve grouping. Clustering on the learned rigidity embeddings instead of k-NN lowers per-pixel error and substantially raises cluster IoU — for example, IoU rises from 0.504 (k-NN) to 0.709 on HO3D and from 0.505 to 0.784 on iTACO at the object/part level.
-
Best average 3D point tracking accuracy across three datasets. Following the Track4World protocol on PointOdyssey, ADT, and PStudio from the TAPVid-3D benchmark, MoSE3 reports the best average accuracy at both L-16 (0.5941) and L-50 (0.5946) horizons, beating Track4World's 0.5847 and 0.5402. It is best on every dataset and horizon except ADT L-16, where it ranks second (0.5789) behind Track4World (0.6250). DriveTrack was omitted because its evaluation split has not been publicly released.
-
Generalization from synthetic training data. MoSE3 shows strong generalization to in-the-wild real-world videos despite being trained solely on synthetic motion data, with qualitative results covering rigid, articulated, and deformable cases.
-
Direct SE(3) regression is worse. In the ablation on iTACO under a reduced schedule, directly regressing SE(3) gives RRE 24.57°, AUC_r 0.525, ADD 19.05 cm, AUC_ADD 0.393, versus the full model's 17.27°, 0.647, 14.09 cm, 0.517.
-
Both supervision signals help. Removing the rigidity loss gives RRE 18.97°, AUC_r 0.614, ADD 15.83 cm, AUC_ADD 0.480; removing the SE(3) loss gives 23.90°, 0.555, 17.82 cm, 0.465; a track-only variant with k-NN gives 31.11°, 0.453, 20.96 cm, 0.382, and with rigid clustering 24.30°, 0.550, 18.01 cm, 0.460 — all worse than the full model (17.27°, 0.647, 14.09 cm, 0.517).
-
Art-Kubric contributes, mainly on articulated data. Removing Art-Kubric degrades SE(3) accuracy on iTACO (RRE 22.89°, AUC_r 0.566, ADD 17.01 cm, AUC_ADD 0.462 versus the full model's 17.27°, 0.647, 14.09 cm, 0.517 under the same reduced schedule), while on HO3D it leaves SE(3) accuracy unchanged and lowers only cluster IoU, per the paper's reference to its Table 4.
-
ProxyPose comparison. ProxyPose is reported only at the object/part level under its own single-query protocol with ground-truth intrinsics: RRE 31.22°, AUC_ADD 0.490 on HO3D and RRE 28.51°, AUC_ADD 0.300 on iTACO, with no IoU because it predicts no mask. Its ADD is omitted.
-
Art-Kubric at a glance. 5,000 scenes containing 13–30 rigid and articulated objects, about two-thirds of them in motion, totaling 163M tracks, each with a rigid-group label, per-frame occlusion flag, and its link's SE(3) transform at every frame. Assets span 46 PartNet-Mobility categories, 226 procedurally authored Articraft categories, and 7 Infinigen-Articulated categories.
Methodology in Plain English
The model builds on π³, a geometric foundation model that estimates cameras and dense geometry feed-forward. MoSE3 keeps π³'s geometry branch frozen and attaches a trainable tracking branch alongside it. The tracking branch reuses π³'s image encoder and first few decoder layers, then splits off, with tracking tokens attending to π³'s geometry tokens at every tracking-branch layer after the shared frozen layers. This keeps the predicted tracks aligned with the backbone's pointmaps in co-visible regions.
From the tracking branch, two heads produce the intermediates. A point-tracking head outputs dense 3D track positions in each frame's camera coordinates plus per-pixel visibility logits, alternating frame-wise and global self-attention and conditioned on the query frame through Adaptive Layer Norm. A rigidity-embedding head outputs L2-normalized embeddings that are query-frame invariant, so pixels on the same rigid body get similar embeddings.
SE(3) itself is never regressed directly. For a given pixel, the model takes a weighted Horn fit over the world-frame 3D tracks of its neighbors, where the weights come from a temperature-softmax over rigidity-embedding inner products. Pixels whose embeddings are similar are treated as moving together, so the fit softly solves for the transform that best explains that pixel's rigid neighbors. Because this fit is closed-form and differentiable, SE(3) supervision flows back into both the track predictions and the embeddings, and the whole thing trains end-to-end in a single stage with all losses applied jointly, run in three consecutive phases that differ only in learning-rate schedule.
Training combines four loss groups: tracking losses (normalized image coordinates, inverse-depth-weighted depth, a query-frame surface-gradient term, and a temporal-displacement term), a binary cross-entropy visibility loss, an embedding loss that matches predicted embedding similarities to target affinities (an SE(3)-based affinity where SE(3) labels exist, and a track-geometry-based affinity that needs no rigid-body labels), and a direct SE(3) loss on the recovered transforms using geodesic angle plus a Huber penalty on where the transform places a weighted query-frame centroid.
To supply the missing SE(3) annotations, the authors built Art-Kubric: articulated assets dropped into a textured environment, simulated with physics so multiple bodies interact, observed by a moving camera under randomized HDRI lighting, with per-link SE(3) ground truth for every kinematic link at every frame, plus dense 2D and metric world-space 3D tracks, a per-pixel partition by shared rigid motion, depth, surface normals, and three-granularity segmentation. Resolution and file layout follow Kubric MOVi-F so existing tracking loaders work unmodified.
For evaluation, since no prior method predicts dense world-space SE(3) from monocular video, the authors adapt 3D trackers: for each query pixel they take its K 3D nearest neighbors in the query frame, treat them as one rigid group, and recover a transform from the predicted tracks via Horn, sweeping K per method per clip and reporting the best result. They also compare against ProxyPose, which predicts a pose per query rather than a dense field. For object/part-level evaluation they cluster pixels with HDBSCAN and match clusters to ground-truth part masks by IoU.
Why This Matters
Impact on research. The paper reframes motion estimation from per-pixel translation curves to per-pixel world-space rigid transforms, which is a more structured and compact representation of scene motion. It shows that SE(3) can be learned through two well-supervised intermediates plus a differentiable closed-form fit rather than direct regression, and it demonstrates that synthetic articulated data with physics-driven interaction can transfer to real-world videos. It also releases a labeled dataset and generation pipeline for articulated scenes, where the paper argues prior corpora are rigid-only or dominated by skinned humans and animals.
Real-world applications:
- Robotic manipulation of articulated objects. The paper's motivating example is a robot opening a hinged door: it needs the rotation axis and joint angle, not the trajectories of a million surface points.
- Physical reasoning. Understanding how parts move and rotate relative to each other underpins reasoning about contact, mechanism, and causality in video.
- Articulated-object understanding. Recovering per-part 6-DoF poses directly from monocular video, without CAD models or depth sensors, supports part-level scene understanding.
- Deformable and hand-object interaction settings. The evaluations on HO3D (hand-object) and YCBInEOAT (robot arm manipulating rigid objects) point toward learning from real manipulation footage.
Industry relevance. Removing the need for CAD models, depth/stereo sensors, category templates, and costly per-video optimization makes this class of method more deployable on ordinary RGB video, which matters for robotics, AR/VR scene understanding, and any pipeline that currently consumes point tracks but needs joint axes and part groupings instead. The authors state they will publicly release the Art-Kubric dataset and generation pipeline, and the project page is mose3-tracker.github.io.
Future Directions
- Scaling training and supervision. The ablations were run at small scale with a reduced schedule due to compute constraints, leaving open how much further the decomposition benefits from larger-scale training and how close direct SE(3) regression can get with more annotations.
- Closing remaining benchmark gaps. MoSE3 ranks second on ADT L-16 behind Track4World, and DriveTrack could not be evaluated because its split is not publicly released — both are open comparisons.
- Extending beyond synthetic training data. The model trains solely on synthetic motion data and generalizes to real videos; whether real-world SE(3) and rigidity annotation can be bootstrapped to further improve accuracy is an open question.
- Richer deformable and articulated modeling. The paper frames deformable objects as smooth fields of locally rigid pieces; refining that approximation and validating it on broader articulated and deformable benchmarks is a natural next step.
Details on some quantities are not reported in the provided content: the specific loss weights, training hyperparameters, embedding dimension D, and the full contents of Tables 4, 6, and 7 reside in the paper's appendices and are not included in this excerpt. No inference speed, runtime, or parameter counts are reported either.
Target Audience
Researchers and practitioners in 3D vision, video motion estimation, and geometric foundation models who already work with SE(3), point tracking, or Procrustes/Horn alignment. It is also relevant to robotics and manipulation researchers who need part-level 6-DoF motion from monocular video, and to dataset builders interested in the Art-Kubric generation pipeline for articulated scenes with physics interaction. Because the material assumes comfort with rigid transforms, transformer architectures, and benchmark protocols, it is best suited to advanced readers rather than beginners.
Authors’ abstract
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.