Research
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Overview Research area: Robotics — generalist mobile manipulation, vision-language-action (VLA) foundation models, robot policy learning from large-scale heterogeneous data. Technical level: Advanced.

- arXiv
- 2609.35652
- Published
- 2026-09-28
- Authors
- Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
AI summary
Overview
- Research area: Robotics — generalist mobile manipulation, vision-language-action (VLA) foundation models, robot policy learning from large-scale heterogeneous data.
- Technical level: Advanced. The paper assumes familiarity with flow matching / rectified flow, transformer architectures, VLM feature conditioning, and embodied benchmarks.
- Scope: MM-ABC is a single foundation model that jointly controls a mobile base and one or two robot arms by combining sparse multi-level VLM visual features, a training-only future-geometry prediction branch, and a dual-stream action transformer with clean-action prediction.
What This Paper Is About
A fixed-base robot arm can only touch what it can reach, so a mobile base turns the reachable workspace itself into something the policy controls. This creates two problems: perception must stay spatially grounded while the camera viewpoint keeps moving, and the arm and base actions must be coordinated even though they operate at very different scales. MM-ABC's goal is to build one pretrained policy that handles both, by seeing the changing workspace, imagining how interaction will reshape it, and coordinating separate arm and base action streams.
Key Contributions
-
MM-ABC foundation model. A large-scale mobile manipulation policy organized around three principles — Seeing, Coordinating, and Imagining Arm–Base Collaboration — evaluated on simulation benchmarks, controlled ablations, and real-world deployment.
-
Sparse DeepStack conditioning and training-only world imagination. Multi-level VLM hidden states are injected as gated residual updates into the action expert, and a future branch predicts geometry-rich features at sparse future horizons from a frozen geometric teacher. The future branch never reads action tokens and is removed at deployment, so it acts purely as extra training supervision.
-
MM-APT (Mobile Manipulation Action Prediction Transformer). A structured dual-stream action transformer that keeps separate manipulation and body streams while letting them exchange information through masked joint attention, using clean-action x-prediction plus asymmetric near–far attention (near-term actions cannot attend to speculative far-future actions). A controlled synthetic study with a known Bayes-optimal denoiser favors clean-action over velocity prediction.
-
Data contributions at scale. A pretraining mixture of 5,166.2 cleaned hours spanning 400K+ episodes, 12 datasets, 51 subsets, and 17 embodiments, plus MM-30, a self-collected real-world mobile manipulation dataset of 30+ hours and 40+ tasks on a HexFellow Trigger-A3 omnidirectional base with two AgileX PiPER-X 6-DoF arms.
Main Findings
- EBench: MM-ABC reaches 44.71% success, exceeding the strongest baseline in that comparison by 3.30 percentage points.
- RoboCasa365: 61.2% task-weighted average, exceeding the strongest baseline by 7.0 percentage points.
- LIBERO and LIBERO-Plus: 99.1% mean success on LIBERO, and the same LIBERO-trained policy reaches 82.8% on LIBERO-Plus with no further training — the highest among compared methods, including ones trained on perturbed demonstrations.
- Real world: 83% mean success across five household, office, workcell, and laboratory tasks, 12 points above π_0.5.
- Clean-action vs. velocity prediction (robustness test): On RoboCasa365 composite-seen tasks, replacing clean-action prediction with velocity prediction drops success from 32.8% to 29.2% (a 3.6-point drop). Removing future supervision or multilevel conditioning causes larger drops, though those specific numbers are not reported in the provided text.
- Synthetic study — target geometry: Clean chunks occupy a prescribed 13-dimensional subspace, while velocity and noise span the ambient space, so a velocity head must reproduce noise directions absent from the clean-action subspace.
- Synthetic study — noise retained in endpoints: With a clean chunk fixed and 24 independent noise draws, x-prediction variance is 0.023 versus 0.54 for velocity prediction at t = 0.01 (a 23× reduction), and 0.35 versus 0.63 at t = 0.2. Ridge probes on final manipulation-token features recover injected noise with held-out R² of 0.62 for x-prediction and 0.94 for v-prediction, against −0.09 for both with permuted rows.
- Synthetic study — accuracy past the Bayes-optimal denoiser: The v/x excess-risk ratio reaches 1.64 at t = 0.01, meaning x-prediction is markedly more accurate in the high-noise regime where every sampling trajectory begins.
- Synthetic study — sampling and coordination: x-prediction lowers task-allocation error at every budget from 1 to 16 Euler steps (0.0106 to 0.0090 at five steps, the budget used by the policy), yields lower off-subspace mass (0.095 versus 0.104 at five steps), and has lower sliced Wasserstein distance to the true latent distribution at 8 and 16 steps (0.075 versus 0.076, and 0.046 versus 0.050).
- Corpus composition: The pretraining corpus contains 5,166.2 cleaned hours and 65K+ unique natural-language instructions; real-world and simulated demonstrations are 80% and 20% of total duration; platforms capable of mobile manipulation contribute roughly 61% of cleaned hours, with 1,371.7 hours (26.6% of the full corpus) both retaining base-control channels and exhibiting nontrivial base motion.
- Embodiment spread: 17 embodiments, valid stored dimensions ranging from 10 to 43, and control frequencies from 15 to 30 Hz.
- Not reported: The provided text lists ManiSkill-HAB among the evaluation benchmarks but gives no ManiSkill-HAB result.
Methodology in Plain English
MM-ABC is trained once on a very large, mixed collection of robot demonstrations and then applied, with the same architecture, to both mobile and fixed-base robots.
-
Put every robot into one shared format. Different datasets use different robots, sensors, and control conventions. The team audited each source, standardized states and actions into a fixed 80-dimensional interface (29 left-arm + 29 right-arm + 22 base dimensions), and kept a validity mask so that unused channels are ignored rather than trained on. Only temporally contiguous, valid trajectories were kept.
-
Balance the data mixture. Raw hours are extremely uneven — some sources contribute 700–900+ cleaned hours and more. Sampling weights scale sublinearly with duration (h_i^0.4) and are capped so no single profile exceeds 0.2 normalized probability, with the excess redistributed. This keeps small but distinctive sources visible during training.
-
Look at several depths of the vision-language model. Rather than conditioning only on the final layer of a Qwen3-VL-4B backbone, the model pulls hidden states from four layers, keeps the last one as the main semantic context, and injects three intermediate ones into the first three expert blocks through zero-initialized gates. Those features are read-only: action tokens read them, but never modify the context sequence.
-
Split control into two collaborating streams. The action transformer keeps separate transformations, normalization, and feed-forward layers for manipulation (58 dimensions) and body motion (22 dimensions), but lets their tokens attend to each other within a segment. Time is divided into a near segment (state token plus the first 32 actions) and a far segment (the remaining 32); far tokens can read near tokens, but near tokens cannot read far tokens.
-
Train on imagination, deploy without it. A separate set of learned future queries predicts frozen VGGT-Omega geometric teacher features at 32 and 64 steps ahead, over an 8×8 grid per camera view. These queries share the perceptual context but never exchange information with the action streams, so the geometry loss shapes shared perception without becoming a policy input. The teacher and the whole future branch are deleted at inference.
-
Predict the clean action, not the velocity. Because the two action streams share a finite network width, the researchers predict the final clean action chunk directly instead of the flow velocity, then convert back to a sampling velocity at inference. They first verified this choice in a synthetic task where the true action geometry and the optimal denoiser are known analytically — so the comparison isolates the prediction target from architecture, conditioning, data, and noise.
-
Run the trained model. At inference, action chunks are generated from Gaussian noise in five Euler steps and a prefix is executed before replanning.
Why This Matters
Impact on research. The paper argues that adding mobility to a manipulation policy is not enough — the model needs representations designed for cross-stream collaboration. It offers a concrete design (masked joint attention over separate arm and base streams, read-only multilevel visual context, a training-only predictive branch) and a controlled way to reason about the prediction target in two-stream action experts. The synthetic study separates the effect of the prediction target from architecture and iterative sampling, and the ablations test the same question inside the full policy.
Real-world applications:
- Household service robots that must fetch, pour, wipe, open, and close objects across rooms rather than a single tabletop — the setting the authors evaluate in five real-world tasks.
- Warehouse, retail, and commercial-space robots that reposition themselves to reach shelves or stock, drawing on the mobility-heavy portion of the pretraining corpus.
- Office and laboratory workcells where one policy drives bimanual manipulation on a repositioning base, including bimanual handovers.
- Fixed-base bimanual manipulation, since the same architecture and LIBERO-trained policy transfer to stationary benchmarks without a mobile base.
Industry relevance. The paper's data engine — auditing heterogeneous sources, mapping them into one masked interface, and rebalancing sampling weights — addresses the practical bottleneck of training robot foundation models on data that was not collected for the same robot. The claim that the future branch and its teacher are removed at deployment matters for latency-sensitive real systems, since no world model is run at test time. The MM-30 collection also documents a specific hardware configuration (HexFellow Trigger-A3 base with two AgileX PiPER-X arms), which is directly useful to teams building on similar platforms.
Future Directions
- Quantifying the remaining ablations. The paper reports that removing future supervision or multilevel conditioning causes drops larger than the 3.6-point clean-action effect, but the provided text does not give those numbers; reporting them would clarify how much each component contributes.
- ManiSkill-HAB and the full benchmark suite. ManiSkill-HAB is named in the experiments but no result appears in the provided text, leaving open how MM-ABC performs there relative to other methods.
- How much mobility is actually learned. Only 26.6% of the corpus both retains base-control channels and shows nontrivial base motion, so a natural question is how performance scales with the share of genuinely mobile data.
- Whether the future branch can be made cheaper or adaptive. The future stream is currently removed entirely at inference; exploring whether an attenuated version at deployment helps in unfamiliar environments, or whether supervision horizons beyond n+32 and n+64 add value, are open design questions.
- Extending the coordination scheme to more heterogeneous subsystems. The masked joint attention and near–far asymmetry are demonstrated on two streams; whether the same structure generalizes to more body subsystems (torso lift, legs, additional arms) is not established.
Target Audience
Robotics and embodied-AI researchers working on mobile manipulation, vision-language-action foundation models, or large-scale multi-embodiment pretraining will get the most from this paper. It is also relevant to engineers building generalist robot policies who care about architecture choices (stream decoupling, visual feature interfaces, prediction parameterization) and about practical data engineering across heterogeneous robot datasets. Readers without a background in flow matching, transformer attention masks, or robot learning will find the technical sections demanding, though the high-level framing of seeing, imagining, and coordinating is accessible.
Authors’ abstract
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.