MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Abstract
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
Community
We introduce MM-ABC, a generalist foundation model for mobile manipulation built around Seeing, Coordinating, and Imagining. It combines sparse multi-level VLM features, a two-stream arm–base action transformer with clean-action x-prediction, and future geometric supervision for jointly learning perception and coordinated whole-body control.
MM-ABC is pretrained on 5,000+ hours, 400K+ episodes, and 17 robot embodiments, and evaluated across five simulation benchmarks and real-world mobile manipulation. It achieves 44.71% on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, and 83% average success in real-world tasks.
Project page and the MM-30 real-world mobile manipulation dataset are publicly available. The code is coming soon
Get this paper in your agent:
hf papers read 2609.35652 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper