Mobile robots that both navigate and manipulate have a hidden coordination problem: the eyes that help a robot walk to a target are not the eyes that help it pick up a glass. Most designs force one perception stack to do both, and the compromises show up as dropped cups and missed doorways. The MoPA preprint splits the problem in two and rebuilds the bridge at the action layer.
Two parallel perception streams read the same visual scene, one tuned for the base that drives the robot and one for the arm that handles objects. They stay coordinated because each perception head and its action head update together at every layer of a shared decoder. Two specialists in the same room with different checklists, handing a unified plan to wheels and gripper at once.
The 76.3% number is the receipt. Across four real-world household tasks, color sorting, fruit collection, a cross-table block transfer, and a microwave retrieval, MoPA beats the next best design by 12.5 percentage points. The result earns the headline because the design is reusable: any mobile-manipulation system running separate perception heads can borrow the joint decoder pattern.
The honest counterweight is short. Four tasks is not a kitchen, and the result is an author-reported preprint, not peer reviewed. What MoPA settles is the mechanism: the field has been treating navigation and manipulation as one perception problem. The better answer is to split perception, not action.
Reported by Samantha for Type0, from MoPA: Coordinated Mobile Manipulation via Subsystem-Specific Perception Alignment. Read the original: arxiv.org