From Instruction to Event: Sound-Triggered Mobile Manipulation
Organizations: University of Macau · Beihang University · Hefei University of Technology
Abstract
Current mobile manipulation research predominantly follows an instruction-driven paradigm, where robots rely on predefined textual commands to execute tasks. However, this setting confines robots to a passive role, limiting robotic autonomy and the ability to react to dynamic environmental events. To address these limitations, we introduce Sound-Triggered Mobile Manipulation (STMM), where robots must actively perceive and interact with sound-emitting objects without explicit action instructions. To support STMM, we develop Habitat-Echo, a simulation platform that integrates sound rendering with physical interaction. We further propose a hierarchical baseline that translates high-level planning into low-level executions, where a task planner predicts a skill chain for policy models to execute sequentially. Experiments indicate the feasibility of perceiving auditory events and executing corresponding physical interactions without explicit instructions. Notably, in challenging multi-event scenarios, the robot successfully isolates the primary sources from overlapping acoustic interference to execute the first interactions, and subsequently proceeds to manipulate the secondary objects. These new challenges of planning-to-execution position STMM as a measurable research direction.The code and datasets will be released upon acceptance.
Figures & tables
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Input Modality | S.R. ( ) | |
| Head RGB | Sound | ||
| I | ✓ | ✗ | 6.25 |
| II | ✓ | ✓ | 68.25 |
| III | ✗ | ✓ | 69.50 |