Current mobile manipulation research predominantly follows an instruction-driven paradigm, where robots rely on predefined textual commands to execute tasks. However, this setting confines robots to a passive role, limiting robotic autonomy and the ability to react to dynamic environmental events. To address these limitations, we introduce Sound-Triggered Mobile Manipulation (STMM), where robots must actively perceive and interact with sound-emitting objects without explicit action instructions. To support STMM, we develop Habitat-Echo, a simulation platform that integrates sound rendering with physical interaction. We further propose a hierarchical baseline that translates high-level planning into low-level executions, where a task planner predicts a skill chain for policy models to execute sequentially. Experiments indicate the feasibility of perceiving auditory events and executing corresponding physical interactions without explicit instructions. Notably, in challenging multi-event scenarios, the robot successfully isolates the primary sources from overlapping acoustic interference to execute the first interactions, and subsequently proceeds to manipulate the secondary objects. These new challenges of planning-to-execution position STMM as a measurable research direction.The code and datasets will be released upon acceptance.
Figures & tables
Figure 2: Task taxonomy and event settings of STMM. (a) STMM comprises two interaction types according to the physical properties of sound-emitting targets: object relocation for rigid sources and state transition for articulated sources. (b) Both interaction types are instantiated under single-event and multi-event settings, and the latter requires robots to resolve concurrent events sequentially according to a specified priority.
Figure 3: (a) Comparison of different simulators. Habitat-Echo supports both sound rendering and physical interaction. (b) Visual distractor objects. STMM-12k includes 77 additional visual distractor objects. (c) Distribution of sound instances. Instances are color-coded by category. (d) Episode statistics. Train and test splits. ( 10.47k and 1.39k episodes) O.R. denotes Object Relocation, and S.T. denotes State Transition. Interaction type of the first sound source → the second one.
Figure 4: Overview of the proposed baseline. The sound-triggered task planner processes the initial observation to reason and generate a high-level skill chain from the skill library (right) . Guided by this chain, specialized policy models are sequentially activated to generate low-level actions and interact with Habitat-Echo.
Figure 5: Results on STMM-12K. (a) Success rates of task planning and skill completion. Fully-supervised SSP achieves the best performance in both settings. Best and second best . Progressive completion success rate of (b) object relocation (single-event), (c) state transition (single-event), and (d) multi-event setting. ∗ denotes the target position that is provided for the navigation policy model.
Figure 6: Qualitative Visualizations of Task Execution. Execution trajectories for (a) Object Relocation, (b) State Transition, and (c) Multi-event Setting. For each task, the left panel maps robot paths and sound sources from a top-down view, while the right panel shows third-person keyframes. Multi-event Setting requires sequential navigation and interaction with two distinct sound sources.
Table 1: Ablation Study. (a) Effect of sound modality on navigate policy model. Cat. and Tgt. Pos. denote the category embedding and target position of the sound source, respectively. (b) Task Planning. Combining Head RGB and sound is the best. (c) Policy model for Navigate (Visual Modality). Using Head RGB and Head Depth yields the highest. (d) Policy model for Pick and Open Washing Machine (O.W.M.). Combining Head and Arm Depth outperforms single-view.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Input Modality
S.R. ( ↑ )
Head RGB
Sound
I
✓
✗
6.25
II
✓
✓
68.25
III
✗
✓
69.50
Appendix
Table 2: Ablation study on input modalities for Setting-Specific Planner (SSP) in the multi-event setting.
Figure 7: Additional Qualitative Results of Different Skills. (a), (b) Success and failure cases of the Pick skill. (c), (d) Success and failure cases of the Place skill. (e), (f) Success and failure cases of the Open Door skill. (g), (h) Success and failure cases of the Close Sink skill.
Table 3: Numerical Results on the STMM-12k. We report the success rate ( % ) of task planning ( SRplan ), individual skill completion ( Individual ), and progressive completion success rate of different skills across different settings. The higher the better. (a) Evaluation of object relocation . (b) Evaluation of state transition . (c) Evaluation of multi-event setting . Success rates are reported in the format “First Sound Source / Second Sound Source”. ∗ denotes the target position that is provided for the navigate policy model.
Figure 8: Architecture of low-level policies and Setting-Specific Planner (SSP) . (a) Navigation policy at the single-event setting. (b) Low-level policy at the state transition task. (c) Navigation policy at the multi-event setting. (d) SSP at the single-event setting.
Department of Computer Science, University College London, United Kingdom. · Department of Artificial Intelligence, University of Malaya, Kuala Lumpur, Malaysia.