cs.ROOct 3, 2026

ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

Authors: Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Beijing Academy of Artificial Intelligence · Beijing Jiaotong University · State Key Laboratory of Multimedia Information Processing, Peking University · AresoX

Abstract

Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

Figures & tables

Explore similar work

CardsList
  1. I-Perceive: A Foundation Model for Vision-Language Active Perception

    Feb 28, 2026Yongxi Huang, Zhuohang Wang, Wenjing Tang +3Active PerceptionRobot Foundation Models

  2. MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

    Jun 16, 2026Xingyuming Liu, Ruichun Ma, Heyu Guo +7Language-Conditioned Robot ManipulationRobotic Manipulation

  3. ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

    Sep 16, 2026Shuai Zhou, Kaisheng Pang, Wenxuan Song +3Active Vision for Robot ManipulationRobotic Manipulation