ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception
Authors: Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu
Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Beijing Academy of Artificial Intelligence · Beijing Jiaotong University · State Key Laboratory of Multimedia Information Processing, Peking University · AresoX
Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.
Figures & tables
Figure 1 : ROMA is an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. Through a physical interface, ROMA actively acquires missing sensory evidence by integrating visual, audio, tactile, and force feedback in a reasoning-interaction-feedback loop. Given an instruction, ROMA can identify the information required for the task and selectively determine the target object, interaction, and sensory modality to acquire the necessary evidence.
Figure 2 : Overview of ROMI-2K and the formulation of multi-sensory active perception tasks . ROMI-2K combines two complementary real-world data subsets: a large-scale handheld object collection for diverse single-object interactions and a robotic-arm tabletop collection for multi-object scenes, with synchronized visual, audio, tactile, and force feedback. We organize the active perception tasks into perception chains and construct three task types: (1) Single-Chain tasks for an individual attribute, (2) Multi-Chain tasks for multiple attributes requiring interaction planning, and (3) Intent-Driven tasks for inferring and perceiving target attributes from implicit user instructions. Multiple perception chains may also involve reusing feedback across chains.
Figure 3 : Overview of the ROMA system for real-world object-centric multi-sensory active perception . Given an initial scene and a task instruction, the Multi-Sensory LLM identifies the objects, determines the missing evidence, and selects the target object, interaction, and modality needed to acquire it. The physical interface, consisting of a grasp interface and 6 pre-defined interactions, translates these decisions into executable robot actions and collects the visual, audio, tactile, and force feedback. The acquired feedback is returned to the LLM for further reasoning, forming an iterative reasoning-interaction-feedback loop until sufficient evidence is obtained to answer.
Figure 4 : Multi-Sensory Alignment and dynamic scene sampling for training the Multi-Sensory LLM . We illustrate the alignment process using tactile-visual-text alignment as an example. We align tactile representations with visual and textual semantics, enabling the model to better leverage complementary cues across modalities. For dynamic scene sampling, we use Beta-distribution-based sampling to smoothly shift the training from interaction-level supervision toward real-world scene training.
Model
Single-Chain
Multi-Chain
Intent- Driven
Total
Har.
Rou.
Tex.
Ins.
Mat.
Wei.
All
Har.
Rou.
Tex.
Ins.
Mat.
Wei.
All
GPT-5.4
34.3
41.0
38.1
23.1
60.1
38.6
45.1
32.0
37.6
38.1
41.3
42.1
35.8
38.4
56.0
43.5
GPT-6 Astra
70.7
88.5
76.2
32.1
75.9
100.0
76.9
62.8
72.1
65.9
56.3
68.6
75.2
66.6
79.9
72.3
Gemini 3.5 Flash
45.5
66.7
50.0
28.2
58.4
60.8
54.4
48.3
58.8
46.9
46.4
51.8
52.9
49.9
59.4
53.0
Qwen 2.5-Omni
19.2
12.8
9.5
5.1
37.8
5.9
21.1
20.1
16.4
18.8
23.3
28.8
15.0
21.2
26.3
22.0
Qwen 3-Omni
4.0
14.1
19.0
6.4
53.3
0.0
24.7
12.3
10.3
16.6
16.8
22.0
0.7
13.3
28.8
19.7
Table 1: Evaluation of multi-sensory active perception on ROMA Bench across three task types. Har. (Hardness), Rou. (Roughness), Tex. (Texture), Ins. (Inside), Mat. (Material), and Wei. (Weight) denote the target object attributes in single-chain and multi-chain perception tasks. We report task success rate, where failures in object bounding box localization are also counted as task failures. Bold and underlined values indicate the best and second-best results, respectively.
Model
Single-Chain
Multi-Chain
Intent-Driven
Total
GPT-5.4
42.6
29.4
55.6
40.2
GPT-6 Astra
64.8
60.8
66.7
63.6
Gemini 3.5 Flash
50.0
29.4
48.1
41.7
Qwen 2.5-Omni
20.4
29.4
18.5
23.5
Qwen 3-Omni
33.3
15.7
18.5
23.5
ROMA-7B
57.4
64.7
63.0
61.4
Table 2: Evaluation of multi-sensory active perception in real-world scenes using our physical interface across single-chain, multi-chain, and intent-driven tasks. Unsuccessful grasps caused by object localization failures are also counted as task failures. Bold and underlined values indicate the best and second-best results, respectively.
Figure 5 : A case study of MLLM baselines and ROMA-7B on multi-sensory active perception tasks in ROMA Bench. GPT-5.4 tends to terminate with fewer active interactions, occasionally skipping critical objects and sensory evidence. Both GPT-5.4 and Gemini 3.5 Flash also produce inappropriate action-modality combinations, such as attempting to observe gravity force during vigorous shaking. In contrast, ROMA-7B more often continues interaction when the available evidence is insufficient and selects complementary actions and modalities to acquire task-relevant information.
Method
ROMA Bench (w/ Oracle Bounding Box)
Real World
Grasp
Further Interaction
Total Acc.
Grasp
Further Interaction
Per Scene
Per Scene
Per Object
Per Scene
Per Scene
Per Object
Exhaustive
3.68
22.09
6.00
–
3.21
19.27
6.00
GPT-5.4
1.10
1.20
1.09
63.3
0.63
0.72
1.14
GPT-6 Astra
2.29
2.61
1.14
73.5
2.08
2.24
1.08
Gemini 3.5 Flash
2.23
2.77
1.24
68.8
2.24
2.92
1.30
Table 3: Comparison of grasp and interaction counts on completed active perception tasks. For ROMA Bench, oracle bounding boxes are provided to isolate interaction behavior from object localization errors. Per Scene denotes the average number of grasps or further interactions per scene, while Per Object denotes the average number of further interactions per grasped object. Exhaustive grasps every object and executes all six interactions on each object without active selection.
Figure 6 : The illustration of the handheld object interaction data collection platform.
Hardware
Model
Role and Frequency
Handheld UMI
Agilex Pika Sense
Handheld Object Interaction
Wrist Camera
Fisheye RGB Camera × 1
Wrist-view RGB Observation (10 fps)
Third-Person Camera
Intel RealSense D455f × 1
External RGB Observation (10 fps)
Tactile Sensor
GelSight Mini × 2
Optical Tactile Collection (10 fps)
Microphone
USB Microphone
Audio Collection
Table 4: Hardware configuration of the handheld data collection platform.
Figure 7 : An example of a multi-sensory data sample and its annotations from the handheld object subset of ROMI-2K.
Figure 8 : An example of an interaction-conditioned QA sample constructed using the handheld object subset of ROMI-2K.
Figure 9 : An example of a single-chain active perception QA sample constructed using the handheld object subset of ROMI-2K.
Figure 10 : An example of a multi-chain active perception QA sample constructed using the handheld object subset of ROMI-2K.
Figure 11 : The illustration of the real-world robotic arm collection and interaction platform.
Hardware
Model
Role and Frequency
Manipulator
xArm 6
Motion Execution
Gripper
Robotiq 2F-140
Grasp + Tactile Sensor Mount
Wrist Camera
Intel RealSense D435 × 1
Wrist-view RGB Observation (10 fps) Initial Scene (RGB + Depth)
Third-Person Camera
Intel RealSense D435i × 1
External RGB Observation (10 fps)
Force/Torque Sensor
Six-axis Force/Torque Sensor × 1
World-Frame Force (30 Hz)
Tactile Sensor
GelSight Mini × 2
Optical Tactile Collection (10 fps)
Table 5: Hardware configuration summary of the robotic data collection and interaction platform.
Figure 12 : Visualization of the 3D point cloud constructed by the grasping module in the ROMA system, the robot base coordinate frame, and the spatial cuboids. The predicted object bounding box is first projected and uniformly expanded into a search-margin cuboid Csearch , which defines the region for detecting and completing the object point cloud. Each object has an object center cobj and is tightly enclosed by an object cuboid Cobj with an object width wobj , length lobj , and height hobj . The grasp pose for each object can only be generated within the workspace cuboid Cwork , which is obtained by uniformly expanding the object cuboid Cobj by a predefined scale factor.
Figure 13 : Visualization of the grasp poses predicted by AnyGrasp, ROMA without point cloud completion, and ROMA with point cloud completion. For each grasping module, we visualize the top-3 predicted grasp poses and highlight the top-1 grasp pose to be executed in red. The original AnyGrasp predictions often fail to provide sufficiently stable grasps for the robot to perform a sequence of interactions, particularly more vigorous actions such as shaking and rotating. Without point cloud completion, the predicted grasp poses are affected by gaps in the observed point cloud and often result in overly shallow grasps, causing objects to slip or fall. With point cloud completion, the full ROMA grasping module produces more stable and reliable grasp poses.
Figure 14 : An example of a multi-sensory data sample and its annotations from the tabletop scene subset of ROMI-2K.
Figure 15 : An example of an intent-driven active perception QA sample constructed using the tabletop scene subset of ROMI-2K.
Figure 16 : Statistics of object materials in the tabletop scene subset of ROMI-2K.
Figure 17 : Statistics of objects with inside contents in the tabletop scene subset of ROMI-2K.
Figure 18 : Number of tasks by type and the relationship between target attributes in Single-Chain and Multi-Chain QA pairs in ROMA Bench.
Figure 19 : Statistics of the handheld object subset in ROMI-2K. Left: distribution of objects with and without contents. Right: distribution of the top 14 material categories.
Hyperparameter
Setting
Stage 1: Audio Alignment
Trainable Components
Audio Branch
Frozen Components
LLM + Vision Branch + Tactile Branch
Precision
BF16
Effective Batch Size
96
Learning Rate
1×10−4
Table 6: Training hyperparameters for the two-stage alignment and supervised fine-tuning of ROMA-7B. The “Branch” refers to the encoder and adapter.
Figure 20 : The system prompt for Gemini 3.5 Flash in ROMA Bench active perception evaluation.
Figure 21 : The system prompt for GPT-5.4 and GPT-6 Astra in ROMA Bench active perception evaluation.
Figure 22 : The system prompt for ROMA-7B in ROMA Bench active perception evaluation.
Figure 23 : Comparison of action frequencies on completed active perception tasks in ROMA Bench.
Model
Single-Chain
Multi-Chain
Intent- Driven
Total
Ha.
Ro.
Te.
In.
Ma.
We.
All
Ha.
Ro.
Te.
In.
Ma.
We.
All
GPT-5.4
68.7
64.1
38.1
33.3
70.1
92.2
68.2
56.5
58.2
53.7
57.0
58.4
64.7
58.5
67.5
63.3
GPT-6 Astra
73.7
94.9
76.2
32.1
76.6
100.0
78.3
64.3
73.3
66.2
57.6
66.0
75.5
66.9
83.9
73.5
Gemini 3.5 Flash
66.7
82.1
73.8
46.2
69.4
100.0
74.5
62.5
70.3
60.8
61.7
67.0
70.1
65.1
67.5
68.8
Qwen 2.5-Omni
33.3
43.6
28.6
14.1
44.7
7.8
31.3
32.0
32.7
33.8
31.2
36.4
26.7
32.1
31.3
31.7
Qwen 3-Omni
7.1
6.4
9.5
10.3
50.2
3.3
23.6
15.2
14.5
16.9
16.8
23.3
2.9
14.8
31.9
20.5
Table 7: Evaluation of multi-sensory active perception on ROMA Bench under relaxed object localization criteria, with an emphasis on target object and interaction selection, as well as multi-sensory reasoning. Ha. (Hardness), Ro. (Roughness), Te. (Texture), In. (Inside), Ma. (Material), and We. (Weight) denote the target object attributes in single-chain and multi-chain perception tasks. We provide oracle bounding boxes among the answer options for single-chain and multi-chain tasks, and relax the bounding box matching criterion for all three task types. These settings largely reduce the impact of localization errors on physical interaction and subsequent multi-sensory reasoning, while retaining target object identification as part of the active perception process. It is important to note that providing oracle bounding boxes is not a realistic setting in real-world active perception. Bold and underlined values indicate the best and second-best results, respectively.
Variant
Single-Chain
Multi-Chain
Intent-Driven
Total
ROMA-7B
71.1
74.1
73.1
72.9
- Handheld Data
69.2
72.0
70.9
70.9
- Token Init
69.0
73.3
73.4
71.8
- Dual Align
70.3
71.3
72.4
71.1
- Sampling
70.0
73.1
70.0
71.5
Table 8: Ablation study of ROMA on ROMA Bench. Each row removes one component or training strategy from the full ROMA model.
Variant
Single-Chain
Multi-Chain
Intent- Driven
Total
Ha.
Ro.
Te.
In.
Ma.
We.
All
Ha.
Ro.
Te.
In.
Ma.
We.
All
ROMA-7B
48.5
67.9
81.0
59.0
70.8
91.5
71.1
62.5
75.8
77.4
75.1
70.7
80.9
74.1
73.1
72.9
- Audio
47.5
67.9
81.0
24.4
65.6
91.5
65.3
57.2
67.9
70.3
58.1
58.4
74.8
64.3
61.0
64.1
- Touch
42.4
25.6
35.7
60.3
70.8
91.5
63.4
60.6
29.7
43.1
64.6
67.5
72.5
59.5
71.5
62.7
- Force
48.5
67.9
81.0
61.5
70.4
46.4
61.9
57.2
70.9
67.3
64.3
61.8
52.7
61.6
65.0
62.2
Table 9: Evaluation of ROMA under missing-modality settings on ROMA Bench. Each variant returns “Not Available” when the model attempts to acquire the corresponding sensory modality.
Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant information - is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural language. We propose I-Perceive, a foundation model for language-conditioned active perception in large-scale indoor environments. Given a query image, a set of context images, and a natural language instruction, I-Perceive predicts a 6D camera pose that fulfills the specified perception intent. The model integrates a vision-language pathway for semantic grounding with a geometric reasoning pathway for multi-view 3D understanding, connected via multi-layer semantic fusion to enable language-conditioned geometric reasoning. To support scalable training, we construct a large-scale dataset of language-viewpoint pairs from both real-world scene-scanning data and simulated environments using an automated pipeline. Extensive experiments demonstrate that I-Perceive significantly outperforms strong baselines on prediction accuracy, viewpoint feasibility, and instructions alignment. The model exhibits strong zero-shot generalization to unseen scenes and instructions, and enables closed-loop active perception, progressively refining viewpoints over sequential interactions.
Yongxi Huang, Zhuohang Wang, Wenjing Tang +3
Shanghai Jiao Tong University · Shanghai Innovation Institute · Beihang University
Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for robotics rely solely on RGB observations. This limits their ability to perceive physical properties that are difficult or impossible to infer from RGB cameras, such as temperature, sound, or radar response. We present MuseVLA, an adaptive multimodal sensing VLA model that integrates novel sensors as on-demand tools for robotic manipulation. Given a task instruction and visual context, MuseVLA first generates a sensor token and target description that select the sensing modality to invoke and what to attend to, analogous to a tool call with arguments. It then converts the selected sensor measurement into a grounded sensor image, a unified intermediate representation that encodes heterogeneous readings for multimodal fusion and action generation. This design decouples sensor-specific processing from the VLA backbone, enabling efficient integration of diverse modalities. To reduce the need for expensive multisensory robot datasets, we further introduce a data synthesis pipeline that augments existing RGB video datasets with grounded sensor images, enabling generalization to unseen sensor-guided tasks. We evaluate MuseVLA on a real-world robot across challenging dexterous hand manipulation tasks that require multimodal sensing inputs, including temperature-guided pick-and-place, audio-driven object search, and radar-assisted hidden object retrieval. MuseVLA achieves 80.6% success rate on average, outperforming RGB-only and multisensory VLA baselines significantly, and exhibits strong zero-shot capabilities on unseen tasks.
Xingyuming Liu, Ruichun Ma, Heyu Guo +7
School of Computer Science, Peking University · 2Microsoft Research Asia · *Work done during internship at Microsoft Research Asia. +2
Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a framework that advances active perception through coordinated model, data, and hardware designs. Our model augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support a coherent understanding of the scene. To learn from the camera motion naturally present in human activity, we introduce a scalable human--robot mid-training recipe using 1000 hours of egocentric and robotic data, adapting the model to temporal inputs and pose supervision. We further introduce Active-perception Mobile-manipulation Platform (AMP), a robotic platform that supports active perception and mobile manipulation through single-operator teleoperation, enabling scalable collection of demonstrations that coordinate viewpoint changes and manipulation. Experiments demonstrate improved success rates on active-perception tasks, while ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training. Together, these components provide an integrated foundation for studying and developing active perception in robotic manipulation.
Shuai Zhou, Kaisheng Pang, Wenxuan Song +3
Robotics Institute, Carnegie Mellon University · The Hong Kong University of Science and Technology (Guangzhou)