R2RI: A Multi-View Event and RGB Dataset for Robot-to-Robot Interaction
Authors: Gabriele Magrini, Riccardo Catalini, Federico Becattini, Guido Borghi, Pietro Pala, Roberto Vezzani, Lorenzo Seidenari
Organizations: Department of Information Engineering, University of Florence, Florence, Italy · Department of Engineering “Enzo Ferrari”, University of Modena and Reggio Emilia, Modena, Italy · Department of Information Engineering and Mathematics (DIISM), University of Siena, Siena, Italy
Understanding and modeling interactions between autonomous agents is a fundamental challenge in robotics, with broad implications for collaborative systems, social robotics, and human-robot coexistence. Although the study of robot interactions has emerged as a compelling research direction, progress has been severely hampered by the absence of large-scale benchmarks. In this paper, we introduce Robot-to-Robot Interaction (R2RI), the first dataset specifically designed to address the Robot-Robot Interaction (RRI) task. R2RI consists of different humanoid robots and realistic interactions modeled on real human social behaviors. Complementary viewpoints are available, \textit{i.e.}, an egocentric perspective from each robot's onboard sensors, and an exocentric perspective from external fixed cameras, thus enabling rich spatial and contextual understanding of the interaction dynamics. The dataset comprises more than 6.5M frames and ≈5000 videos at 120 fps, including Event and RGB domains. We investigate pros and cons of each domain, comparing state-of-the-art approaches for a number of key sensing and interaction based tasks. We publicly release the dataset and its annotations for all tasks and modalities at https://github.com/MagriniGabriele/R2RI.
Figures & tables
Figure 1 : Overview of R2RI, a multi-view, multimodal dataset for Robot-to-Robot Interaction. R2RI provides synchronized RGB and event data from complementary exocentric and egocentric viewpoints, includes challenging illumination conditions, and supports multiple perception tasks, including robot detection, 2D/3D pose estimation, collision detection, and action recognition.
Dataset
Year
#Frames
Images
Multiview
#Robot
Robot type
Tasks
Open-X [ 12 ]
2024
–
RGB-D
22
Arm
RM
Hum. Everyday [ 13 ]
2025
–
RGB-D, Lidar
✓
2
Humanoid
RM, HRI
RoboInter [ 14 ]
2026
86M
RGB
✓
5
Arm
RM
CRAVES-lab [ 15 ]
2019
20k
RGB
1
Arm
RPE
CRAVES-youtube [ 15 ]
2019
275
RGB
1
Arm
RPE
Panda-3Cam [ 16 ]
2020
17k
RGB-D
1
Arm
RPE
Table 1 : Summary of the datasets available in literature proposed for Robot Manipulation (RM), Robot Pose Estimation (RPE), Human-Robot Interaction (HRI), and Robot-Robot Interaction (RRI). Further details are reported in Section 2 .
Figure 2 : Overview of the R2RI \xspace dataset generation pipeline from synthetic robot meshes and real human interactions to the rendering of RGB and event frames.
Figure 3 : Exo examples for both RGB and Event data. R2RI \xspace spans across a variety of environments and interactions, comprising both indoor and outdoor scenarios.
Dataset overview
Attribute
Value
Sequences
≈ 5000
Interaction classes
20
Robot platforms
Atlas, G1, iCub, NAO
Viewpoints
Exocentric, Egocentric (2 egos)
Modalities
RGB, Depth, Event
Table 2 : R2RI dataset statistics. R2RI provides synchronized RGB and event streams from exo and ego viewpoints at 120 fps.
Table 6Table 7
Task
Model
Mod .
Atlas Out
G1 Out
iCub Out
NAO Out
mAP 50 ↑
mAP 50:95 ↑
mAP 50 ↑
mAP 50:95 ↑
mAP 50 ↑
mAP 50:95 ↑
mAP 50 ↑
mAP 50:95 ↑
Det.
YOLO26l [ 59 ]
Event
91.5
71.6
93.4
69.1
92.1
67.9
91.6
67.5
RGB
50.6
21.1
76.5
38.5
79.0
39.0
93.8
66.6
Pose
YOLO26l-pose [ 59 ]
Event
26.7
6.7
53.5
1 6.3
60.9
25.0
56.6
21.1
RGB
25.6
4.6
36.1
11.7
48.9
8.2
22.1
7.7
Table 7 : Results for the “Humanoid Generalization” setting, using a leave-one-domain-out protocol.
Table 9
2D pose
3D pose
Horizon h=15
Horizon h=30
Horizon h=15
Horizon h=30
Model
ADE ↓
FDE ↓
ADE ↓
FDE ↓
ADE ↓
FDE ↓
ADE ↓
FDE ↓
Linear Velocity
25.2
50.2
41.0
82.0
47.3
99.7
100.4
208.6
MLP (512-256)
58.8
68.4
61.2
77.3
134.8
148.5
159.1
191.2
TCN [ 70 ]
47.7
63.9
57.6
75.2
77.4
103.2
117.2
173.9
GRU [ 71 ]
45.0
56.3
53.7
76.0
89.0
112.4
130.9
182.4
Table 10 : 2D and 3D pose forecasting. 2D units: pixels; 3D units: millimetres.
Table 11 : Collision detection(a) and Action recognition(b) results.
Model
All ↑
H1 ↑
EVE ↑
Kepler ↑
DARwIn ↑
Atlas ↑
FIG01 ↑
Toro ↑
Apollo ↑
Optimus ↑
TALOS ↑
Phoenix ↑
Faster R-CNN [ 29 ]
29.6
39.1
76.2
35.6
71.5
56.0
20.4
73.6
27.1
10.5
14.2
15.5
Def. DETR [ 61 ]
24.8
38.4
40.4
17.2
53.7
54.7
28.4
54.0
30.2
13.3
4.0
10.6
RT-DETR [ 63 ]
44.8
99.0
77.9
69.0
49.8
53.5
50.5
32.8
43.9
56.5
30.6
16.9
DINO [ 62 ]
55.4
88.3
87.6
59.8
67.8
80.4
78.3
48.2
40.1
90.3
66.1
3.1
DETR [ 60 ]
19.1
51.6
78.6
7.7
41.6
51.7
9.1
34.1
12.3
8.8
33.5
3.1
YOLO11l [ 30 ]
21.2
99.2
0.0
51.1
17.3
1.7
49.9
0.1
7.2
0.3
3.4
8.8
Table 12 : Zero-shot 2D robot detection AP 50 (%) on DHRP. Models are trained on synthetic exocentric RGB data without fine-tuning on DHRP.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1 : Examples of movement sequences, captured each 10 frames. Annotations for both detection and 2D pose are displayed.
Figure E.1: Robustness to illumination degradation. Qualitative comparison under progressively reduced visibility (100% → 5% lightness). RGB frames become nearly unintelligible at low light levels, while the event modality preserves clear subject silhouettes regardless of illumination.
Half brightness
Quarter brightness
Min brightness
Model
mAP 50
mAP 50:95
mAP 50
mAP 50:95
mAP 50
mAP 50:95
Event Modality
Faster R-CNN [ 29 ]
79.6
49.7
80.2
50.2
79.7
49.0
Deformable-DETR [ 61 ]
83.1
50.0
83.4
50.4
79.5
47.2
RT-DETR [ 63 ]
88.8
72.0
88.9
72.5
67.2
50.0
DINO [ 62 ]
91.3
60.4
91.8
61.0
89.6
58.6
Appendix
Table E.1 : Night Split – Detection . Performance under different illumination levels comparing Event and RGB modalities.
Half brightness
Quarter brightness
Min brightness
Model
mAP 50
mAP 50:95
mAP 50
mAP 50:95
mAP 50
mAP 50:95
Event Modality
RTMPose [ 64 ]
82.4
59.9
82.4
60.1
82.6
60.3
DWPose [ 65 ]
82.3
60.1
82.3
60.3
83.4
60.9
HRNET [ 66 ]
61.6
39.9
62.4
40.5
62.3
42.9
YOLO11n-pose [ 30 ]
76.3
45.8
76.8
46.3
78.6
48.0
Appendix
Table E.2 : Night Split – Pose Estimation . Performance under different illumination levels comparing Event and RGB modalities.
Figure F.1 : Exocentric and Egocentric (for both robots) views of the same scene in different indoor environments.
Figure F.2 : Exocentric and Egocentric (for both robots) views of the same scene in different outdoor environments.