The future of autonomous robots in human-oriented environments depends on their ability to navigate in the absence of prebuilt maps; exploiting navigational aids, such as navigational signs, designed by humans for humans is critical for enabling autonomy. This manuscript introduces a unique dataset collected using a handheld device, at various public spaces across Singapore, representing environments that humans frequent daily, and focusing on sign-centric decision making for mapless navigation. The dataset includes RGB, sparse depth, odometry and IMU measurements of scenarios centered around navigational signs and the complex environments where they are placed. Additionally, we provide 56 long-horizon navigation missions and over 450 sign-centric scenarios. All data is provided in a human- readable format, as well as utility scripts for conversion to ROS 2. Lastly, we provide venue maps and GPS aligned scene graphs of the test environments. We discuss the potential use cases for this dataset.
Figures & tables
Figure 1 : Diversity of missions and environments in dataset. We collate a range of missions across indoor and outdoor environments, including common public spaces from malls to hospitals to parks. Indoor missions extend across multiple floors and interconnected buildings.
Figure 2 : Diversity of signs and scenes. Our dataset has diverse mission scenarios spanning a range of complex signs and scenes.
Figure 3 : Sign-optimal path. Given the mission goal ”COM1”, the robot must follow the sign-optimal path, which agrees with the navigational instruction associated with ”COM1” at every step of the navigation process. Other paths can also lead to the goal, but they are not considered sign-optimal, despite having equal (or lesser) step length.
Figure 4 : VQA format . The VQA options are provided in 3 different formats, to accommodate a variety of methods, from modular to end-to-end approaches: (a) as 3D points in the robot’s coordinate frame; (b) in pixel coordinates across multiple different view frames; (c) in a JSON format.
Sensor
Data Type
Data logged
Frequency
iPhone Camera
1920 × 1440 RGB images
274742
15 Hz
iPhone Lidar
256 × 192 depth images
274742
15 Hz
iPhone IMU
Tri-axial linear accelerations, tri-axial angular velocities
1772208 sets
100 Hz
iPhone Odometry
(x, y, z, q x , q y , q z , q w )
274742
15 Hz
Table 1 : Raw data. Summary of the raw data collected in this dataset.
Figure 5 : GPS-aligned SceneGraphs. When available, we provide GPS-aligned SceneGrpah of the environment, annotated with building, floor and room nodes.
Environment
Scenarios
Area ( m2 )
Type
Missions
multi-building
multi-floor
Outdoor
Length
E1
74
-
campus
M1
✗
✓
✓
4
M2
✓
✓
✗
8
M3
✓
✓
✗
6
M4
✓
✓
✗
9
M5
✓
✓
✗
7
M6
✓
✓
✗
7
Table 2 : Environments and missions. For each environment, we provide information about its area size, type, number of grounding scenarios and the missions. Each mission can span across different floors, interconnected building and urban outdoor spaces. The mission length refers to number of scenarios in the sign-based optimal path.
File Name
Content
Format
images/000000-999999.png
1920 × 1440 RGB image
PNG
depth/000000-999999.png
256 × 192 depth image
PNG
confidence/000000-999999.png
256 × 192 depth confidence map
PNG
imu.csv
Accelerometer and gyroscope measurements
Timestamp ( s ), (x, y, z) linear accelerations m2s−1 , (x, y, z) angular velocities ( rads−1 )
odometry.csv
Local pose estimates
Timestamp ( s ), frame id, (x, y, z) linear translation m , orientation (q x , q y , q z , q w ) rad
camera_matrix.csv
3 × 3 camera intrinsics
(f x , 0, c x , 0, f y , c y , 0, 0, 1)
Table 3 : Scenarios. Human-readable and image data formats and naming conventions for each scenario.
Environment
Success Rate (strict)
Success Rate
VQA Accuracy
Parsing Success
E1
2/13
2/13
57/93
155/166
E2
0/4
0/4
14/26
85/103
E3
2/10
2/10
52/87
109/134
E4
4/6
4/6
32/34
79/90
E5
1/7
1/7
21/43
8/8
E6
0/7
0/7
23/48
63/73
Table 4 : Evaluation. The performance of SignScene on the dataset, on the tasks of sign-based mapless navigation, visual sign grounding and sign parsing.
Collecting real-world navigation data for mobile robots typically requires platform-specific teleoperation, making large-scale collection expensive and difficult to scale. We introduce Universal Navigation Interface (UNI), a robot-free data collection paradigm that uses a four-wheeled rollator walker (rollator) and smartphone to collect physically constrained human demonstrations. Because the rollator cannot climb stairs, negotiate uncut curbs, or pass through narrow gaps, demonstrations are naturally biased toward wheeled-feasible routes. Using UNI, we collect 37.2 km of real-world navigation data and recover metric trajectories that directly supervise goal-conditioned navigation models. Fine-tuning visual-navigation models on UNI reduces trajectory prediction error by 17.4-24.8% on held-out UNI demonstrations. Evaluation on other navigation datasets shows benefits that vary by dataset and metric. We further demonstrate closed-loop transfer to a powered wheelchair in curb, staircase, and curb-cut scenarios. These results support low-cost physical proxies as a practical source of navigation supervision collected without the target robot.
Sarvesh Prajapati, Ananya Trivedi, Lorena Maria Genua +3
Institute for Experiential Robotics, Northeastern University, Boston, MA. · Khoury College of Computer Science, Northeastern University, Seattle, WA, USA.
We contribute Bi3, a dataset of social robot navigation among groups of people in a constrained lab space. Compared to prior data collection efforts for social robot navigation, our dataset is unique in that it features: an original experiment design giving rise to close navigation encounters between two humans and a robot; five different navigation algorithms; two different robot platforms; a diverse participant pool of 74 people recruited from two sites in the USA and France; multimodal data streams including 10.5 hours of human and robot ground-truth motion tracks, RGB video, and user impressions over robot performance. Our analysis of the collected dataset through metrics like interaction density and human velocity suggests that Bi3 represents a benchmark of unique diversity and modeling complexity. Bi3 contributes towards understanding how humans and robots can productively mesh their activities in constrained environments, and can be a resource for training models of human motion prediction and robot control policies for navigation in densely crowded spaces.
Andrew Stratton, Phani Teja Singamaneni, Pranav Goyal +2
Department of Robotics, University of Michigan, Ann Arbor, USA · LAAS-CNRS, University of Toulouse, Toulouse, FR · INRIA, University of Lorraine, Nancy, FR
Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigation (VLN) benchmarks of POI-goal navigation often suffer from coarse granularity or significant sim-to-real gaps due to generated scene. To bridge this gap, we present POINav-Bench, the first benchmark designed for closed-loop evaluation of real-world POI-goal navigation. It comprises 11 commercial areas reconstructed from real-world captures using 3D Gaussian Splatting (3DGS), covering 126,398 m2 in total and spanning 163 distinct POIs. With traversability-aware annotations and reference trajectories, POINav-Bench enables high-fidelity evaluation of navigation agents in realistic, POI-rich real-world environments. Building on this, we propose the POINav Brain-Action Framework where a Brain module performs POI-grounded reasoning to guide an Action module in predicting continuous waypoints for real-world execution. We further curate the POINav-Dataset, containing 70K real-world signage-entrance pairs. Experiments show that our framework provides a viable path toward refining real-world POI-goal navigation.