Demonstrating Arena 5.0: A Photorealistic ROS2 Simulation Framework for Developing and Benchmarking Social Navigation
Authors: Volodymyr Shcherbyna, Linh Kästner, Duc Anh Do, Hoang Tung, Huu Giang Nguyen, Maximilian Ho-Kyoung Schreff, Tim Seeger, Eva Wiese, +8 more
Organizations: Technical University Berlin (TUB), Germany · National University of Singapore (NUS), Singapore · Technical University Munich (TUM), Germany
Building upon the foundations laid by our previous work, this paper introduces Arena 5.0, the fifth iteration of our framework for robotics social navigation development and benchmarking. Arena 5.0 provides three main contributions: 1) The complete integration of NVIDIA Isaac Gym, enabling photorealistic simulations and more efficient training. It seamlessly incorporates Isaac Gym into the Arena platform, allowing the use of existing modules such as randomized environment generation, evaluation tools, ROS2 support, and the integration of planners, robot models, and APIs within Isaac Gym. 2) A comprehensive benchmark of state-of-the-art social navigation strategies, evaluated on a diverse set of generated and customized worlds and scenarios of varying difficulty levels. These benchmarks provide a detailed assessment of navigation planners using a wide range of social navigation metrics. 3) Extensive scenario generation and task planning modules for improved and customizable generation of social navigation scenarios, such as emergency and rescue situations. The platform's performance was evaluated by generating the aforementioned benchmark and through a comprehensive user study, demonstrating significant improvements in usability and efficiency compared to previous versions. Arena 5.0 is open source and available at https://github.com/Arena-Rosnav.
Figures & tables
Fig. 1 : Sample scenes from the Arena 5.0 platform, which provides tools to develop social navigation approaches in highly dynamic and crowded environments. It focuses on social navigation and provides a number of modules to achieve realistic simulation of human-centric environments, developing and testing navigation algorithms on various robotic systems and setups, and simplified extension with new modules.
Fig. 2 : Data flow of the Generation Stage and the Population Stage . The Generation Stage combines multiple SotA technologies to process text inputs into a floor plan image and room asset locations. 3DSGs are used as an intermediate data structure to divide the problem into a text transformation task solvable by an LLM, and a graph transformation task solvable by a spatial GNN. The Population Stage populates the floor plan’s asset zones with 3D models by employing the Asset Placer . A pre-built semantic vector Model Database is queried for a related model, which is arranged into the zone by a Fitter algorithm. After a final post-processing step, the end result is a finished environment consisting of 3D walls and models.
Fig. 3 : System Design of Arena 5.0 : Our central modules are arena_simulation_setup and task_generator and, which provide interfaces for loading worlds, defining world semantics, providing scenarios and benchmarking configurations, managing robot and non-robot entities, interacting with the parameter server, interfacing the navigation stack, photorealistic simulators, pedestrian simulators. The Python-only modules are highly extensible and provide a convenient API that simplifies interactions with both Arena and the ROS2 ecosystem. We provide additional tooling in the form of installers and build/version management tools, building on top of widely used technologies colcon and vcstool . Our peripheral modules Arena-Gen and Arena-Evaluation are installed alongside the core modules and are interfaced directly and implicitly based on the user’s intentions. The lifecycle management of the system is tied together in our separate centralized module arena_bringup .
Fig. 4 : System Design and Data Flow of arena_gen : (1) a user prompt (text or floorplan image) is converted into a 3DSG and edited by the user; (2) a floorplan is generated by the HouseDiffusion [ 25 ] and post-processed; (3) objects are placed into the rooms and realistically arranged by the MiDiffusion network, creating a world; (4) a scenario is generated and edited by the user, selecting specific tasks for that world; and finally exported to arena_simulation_setup for use in training and evaluation in the rest of the Arena ecosystem. The world consists of a classic map definition, as well as additional structural and semantic annotations used to load the world into a simulator and construct semantically relevant tasks.
Fig. 5 : 3dsg_from_text architecture : The pipeline architecture features a staggered design that processes different parts of the prompt with different LLMs. Language Model 1 generates the 3DSG upper half (room and doorway lists) from the user prompt. Language Model 2 then infers the objects in the rooms from both the user prompt and semantic reasoning capabilities. The lists of possible room and asset types are provided as system prompts and can be changed without re-training.
Fig. 6 : Prompt Examples : Text prompts and generated graphs to illustrate the geometric and semantic reasoning capabilities of the 3dsg_from_text stage. Our architecture is capable of understanding and consistently generating multiple graph types, while also retrieving objects from the prompts and inferring semantically sound objects at the same time. We achieve this with a divide and conquer approach by dividing both reasoning tasks between two separate LLMs.
Fig. 7 : Sample outputs of the population stage. Left , Center : Single-room inference for a given room type and a list of objects. The objects are placed in an intuitive way that allows humans to interact with objects in a realistic way. Right : Result of population stage applied to an entire residential floorplan. Different room types (bedrooms, living room) are populated alongside each other. The object positioning takes doorways into account and leaves walking space in every room.
Fig. 8 : User Story of Arena-Gen Web App : We provide a web frontend for our Arena-Gen module, which provides a convenient way for users to create synthetic worlds in an interactive chat. The user is (1) prompted to either write an input prompt or upload a floor plan which is converted into a (2) 3D Scene Graph that can be edited by the user (changing room layout, adjusting obstacles). When satisfied with the resulting 3DSG, the user can (3) generate floorplans based on the 3DSG until a suitable layout is achieved, then the population stage places the desired objects into sound positions within the room. Once the user is satisfied with the world generated in step (3), a default scenario is machine-created and then edited by the user (4). In the final step, (5) benchmarking tasks for this world can be enabled and weighted based on possible world situations. Finally, the world and configurations can be downloaded or exported to Arena directly. Overall, the user has a compact web interface that combines access to the various technologies used across all pipeline stages. All functionalities are exposed through a stateless HTTP backend, with a central /pipeline backend that takes a text prompt and automatically passes through all following steps, for use in our simple CLI client.
Fig. 9 : Example worlds generated using the Arena-Gen module for benchmarking and competition purposes. The worlds were generated using the text "generate 4 difficulty levels of a [hospital, residential, office] environment". The assets are automatically taken from the arena model database and pedestrians spawned with HuNavSim. Notably, a large variety of worlds for each environment type and level can be generated, e.g. 500 environments of hospital level 2. This feature aids quantitative benchmarking and in training new models. Users can also customize room layouts, pedestrian interactions with each other, asset placements, and specific situations using the Arena Architect GUI (shown in the supplementary video).
Fig. 10 : Example plots generated with Arena evaluation module for the conducted benchmark of all planners available in Arena 5.0, plotted on several social navigation metrics
Planner
Type
Input
Description
Applr [ 35 ]
Hybrid
2D Lidar
A hybrid planner combining different approaches for adaptive planning.
Cohan [ 36 ]
Classic
2D Lidar
A traditional planner focusing on human-aware navigation strategies.
Dragon [ 37 ]
Hybrid
2D Lidar
A hybrid navigation system designed for dynamic environments.
DWB [ 38 ]
Classic
Costmap
Modified version of the DWA Dynamic Window Approach, for ROS2
TEB [ 39 ]
Classic
Costmap
Timed Elastic Bands, optimizing a global path by considering kinematic and dynamic constraints.
Graceful [ 40 ]
Classic
Costmap
A Smooth Control Law for Graceful Motion of Differential Wheeled Mobile Robots in 2D Environment.
TABLE I : Overview of the robot and planner suites with descriptions of the planners and indications on which robot platform they can be deployed. Costmap planners are fully integrated with the nav2 framework and accept any number and types of sensors.
Metric
Unit
Explanation
Performance
Success Rate
%
Runs with < 2 collisions
Collision
-
Total number of collisions
Time to reach goal
[ s ]
Time required to reach the goal
Path Length
[ m ]
Path length in m
TABLE II : Overview of evaluation metrics
Parameter
Value
Dataset Size
≈60000
Batch Size
16
Epochs
10
Optimizer
AdamW
Learning Rate
3⋅10−5
Warmup Steps
500
TABLE III : 3dsg_from_text training parameters
Prompt
Generate a house with 3 rooms. Add a toilet to the bathroom
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning. Project page: https://jiaming.im/VideoSocNav
Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen +4
National University of Singapore, Singapore · Singapore Management University, Singapore · The Chinese University of Hong Kong, Hong Kong SAR, China +1
Developing navigation policies requires simulation scenarios that support repeatable training and evaluation. Despite the availability of numerous simulators, constructing diverse navigation scenarios often involves writing simulator-specific code, using complex graphical interfaces, performing substantial manual configuration, or relying on resource-intensive computing platforms, making scenarios difficult to reproduce and reuse across experiments. To this end, we develop the Intelligent Robot Simulator (IR-SIM), a lightweight declarative simulator implemented as a Python library to support navigation learning and benchmarking through rapid construction of reusable and diverse scenarios and efficient execution on accessible platforms. IR-SIM represents scenarios as human-readable YAML configurations, in which necessary objects, behaviors, sensors, maps, and environment parameters can be flexibly specified and composed. This declarative representation makes IR-SIM friendly to large language model (LLM)-powered agents, enabling them to compose and modify executable navigation scenario files through the provided agent skills, instead of writing a large amount of code that is difficult to reuse. Despite its lightweight implementation, IR-SIM provides the essential components for navigation simulation, while retaining interfaces to high-fidelity simulators for downstream validation. Experiments demonstrate that IR-SIM runs up to several hundred times faster on the evaluated CPU platform, while agent skills reduce mean LLM-based scenario construction time by more than 50% across both evaluated models. The experiments further show ability to support reproducible benchmarking and RL-based navigation policy learning.
Ruihua Han, Shuai Wang, Chengyang Li +8
The University of Hong Kong · Shenzhen Institutes of Advanced Technology · Southern University of Science and Technology +2
Pedestrian simulation is a critical component for training and deploying social robot navigation approaches, yet it remains a largely rigid system that repeatedly requires manual data generation to define even simple scenarios. We propose GROVE, a text-to-scenario pedestrian simulation framework that combines state-of-the-art approaches to produce realistic, socially challenging scenarios for social robot navigation. Our framework allows users to customize one of several common presets (emergency, queuing, normal) or even enter a fully independent prompt to generate a highly customizable pedestrian simulation. Multiple modules separately ensure the realism and soundness of long-horizon human behavior, medium-horizon pedestrian navigation, and short-horizon robot/social interactions. Each module is tuned by the prompt in a way that reflects the user intent across all aspects of pedestrian simulation. By dynamically selecting one of several state-of-the-art (SotA) approaches in our modules based on the scenario, we capture many situational nuances of pedestrian behavior in order to narrow the simulation-to-real (sim2real) gap. The human simulation is directly integrated into Isaac Sim, Gazebo, and RViz simulators for robot deployment in highly social environments. We validate our approach through qualitative comparison against existing pedestrian simulation baselines across scenarios of varying complexity in residential, hospital, and office environments. The result is a high-fidelity pedestrian simulation that challenges social robot navigation with complex, diverse, realistic human behaviors.
Duc Tai Nguyen, Volodymyr Shcherbyna, Anh Do Duc +3
1Singapore Management University · 2Technical University Berlin · 3National University of Singapore