VehicleArena: A Realistic Urban Environment for Multi-Agent Driving
Organizations: Fudan University · Shanghai Innovation Institute
Abstract
Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.
Figures & tables
| Basic (80 tasks) | Multi-Agent (32 tasks) | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | Arr. (%) | Drive | Req. | Cabin | Arr. (%) | Drive | Req. | Cabin |
| DeepSeek-V4.1-Flash | 62.5 | 85.0 | 88.5 | 76.6 | 56.3 | 81.7 | 82.1 | 73.9 |
| MIMO-V2.6-Pro | 40.0 | 79.5 | 90.0 | 57.5 | 62.5 | 80.0 | 88.0 | 66.7 |
| Kimi-K3 | 51.3 | 78.9 | 84.9 | 41.5 | 40.6 | 71.4 | 76.4 | 49.2 |
| GLM-5.3-Flash | 45.0 | 78.6 | 85.7 | 55.6 | 62.5 | 80.9 | 85.7 | 63.1 |
| Qwen3.8-Max | 65.0 | 83.1 | 86.7 | 72.0 | 65.6 | 82.4 | 88.2 | 65.4 |
| Model | Speed (km/h) | Collisions |
|---|---|---|
| DeepSeek-V4.1-Flash | 19.0 | 4 |
| MIMO-V2.6-Pro | 16.8 | 0 |
| Kimi-K3 | 18.6 | 4 |
| GLM-5.3-Flash | 18.5 | 3 |
| Qwen3.8-Max | 19.2 | 3 |
| Qwen3.8-27B | 16.9 | 5 |
| Basic (80 tasks) | Multi-Agent (32 tasks) | |||||
|---|---|---|---|---|---|---|
| Model | Arr. (%) | Drive | Cabin | Arr. (%) | Drive | Cabin |
| Kimi-K3 | 61.3 ( ) | 74.6 ( ) | 71.6 ( ) | 56.3 ( ) | 77.2 ( ) | 81.6 ( ) |
| GLM-5.3-Flash | 62.5 ( ) | 79.1 ( ) | 75.6 ( ) | 65.6 ( ) | 78.8 ( ) | 83.3 ( ) |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Model | Temperature | Thinking | Effort | Context budget | Output limit |
|---|---|---|---|---|---|---|
| Ego DA | DeepSeek-V4.1-Flash | 0.7 | Enabled | max | 1000000 | 32k |
| MiMo-v2.6-Pro | 1 | Enabled | N/A a | 1000000 | 32k | |
| Kimi-K3 | 0.7 | Enabled | max | 1000000 | 32k | |
| GLM-5.3-Flash | 0.7 | Enabled | max | 1000000 | 32k | |
| Qwen3.8-Max | 0.7 | Enabled | xhigh | 1000000 | 32k | |
| Qwen3.8-27B | 0.7 | Enabled | xhigh | 1000000 | 32k |
| Statistic | Train | Test Basic | Test MultiLLM | Overall |
|---|---|---|---|---|
| Tasks | 100 | 80 | 32 | 212 |
| Maps | 57 | 48 | 13 | 89 |
| LLM vehicles | 1.0 | 1.0 | 4.0 | 1.5 |
| NPC vehicles | 13.1 (8–24) | 13.2 (8–24) | 10.3 (8–13) | 12.7 (8–24) |
| Pedestrians | 0.5 (0–8) | 0.5 (0–8) | 0.0 | 0.4 (0–8) |
| Ego route (m) | 373.8 (22.0–2620.7) | 402.4 (63.4–4867.2) | 199.6 (75.8–427.9) | 358.3 (22.0–4867.2) |
| Wake source | Default rule |
|---|---|
| Shared passenger-visible events | Simulation start, acoustic cue, weather change, and day–night change. |
| Hard braking edge | Acceleration at most for 0.3 seconds; detection is rearmed only after acceleration exceeds . |
| Prolonged-stop edge | Speed below for 15 seconds; detection is rearmed after speed exceeds . |
| Independent random wake | A delay sampled from simulation seconds, unaffected by the Driving-Agent heartbeat. |
| Coalescing | A five-second cooldown merges repeated events by category. A PA may also finish without issuing a request; its update still wakes the Driving Agent once at that boundary. |
| Family | Available predicates |
|---|---|
| Time | A bounded delay after request creation. |
| Route and map | Approaching, entering, or exiting the next intersection; distance to destination below a bounded threshold. |
| Vehicle state | Vehicle stopped, vehicle resumed moving, or a speed threshold held for a bounded duration. |
| Environment | A scheduled weather transition to an exposed condition, or a scheduled transition to darkness. |
| State | Meaning |
|---|---|
| Created / waiting | The frozen request exists; a deferred request is waiting for a feasible selected predicate. |
| Activated / checking | The predicate occurred and the Judge gathers the bounded time-window evidence. |
| Completed | Evidence establishes fulfillment within the request validity period; one-shot requests may close early once all core criteria are established. |
| Uncompleted / unverified | The check budget ends with unmet criteria or insufficient evidence, respectively. |
| Superseded / episode ended | A later PA request replaces the request, or the physical episode ends before its pending lifecycle is resolved. |
| NA / infrastructure failure | An invalid request or insufficient evidence is unscored (NA); operational failures are tracked separately. |
| Metric | Unit and denominator | Definition and companion report |
|---|---|---|
| Trip completion | Percentage of tasks | Ego arrival by without collision or route failure. Peer arrivals are reported separately, not required for ego success. |
| Driving quality | Score in per evaluated vehicle | Starts at 100 with deductions for driving violations; an at-fault collision or red-light violation sets the score to zero. |
| Passenger fulfillment | Score in over valid scored requests | A–F Judge grades mapped to 100, 80, 60, 40, 20, and 0. Request-level and scenario-level means use explicit denominators. |
| Cabin compliance | Score in over checked fields | Field-weighted accuracy pooled across ego checkpoints, including sentinel checks. Rules requiring unavailable equipment are excluded. |
| Externality | Counts and vehicle-seconds per valid pair | Changes in NPC collisions and non-arrivals, plus positive increases in queueing time and trip-completion time relative to SUMO. |
| Efficiency | Input and output tokens per task | Total model-token use, reported separately for input and output. |
| Grade | Score | Evidence-based interpretation |
|---|---|---|
| A | 100 | All requirements are resolved correctly within the expected response time. |
| B | 80 | All requirements are resolved, but after the expected response time and before expiry. |
| C | 60 | All core requirements are resolved, with minor secondary omissions or missing secondary evidence. |
| D | 40 | Useful partial fulfillment of core requirements, or failure to maintain a sustained requirement. |
| E | 20 | Relevant physical action is observed, but it produces no useful requested outcome. |
| F | 0 | No relevant action, an unrelated action, or an outcome contrary to the request. |
| Successful skill loads (tasks) | ||||
|---|---|---|---|---|
| Model | Day/night | Weather | Driving | Safety |
| DeepSeek-V4.1-Flash | 51 | 41 | 62 | 5 |
| MIMO-V2.6-Pro | 37 | 35 | 7 | 0 |
| Kimi-K3 | 14 | 23 | 2 | 0 |
| GLM-5.3-Flash | 30 | 30 | 4 | 0 |
| Qwen3.8-Max | 46 | 48 | 0 | 1 |
| Scene at request time | Passenger request |
|---|---|
| Urban lead-vehicle following in Beijing, in cloudy weather. At 5.0 s, a hard-braking event prompts a request for more separation from the car ahead. | After that hard brake, slow to about 25 km/h so we keep more distance from the car ahead. |
| The same following episode, after rain begins. At 20.4 s, the passenger updates rain-related device settings and plans a quiet final approach. | Once we are within about 130 m of the destination, slow to about 20 km/h for the final approach and turn the music off. |
| Platoon-following traffic in Shanghai. At 14.1 s, fog appears as afternoon changes to dusk, reducing visibility. | It is foggy and we are approaching the destination; slow to about 20 km/h. |
| The same foggy episode at dusk. At 19.1 s, a later request specifies a further slowdown for the final approach and preparation for arrival. | When navigation shows less than 70 m to the destination, slow to about 10 km/h and pause the music to prepare for arrival. |
| An urban crosswalk route in Shanghai at dusk. Rain begins at 15.1 s; the passenger requests wipers, fog lights, and a cautious final approach. | Once we are within 50 m of the destination, slow to below 15 km/h, then stop smoothly after arrival. |
| An urban road with oncoming traffic in Sydney, in the afternoon. At 10.0 s, a passenger who had been reading asks for quieter audio and prepares to pack up. | Within 100 m of the destination, approach gently at under 10 km/h and turn off the passenger reading light so I can pack up. |
| Violation | Deduction |
|---|---|
| At-fault collision or red-light violation | Score becomes 0 |
| Sustained speeding, at least 5 seconds | 20 |
| Short speeding episode, 1 to less than 5 seconds | 5 |
| Unjustified stopping | 5 each |
| Unjustified delay in starting after a green light | 5 |
| Lane change without signaling | 10 |
| Field | Effect |
|---|---|
| Equipment profile | Selects a named set of cabin and sensor modules. |
| enable_modules / disable_modules | Adds or removes named optional modules for one vehicle. Required world and driving modules cannot be removed. |
| Chassis profile | Selects vehicle length, width, acceleration, braking, and lane-change duration. |
| Chassis overrides | Replaces individual positive chassis values for a task-specific vehicle variant. |
| Trusted extension | At run initialization, registers additional modules, equipment profiles, chassis profiles, scenario layers, or rule directories. Scenario data cannot import code. |
| Group | Registered modules |
|---|---|
| World and driving | navigation, map, weather, dayNight, speedLimit |
| Sensors | frontRadar, rearRadar, lidar |
| Climate and comfort | airConditioner, seat, sunshade, readingLight |
| Display, media, and communication | bluetooth, broadcast, centerInformationDisplay, conversation, HUD, instrumentPanel, music, overheadScreen, radio, video, horn |
| Lighting and safety | fogLight, hazardLight, highBeamHeadlight, lowBeamHeadlight, positionLight, tailLight, turnSignal, wiper |
| Body and physical controls | door, footPedal, frontTrunk, fuelPort, rearviewMirror, steeringWheel, sunroof, trunk, window |
| Region | City | Local area(s) |
|---|---|---|
| Mainland China (57) | Beijing | Guomao; Sanlitun; Tiananmen; Wangjing; Wudaokou; Xidan; Yizhuang; Zhongguancun |
| Changchun | Chaoyang | |
| Changsha | Furong; Wuyi | |
| Changzhou | Tianning | |
| Chengdu | Chunxi; Gaoxin; Jinjiang | |
| Chongqing | Jiefangbei |
| Models | #Para | Launch date | Context (tokens) | Scaling | Corporation | License |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | 748B | 2026-09-10 | 1,048,576 | Effort | DeepSeek | MIT |
| MiMo-v2.6-Pro | 1.02T | 2026-09-22 | 1,048,576 | Budget | Xiaomi | MIT |
| Kimi-K3 | 2.8T | 2026-07-17 | 1,048,576 | Effort | Moonshot AI | Kimi K3 |
| GLM-5.3-Flash | 320B | 2026-09 | 1,048,576 | Effort | Z.ai | MIT |
| Qwen3.8-Max | 2.4T | 2026-08-02 | 1,000,000 | Effort | Alibaba | Proprietary |
| Qwen3.8-27B | 27B | 2026-08-14 | 262,144 a | Effort | Alibaba | Apache 2.0 |
| Skill | Use |
|---|---|
| driving_control | Motion commands, perception, routing, and command receipts. |
| weather_transition | Equipment changes for weather and visibility. |
| daynight_transition | Lights, displays, and mirrors around darkness. |
| safety_refusal | Unsafe-request refusal and safe alternatives. |
| Tool | Function |
|---|---|
| navigation__navigation_route_plan | Display a suggested route. |
| navigation__navigation_minimap | Show a route or local map without replanning. |
| navigation__navigation_set_speed | Set target speed and optional acceleration limits. |
| navigation__navigation_emergency_stop | Request emergency braking. |
| navigation__navigation_change_lane | Request one adjacent lane. |
| navigation__navigation_select_maneuver | Select a legal junction connector. |
| Tool | Function |
|---|---|
| get_module_api | List a module’s callable methods. |
| load_tools | Load selected module APIs. |
| load_skill | Load procedural guidance. |
| todo_manage | Maintain tasks across wakes. |
| memory_search | Search passenger, event, and action history. |
| set_heartbeat_interval | Set periodic observation interval ( – s). |
| Tool | Function |
|---|---|
| send_immediate_request | Create or replace an immediate request with outcomes and timing. |
| send_triggered_request | Create or replace a two-phase request with a typed trigger. |
| cancel_passenger_request | Cancel the active request when no pending outcome is wanted. |
| finish | End the wake; keep the active request unchanged. |
| Tool | Function |
|---|---|
| submit_passenger_check | Record intermediate criterion statuses and evidence; no grade. |
| submit_passenger_judgement | Submit a final A–F/NA grade, statuses, and fulfillment times. |