DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
Abstract
Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We also detail the design principles behind our action interface, and show how the tool output format and the framing of the task combined to determine whether models would drive at all or refuse. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.
Figures & tables
| Initial interfaces (paths) | Final interface | |
|---|---|---|
| Model output | route geometry (waypoints, image-pixel paths or curvature segments) | steering %, speed, duration |
| Tools | observe , submit_path , wait_held , final_stop | observe , set_motion , stop_now |
| Harness owns | path tracking, accepting commands (as valid), turn shaping, endpoint braking | native openpilot limits only |
| Latency | absorbed: the car moves along the route while the model plans, no explicit exposure to the models | explicit: commands only start when accepted, and every response is timestamped |
| New command behavior | gets merged with the rest of the old existing/active path | replaces the running command, and expiration brakes |
| Infeasible/invalid request | rejected (explicitly) | clamped by default, only speeds above the ceiling are rejected |
| Progress per attempt (%) | ||||||||
|---|---|---|---|---|---|---|---|---|
| Model | Harness | 1 | 2 | 3 | Best (%) | Finish time | Tokens | Cost ($) |
| GPT-6 Astra | Codex | 49 | 100 | – | 100 | 5:22 | 7.8M | 9.75 |
| Claude Fable 5.1 | Claude Code | 9 | 10 | 45 | 45 | DNF | 3.6M | 3.95 |
| Grok 4.6 | Cursor | 8 | 11 | 10 | 11 | DNF | 1.0M | 0.65 |
| GPT-5.6 Sol | Codex | 6 | 6 | 6 | 6 | DNF | 1.6M | 1.05 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | # | Progress (%) | End of attempt |
|---|---|---|---|
| GPT-6 Astra | 1 | 49 | Operator brake: drifted onto the left cone line and island at the right turn into the cross aisle (D) |
| 2 | 100 | Finished: the model’s own stop_now inside the finish zone | |
| Claude Fable 5.1 | 1 | 9 | Operator brake: swung left across the cone line toward the median curb (A) |
| 2 | 10 | Operator brake: outside the cone line again, pointed at the median island’s curb (A) | |
| 3 | 45 | Operator brake: too little room for the right turn at D, pointed at the planter island | |
| Grok 4.6 | 1 | 8 | Operator brake: drove straight into the mini-cone gate toward the landscaped island (A) |
| Model | # | Prog. (%) | Dist. (m) | Dur. (s) | Cmds | Cmd/min | Moving (%) | Repl. | Med. gap (s) | Peak (m/s) |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | 1 | 49 | 67.3 | 78 | 8 | 6.2 | 85 | 2/7 | 10.0 | 1.8 |
| 2 | 100 | 134.7 | 322 | 24 | 4.5 | 65 | 2/23 | 12.1 | 1.6 | |
| Claude Fable 5.1 | 1 | 9 | 17.5 | 41 | 3 | 4.4 | 45 | 0/2 | 18.9 | 1.8 |
| 2 | 10 | 27.3 | 190 | 4 | 1.3 | 15 | 0/3 | 60.8 | 1.9 | |
| 3 | 45 | 73.7 | 260 | 8 | 1.8 | 27 | 0/7 | 32.7 | 2.2 | |
| Grok 4.6 | 1 | 8 | 14.4 | 35 | 2 | 3.5 | 35 | 0/1 | 28.7 | 2.0 |