PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
Organizations: Computer Science and Engineering University of Oulu Oulu, Finland · Centre for Applied Computing University of Oulu Oulu, Finland
Abstract
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent's goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.
Figures & tables
| Tool | Description | Tool | Description |
|---|---|---|---|
| draw_dots | Place dots at specific coordinates | draw_rounded_rectangles | Rectangles with a corner radius |
| draw_lines | Segments between two points | draw_circles | From a center and a radius |
| draw_polylines | Connected stroke through points | draw_ellipses | From a bounding box |
| draw_polygons | Closed shape through points | draw_arcs | From a box and angle range |
| draw_regular_polygons | From center, radius, sides, rotation | draw_pieslices | Arc closed to the center |
| draw_rectangles | From two opposite corners | draw_chords | Arc closed by a straight line |
| Metric | human baseline (N=300) | Claude Fable 5 | Claude Opus 5 | Claude Sonnet 5 | GPT-6 Astra | GPT-5.6 Luna | GPT-5.6 Sol | Gemini 3.8 Flash | Gemini 3.7 Flash | Gemini 3.5 Flash Lite | Llama 4 Maverick | Muse Spark 1.3 | Grok 4.5 | Mistral Large 3 675B Instruct | Qwen3.5-9B |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rated creativity | 0.41 | 0.52 | 0.56 | 0.42 | 0.81 | 0.61 | 0.68 | 0.71 | 0.67 | 0.35 | 0.40 | 0.59 | 0.60 | 0.55 | 0.37 |
| on 1–5 scale | 2.39 | 2.79 | 2.95 | 2.43 | 3.95 | 3.12 | 3.45 | 3.55 | 3.40 | 2.12 | 2.33 | 3.11 | 3.12 | 2.94 | 2.21 |
| 0.57 | 0.56 | 0.47 | 0.50 | 0.34 | 0.47 | 0.41 | 0.36 | 0.37 | 0.44 | 0.48 | 0.41 | 0.42 | 0.41 | 0.57 | |
| Percentile of human sample | – | 77.3 | 83.3 | 57.3 | 99.7 | 89.7 | 96.7 | 97.3 | 96.0 | 35.7 | 52.0 | 89.0 | 89.3 | 83.0 | 39.3 |
| Rated recognizability | 0.53 | 0.45 | 0.37 | 0.39 | 0.73 | 0.33 | 0.44 | 0.70 | 0.65 | 0.33 | 0.32 | 0.55 | 0.50 | 0.30 | 0.25 |
| on 1–5 scale | 3.21 | 2.85 | 2.41 | 2.50 | 4.14 | 2.21 | 2.75 | 3.99 | 3.75 | 2.24 | 2.15 | 3.32 | 3.04 | 2.06 | 1.88 |
| Manipulation | Manipulated condition | Baseline | Repl. | Trials | Question |
|---|---|---|---|---|---|
| Interface | SVG markup | tool calls (primary study, 150) | 5 | 150 | Does the drawing interface affect rated creativity? |
| Vocabulary | minimal vocabulary | full vocabulary | 1 | 60 | Does the number of drawing tools affect rated creativity? |
| Blank canvas | no stimulus | stimulus (framing baseline, 30) | 1 | 30 | Does the starting stimulus affect rated creativity? |
| Oracle | subject given | subject withheld (framing baseline, 30) | 1 | 30 | Does the choice of subject affect recognizability? |
| Framing | depictive, example | canonical (primary study’s context) | 1 | 90 | Does the wording of the instruction affect recognizability? |
| Context | context components | primary study’s context | 1 | 240 | How does the per-turn context affect the rated creativity? |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Provider | Identifier | Released |
|---|---|---|---|
| GPT-6 Astra | OpenAI | openai/gpt-6-astra | 2026-09-04 |
| GPT-5.6 Sol | OpenAI | openai/gpt-5.6-sol | 2026-07-09 |
| GPT-5.6 Luna | OpenAI | openai/gpt-5.6-luna | 2026-07-09 |
| Claude Fable 5 | Anthropic | anthropic/claude-fable-5 | 2026-06-09 |
| Claude Opus 5 | Anthropic | anthropic/claude-opus-5 | 2026-07-24 |
| Claude Sonnet 5 | Anthropic | anthropic/claude-sonnet-5 | 2026-06-29 |
| Tool | Arguments | Fields of one operation |
|---|---|---|
| draw_dots | dots, erase? | x, y |
| draw_lines | lines, erase? | x1, y1, x2, y2 |
| draw_polylines | polylines, erase? | points, closed? |
| draw_polygons | polygons, erase? | points |
| draw_regular_polygons | polygons, erase? | x, y, radius, n_sides, rotation? |
| draw_rectangles | rects, erase? | x1, y1, x2, y2 |
| Scorer | with ratings | with ratings | with ink (human) | with ink (agent) |
|---|---|---|---|---|
| ViDrA-adapted (ours) | 0.84 [0.81, 0.87] | 0.82 [0.78, 0.86] | 0.72 [0.71, 0.73] | 0.64 [0.61, 0.67] |
| ViDrA (ours) | 0.75 [0.73, 0.77] | 0.72 [0.69, 0.74] | 0.71 [0.70, 0.72] | 0.77 [0.74, 0.78] |
| AuDrA | 0.63 [0.60, 0.66] | 0.61 [0.58, 0.64] | 0.68 [0.67, 0.69] | 0.86 [0.85, 0.87] |
| Rounds | Drawing operations | Operations not drawn | Rated creativity | Rated recognizability | |
|---|---|---|---|---|---|
| SVG (n = 150) | 3.3 (1.0) | 57.8 (20.6) | 0.4 (1.1) | 0.68 (0.12) | 0.40 (0.15) |
| Tool calls (n = 150) | 7.5 (1.7) | 70.5 (22.6) | 0.0 (0.2) | 0.61 (0.11) | 0.33 (0.12) |
| Marker | Substitutes for MTCI | Definition |
| Exploration rounds | Exploration phase | Rounds before the round of the first ink operation |
| Tool calls | Production phase | Ink calls from the first to the last ink operation |
| Drawing operations | Production phase | Operations contained in those calls |
| Verification rounds | Verification phase | Rounds after the last inking round that contain no ink call and do not finish the trial |
| Rounds to completion | Response time | Rounds in the trial |
| Tool diversity | Flexibility | Shannon entropy of trial’s tool-type distribution |
| Pooled | Within-model | |||
|---|---|---|---|---|
| Marker | 95% CI | 95% CI | ||
| Tool calls | 0.06 | [0.01, 0.11] | -0.01 | [-0.05, 0.04] |
| Drawing operations | 0.53 | [0.50, 0.56] | 0.06 | [0.02, 0.11] |
| Verification rounds | -0.05 | [-0.11, 0.01] | -0.06 | [-0.11, 0.02] |
| Rounds to completion | 0.12 | [0.08, 0.17] | -0.03 | [-0.08, 0.01] |
| Tool diversity | 0.33 | [0.29, 0.37] | 0.13 | [0.09, 0.17] |
| Single-rater | Reliability | Attenuation | Ratings | ||
|---|---|---|---|---|---|
| ICC | demanded | collected | total | ||
| .05 | 57 | 25 | .568 | .754 | 75,000 |
| .10 | 27 | 25 | .735 | .857 | 75,000 |
| .15 | 17 | 17 | .750 | .866 | 51,000 |
| .20 | 12 | 12 | .750 | .866 | 36,000 |
| .21 | 12 | 12 | .762 | .873 | 36,000 |
| Agent term | Share | Example title | Human term | Share | Example title |
|---|---|---|---|---|---|
| cosmic | 10.0% | Cosmic Clockwork Fish | house | 6.3% | a house |
| face | 6.9% | a face | face | 6.5% | a face |
| robot | 8.1% | Robot Face | bird | 4.0% | a bird |
| abstract | 5.4% | Abstract Composition | person | 3.4% | a person |
| sea | 4.8% | Lighthouse by the sea | boat | 2.7% | a boat |
| geometric | 4.2% | Geometric Composition | man | 3.0% | a man |