cs.CVOct 6, 2026
SaveA Stevens's Power Law Check-up of GPT-5.5's Image-Based Visualization Reading
Organizations: The Ohio State University
Abstract
We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models see no legend. A model first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more interpretable to humans.
Figures & tables
Figure 1 : A new psychophysical measurement of how neural network models respond to visual variables. We adapt Stevens’s measurement of humans’ subjective sense of physical stimulus intensity to measure how machines (GPT-5.5) respond to visual variables, e. g., length, area, color, and texture, and so on. The value describes how observers perceive magnitude changes as the intensity of a visual variable changes. The model’s response scaling is compressive when , linear when , and expansive when . denotes the exponent for the model counterpart, and for human observers where available. Observation. GPT-5.5’s responses scaled from linear to mildly compressive across visual variables. The three colormaps, Greys , plasma , and jet , had negative , indicating that the model reversed the direction of the data values.
Figure 2 : Visual variable used in the experiment. The twelve variants span angle, area, length, luminance, slope, shape, texture, and colormap-based visual representations. Each panel was generated using the input value before the conversation. , each visual variable value linearly maps to value except for Slope, such as Angle , Area (Circle) , Area (Square) , Length , Shape spikes, Texture (point-based) points, Texture (line-based) lines, Luminance . For each colormap, we divided the full colormap to 1000 equal interval, represented each interval by its midpoint color, and map to every tenth interval.
Figure 3 : Data sampling. Data values ranged from 1 to 100 and were divided into 10 bins. Reference (R) and target (T) values were sampled from non-identical and non-adjacent bins. This pilot study used eight reference-target (R-T) bin pairs, , , and . Each R-T pair represents six different random sampled values.
Figure 4 : Measurement Method II: Errors across visual variables. Points show the mean log absolute error and bars show 95% bootstrap confidence intervals. Letters indicate Tukey HSD groups; variables sharing a letter do not differ significantly in error.
Figure 5 : Innate Colormap produced by MLLM vs. the ground truth colormap used in our experiment. Observations. MLLM flipped some values along the map, which caused the negative values in the Stevens’s power law modeling.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| model | gpt-5.5-2026-04-23 |
| reasoning.effort | low |
| text.verbosity | low |
| detail | original |
Table 1: GPT-5.5 configuration used in the pilot study.
Figure 6 : Ten examples for each of twelve visual variables with two examples randomly selected from each of the five sampled bins. Rows correspond to visual variables, and columns correspond to the sampled values of each bin. Black borders are added for clarity and are not part of the original image.
Explore similar work
Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.
Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.