cs.CVOct 6, 2026
SaveA Stevens's Power Law Check-up of GPT-5.5's Image-Based Visualization Reading
Organizations: The Ohio State University
Abstract
We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models see no legend. A model first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more interpretable to humans.
Figures & tables
Figure 1 : A new psychophysical measurement of how neural network models respond to visual variables. We adapt Stevens’s measurement of humans’ subjective sense of physical stimulus intensity to measure how machines (GPT-5.5) respond to visual variables, e. g., length, area, color, and texture, and so on. The value describes how observers perceive magnitude changes as the intensity of a visual variable changes. The model’s response scaling is compressive when , linear when , and expansive when . denotes the exponent for the model counterpart, and for human observers where available. Observation. GPT-5.5’s responses scaled from linear to mildly compressive across visual variables. The three colormaps, Greys , plasma , and jet , had negative , indicating that the model reversed the direction of the data values.
Figure 2 : Visual variable used in the experiment. The twelve variants span angle, area, length, luminance, slope, shape, texture, and colormap-based visual representations. Each panel was generated using the input value before the conversation. , each visual variable value linearly maps to value except for Slope, such as Angle , Area (Circle) , Area (Square) , Length , Shape spikes, Texture (point-based) points, Texture (line-based) lines, Luminance . For each colormap, we divided the full colormap to 1000 equal interval, represented each interval by its midpoint color, and map to every tenth interval.
Figure 3 : Data sampling. Data values ranged from 1 to 100 and were divided into 10 bins. Reference (R) and target (T) values were sampled from non-identical and non-adjacent bins. This pilot study used eight reference-target (R-T) bin pairs, , , and . Each R-T pair represents six different random sampled values.
Figure 4 : Measurement Method II: Errors across visual variables. Points show the mean log absolute error and bars show 95% bootstrap confidence intervals. Letters indicate Tukey HSD groups; variables sharing a letter do not differ significantly in error.
Figure 5 : Innate Colormap produced by MLLM vs. the ground truth colormap used in our experiment. Observations. MLLM flipped some values along the map, which caused the negative values in the Stevens’s power law modeling.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| model | gpt-5.5-2026-04-23 |
| reasoning.effort | low |
| text.verbosity | low |
| detail | original |
Table 1: GPT-5.5 configuration used in the pilot study.
Figure 6 : Ten examples for each of twelve visual variables with two examples randomly selected from each of the five sampled bins. Rows correspond to visual variables, and columns correspond to the sampled values of each bin. Black borders are added for clarity and are not part of the original image.