Multimodal LLMs can answer many questions in vision-centric benchmarks without looking at the image, using linguistic priors and skewed answer distributions. Blind evaluation shows what a model already answers this way, but not what it could learn from the test set's own regularities. Extending partial-input auditing (e.g., hypothesis-only baselines in natural language inference), we argue that benchmark designers should "train on the test set": probe the artifact they release for exploitable patterns. Our Test-set Stress-Test (TsT) cross-validates a text-only Qwen2-7B on the test set's questions and answer options, yielding a benchmark-level score and a per-question bias score s(x); a random forest on hand-crafted features adds a fast, interpretable audit. On the template-based VSI-Bench and CV-Bench, the held-out score is 17.9 and 13.1 points above the model's own zero-shot score. On MMMU it learns little, even though GPT-4o correctly answers 52.3% of its multiple-choice questions without the image, so blind success there reflects pretrained knowledge rather than learnable test-set patterns. Iterative Bias Pruning (IBP) removes the questions with the highest s(x) and re-diagnoses; on VSI-Bench it widens a fine-tuned model's vision-blind gap more than random removal at the same rate. We also release VSI-Bench-Debiased, which lowers a fine-tuned model's blind score from 44.7 to 32.0 while its vision score falls only from 57.1 to 48.7.
Figures & tables
Figure 1 : Visual benchmarks have moved from controlled tasks to open-ended VQA. Asking questions in language makes benchmarks more expressive, but it also exposes them to non-visual shortcuts: models can score well by exploiting linguistic patterns without looking at the image.
Figure 2 : Blind scores rise with LLM scale on MMMU but not on VSI-Bench. Blind and vision-enabled scores of LLaVA-OneVision [ 29 ] at three LLM sizes. Red arrows give the blind gain over the smallest model’s blind score (dashed line); blue arrows, the gain from enabling vision. On MMMU, scaling the LLM adds more than enabling vision does; on VideoMME and CV-Bench, both help; on VSI-Bench, blind scores stay flat and the gains come from vision. VSI-Bench values are per-task averages from Yang et al. [42] .
Figure 3 : Statistical biases create non-visual shortcuts. (a) Share of counting questions with each gold count. (b) How often each object is the closer one, for objects in at least 10 questions (dashed line: chance). (c) How often each object occupies each position of the gold appearance order, partly reflecting questions that share a video. (d) Room-size distribution, with geometric mean and geometric standard deviation.
Figure 4 : TsT probes the test set directly. (a) A diagnostic trained on separate data from the source distribution (dotted blue) finds only the test-set biases (pink) that its training data shares and can miss those specific to the test set. (b) The test set is split into k folds; a blind diagnostic trains on k−1 folds and scores the held-out fold, k times over, giving a benchmark-level TsT score and a per-question bias score s(x) .
Qwen2-7B, blind
Benchmark
ZS
TsT
ΔTsT
GPT-4o, blind top-1 acc.
CV-Bench
47.3
60.3
+13.1 ∗
44.8
VSI-Bench (MC)
25.5
39.1
+13.6
34.0
VSI-Bench (NUM)
24.0
46.0
+21.9
–
VSI-Bench (MC+NUM)
24.7
42.7
+17.9
–
VideoMME
29.2
33.4
+4.1
46.6
Table 1 : Learnable non-visual shortcuts are large on template-based benchmarks and small elsewhere. A blind Qwen2-7B is scored zero-shot (ZS) and after k -fold LoRA fine-tuning on the other folds of the test set (TsT), with the same rule ( Section C.1 ): mean gold-option probability for MC and MRA for NUM questions, averaged over questions; VSI-Bench (MC+NUM) weights the two formats by question count. ΔTsT is TsT minus ZS, computed from unrounded scores. GPT-4o blind top-1 accuracy is shown for context and is not directly comparable with the Qwen columns.
VSI-Bench (Original)
VSI-Bench-Debiased
Model
Vis.
Blind
Gap
Vis.
Blind
Gap
LLaVA-Video-7B (base)
36.7
25.9
10.8
31.3
20.3
11.0
+ VSI-Train-10k FT
57.1
44.7
12.4
48.7
32.0
16.7
Gain from FT
+20.4
+18.8
+1.6
+17.4
+11.7
+5.7
Table 2 : VSI-Bench-Debiased (designer-in-the-loop pilot). LLaVA-Video-7B scores (MC accuracy and NUM MRA, averaged over questions) with vision (Vis.) and without (Blind), before and after fine-tuning on VSI-Train-10k; Gap is Vis. minus Blind. On the original set, fine-tuning raises vision and blind scores almost equally (+20.4 vs. +18.8); on the debiased set it raises vision much more than blind (+17.4 vs. +11.7), so more of the gain there requires vision.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
TsT-RF CV score
Benchmark
Uniform chance
Frequency baseline
Random folds
Grouped folds
CV-Bench
41.8
–
56.1
55.7
VSI-Bench
–
34.0
43.0
38.8
Appendix
Table 3 : TsT-RF cross-validation scores. Features are label-conditioned ( Section 3.3 ), so these are designer-side audit statistics. Baselines are for scale: uniform chance over each question’s options, averaged over questions, and VSI-Bench’s per-task frequency baseline [ 42 ] . Scores are averaged over questions (MC accuracy, plus NUM MRA for VSI-Bench). Grouped folds keep all questions about one scene (VSI-Bench) or image (CV-Bench) in the same fold. Features marked ‡ in Tables 4 and 5 are excluded. Dataset revisions: CV-Bench bc284db , VSI-Bench bc96b17 .
Task
Format
Feature Type
Feature Name
Description
Object Counting
NUM
Categorical
object
Object category
Numerical
obj_count
Count of this object
obj_val_log_mean
Mean of log gold values
obj_val_log_std
Std dev of log gold values
global_mean_log
Global mean (log space)
global_std_log
Global std dev (log space)
Appendix
Table 4 : Representative VSI-Bench TsT-RF features by question type. Features come from the question text and, for MC questions, the answer options; none uses visual input. ‡ Keyed on the held-out question’s gold answer; excluded from all reported TsT-RF scores and automated selections.
Task
Feature Type
Feature Name
Description
2D Count
Categorical
object
Object category
Numerical
n_options
Number of MC choices
obj_count
Count of this object
obj_val_log_mean
Mean of log gold values
obj_val_log_std
Std dev of log gold values
global_mean_log
Global mean (log space)
Appendix
Table 5 : Representative CV-Bench TsT-RF features by question type. All tasks are MC; no feature uses visual input. ‡ Keyed on the held-out question’s gold answer; excluded from all reported TsT-RF scores.
Subset
n
Vis.
Blind
Gap
Original
900
49.4
42.8
6.7
Δs(x) -ranked removal
800
47.9
40.4
7.5
Random removal
800
–
–
6.69±0.58
Random removal: mean ± s.d. over 1,000 seeds; range 4.62 to 8.75.
Appendix
Table 6 : Pruning MMMU by Δs(x) behaves like random removal. LLaVA-OneVision-7B accuracy averaged over questions; gaps use unrounded scores.
Question Type
Format
Original
Removed
Removed (%)
Kept
Object size
NUM
953
600
63.0
353
Absolute distance
NUM
834
400
48.0
434
Relative distance
MC
710
400
56.3
310
Appearance order
MC
618
300
48.5
318
Object counting
NUM
565
314
55.6
251
Relative direction (med.)
MC
378
324
85.7
54
Appendix
Table 7 : VSI-Bench-Debiased removes questions unevenly across types.
Model
Full gap
Pilot
Random
LLaVA-Video-7B + FT
12.4
16.7
12.4±0.6
Cambrian-S
19.2
26.0
19.4±0.8
Appendix
Table 8 : Random removal at the pilot’s rate barely changes the gap. Gap on the full set, on the debiased set, and after randomly removing 54% of questions (mean ± s.d. over 10 seeds).
Original
Debiased
Model
LLM backbone
V
B
Gap
V
B
Gap
Δ gap
Cambrian-S
Qwen2.5
66.1
46.9
19.2
58.3
32.3
26.0
+6.8
LLaVA-Video-7B + FT
Qwen2
57.1
44.7
12.4
48.7
32.0
16.7
+4.3
InternVL3-9B
InternLM3
43.6
32.5
11.1
41.5
28.2
13.2
+2.2
InternVL2.5-26B
InternLM2.5
45.3
32.5
12.8
38.4
25.1
13.3
+0.5
InternVL2.5-8B
InternLM2.5
44.0
31.6
12.4
37.5
24.2
13.3
+0.9
Appendix
Table 9 : The gap widens most for Cambrian-S and the fine-tuned LLaVA-Video-7B. V and B are vision and blind scores (MC accuracy and NUM MRA averaged over questions; NUM answers normalized for number words and units), sorted by original blind score. Gaps use unrounded scores. Cambrian-S is a pre-release checkpoint; base LLaVA-Video-7B comes from a separate evaluation, so its scores differ slightly from Table 2 .
Selection
Removed (%)
Kept
Gap
Original
0.0
5,130
12.4
Per-format, B=1000
19.5
4,130
12.4
Per-type, uniform
54.0
2,360
14.1
Per-type, pilot budgets (seeds 1, 42)
54.0
2,362
14.5, 14.7
Pilot
54.0
2,362
16.7
Appendix
Table 10 : Automated per-type IBP widens the gap less than the pilot, and per-format IBP leaves it unchanged. Gap of LLaVA-Video-7B fine-tuned on VSI-Train-10k, averaged over questions. Random removal at the same rates: 12.2 (per-format), 12.2 (per-type, uniform), 12.4 (per-type, pilot budgets).
Mean s(x)
B
Removed (%)
MC
NUM
200
3.9
0.337
0.421
500
9.7
0.308
0.385
1,000
19.5
0.270
0.329
1,500
29.2
0.246
0.270
2,000
39.0
0.219
0.210
Appendix
Table 11 : Mean s(x) falls steadily as the per-format IBP budget grows. Mean TsT-RF s(x) after pruning.
Question type
TsT-LLM
TsT-RF
Appearance order
15.5
40.1
Object counting
15.9
15.2
Object size
0.4
31.2
Absolute distance
0.2
12.2
Appendix
Table 12 : TsT-LLM and TsT-RF remove most question types at very different rates. Removal rates (%) by question type. Budgets are 200 for TsT-LLM and 1,000 for TsT-RF. At a matched budget of 200, Jaccard is 0.050.
Object
Mean (cm)
Std Dev (cm)
CV
Low-variance objects (easily exploitable)
Dishwasher
90.4
3.4
0.037
Bed
216.1
17.2
0.080
Washer
87.1
5.0
0.058
Kettle
23.8
1.5
0.062
Mouse
11.6
1.1
0.091
Appendix
Table 13 : Categories with near-constant sizes are easy to exploit. Size statistics for selected categories. CV is the coefficient of variation (standard deviation over mean).
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by model progress. This creates a need for difficult benchmarks that remain relevant for longer. We introduce ZeroBench - a lightweight visual reasoning benchmark curated using adversarial filtering to be "impossible" for frontier LMMs at its original release, with initial SotA scores of 0% pass@1 and pass^5. We track progress on ZeroBench over the subsequent year, observing SotA reaching 6% pass^5 and 19% pass@5, indicating the potential longevity of the benchmark. We evaluate 46 LMMs on ZeroBench, compare performance to a human baseline, analyse strengths and weaknesses, chart a year of progress in visual capabilities, and publicly release ZeroBench at https://zerobench.github.io.
Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31
Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underexplored. Existing VLM TTS methods largely require open-weight model access or expensive repeated sampling, and are evaluated primarily on multimodal mathematical and scientific reasoning benchmarks rather than general visual understanding tasks. In this paper, we propose Test-Time Hinting, a method that improves VLM performance via a single VLM call and requiring only black-box API access, which makes it broadly applicable to frontier closed-weight models. Our method is motivated by the observation that VLM errors tend to cluster around recurring failure patterns. We therefore train a lightweight hint generator model to predict, for a given test input, which "hint" should be prepended to the prompt, providing targeted contextual or procedural guidance that steers the VLM away from its characteristic failure modes. We show that Test-Time Hinting improves the accuracy of multiple closed-weight VLMs on natural-image VQA benchmarks and that these gains generalize to unseen benchmarks and VLMs without retraining the hint generator.
We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results. To this end, we propose MMGist, a curated benchmark that covers seven capability dimensions and contains 7,262 items. MMGist is constructed through a three-stage pipeline, which sequentially combines text-ablation filtering, cross-model saturation filtering, and anomaly detection filtering. We conduct extensive experiments on 27 leading LVLMs and compare MMGist with the raw pool of 23,250 items. The results show that MMGist preserves model rankings with high fidelity, with Spearman ρ=0.98, while reducing evaluation items by 69% and improving cross-model discrimination by 78%. Further results indicate that Visual Logic remains a systematic weakness of current LVLMs, while knowledge-intensive dimensions such as Expert Knowledge dimensions remain important factors for distinguishing closed-source models from open-source models. These findings suggest that high-quality evaluation should prioritize visual dependency, discriminative power, and reliability, rather than simply pursuing benchmark scale.
Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong +3
1Shanghai Jiao Tong University · 2Sichuan University