cs.CVNov 6, 2025

Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts

Authors: Ellis Brown, Jihan Yang, Shusheng Yang, Rob Fergus, Saining Xie

Organizations: New York University

Abstract

Multimodal LLMs can answer many questions in vision-centric benchmarks without looking at the image, using linguistic priors and skewed answer distributions. Blind evaluation shows what a model already answers this way, but not what it could learn from the test set's own regularities. Extending partial-input auditing (e.g., hypothesis-only baselines in natural language inference), we argue that benchmark designers should "train on the test set": probe the artifact they release for exploitable patterns. Our Test-set Stress-Test (TsT) cross-validates a text-only Qwen2-7B on the test set's questions and answer options, yielding a benchmark-level score and a per-question bias score s(x); a random forest on hand-crafted features adds a fast, interpretable audit. On the template-based VSI-Bench and CV-Bench, the held-out score is 17.9 and 13.1 points above the model's own zero-shot score. On MMMU it learns little, even though GPT-4o correctly answers 52.3% of its multiple-choice questions without the image, so blind success there reflects pretrained knowledge rather than learnable test-set patterns. Iterative Bias Pruning (IBP) removes the questions with the highest s(x) and re-diagnoses; on VSI-Bench it widens a fine-tuned model's vision-blind gap more than random removal at the same rate. We also release VSI-Bench-Debiased, which lowers a fine-tuned model's blind score from 44.7 to 32.0 while its vision score falls only from 57.1 to 48.7.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    Feb 13, 2025Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31Visual ReasoningVLM Reasoning

  2. Test-Time Hinting for Black-Box Vision-Language Models

    May 13, 2026Kaihua Hou, Abhijith Varma Mudunuri, Jiaxing Qiu +3Visual Question AnsweringVLM Adaptation

  3. MMGist: A Comprehensive Multimodal Benchmark for 2027

    Jun 21, 2026Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong +3VLM EvaluationLarge Vision-Language Models