cs.CVSep 28, 2026

What Paired Evaluations Reveal under Visual Perturbations

Authors: Yongda Wei, Chen Zhang, Yifei Wang, Xinyu Wang, Bosen Shao, Hanxi Li, Liping Di

Abstract

Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.

Figures & tables

Appendix figures & tables49 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models

    Nov 24, 2025Xiangjie Sui, Songyang Li, Hanwei Zhu +3Overconfidence

  2. Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

    Apr 23, 2026Mohammed Safi Ur Rahman Khan, Sanjay Suryanarayanan, Tushar Anand +1Large Language Model EvaluationBlind Spots

  3. How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

    Jun 24, 2026Yuxing Cheng, Yuan Wu, Yi ChangLingdt-Vl-OcrOptical Character Recognition