What Paired Evaluations Reveal under Visual Perturbations
Abstract
Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.
Figures & tables
| State | Response | Labels | ||
|---|---|---|---|---|
| Harmful flip | ||||
| Correction | ||||
| Persistent error | ||||
| Changed error | ||||
| Retained correctness |
| Target | Definition | Interpretation |
| Mean true-class probability decrease among inputs correct on both sides. | ||
| Mean decrease in maximum probability within the specified wrong-answer state. | ||
| Fraction of inputs with an accepted error after perturbation. | ||
| Fraction remaining correct but changing from accepted to rejected. | ||
| Fraction changing from an accepted correct prediction to an accepted error. |
| Target | Information used | Supported conclusion |
|---|---|---|
| Perturbed records | Exact accepted-error mass. | |
| Classwise overlaps | Sharp bounds equal those under every feasible five-state refinement. | |
| Five-state counts | Can tighten overlap-conditioned bounds. | |
| Shared correspondence across policies | Can certify an ordering throughout the feasible set despite overlapping cost intervals. |
| Question | Comparison | Result |
|---|---|---|
| Shared correspondence: more orderings? | Shared correspondence vs. separate intervals ( ) | 312/450 (69.3%) |
| Correctness overlap: more orderings? | 241/450 (53.6%) | |
| Five-state counts: more orderings? | 42/450 (9.3%) | |
| Minimax regret: lower realized cost? | Minimax regret vs. feasible mean ( ) | 4 lower / 3 higher 443 equal |
Appendix figures & tables49 assets
Supplementary material from the paper’s appendix.
Appendix
| Record | Clean | Perturbed |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 |
| Source | Cells | |||||
|---|---|---|---|---|---|---|
| Syn | 792 | 8.44 | 3.44 | 6.73 | 4.63 | 6.01 |
| Syn-B | 792 | 1.96 | 0.66 | 6.33 | 5.42 | 2.11 |
| ES | 2,376 | 1.56 | 0.79 | 5.52 | 4.44 | 3.52 |
| Diverse | 7,128 | 0.86 | 0.69 | 2.86 | 1.99 | 5.76 |
| Source | Cells | Extra | Narrower | |||
|---|---|---|---|---|---|---|
| Syn | 90 | 8.45 | 3.47 | 2.40 | 1.07 | 90/90 |
| Syn-B | 90 | 1.92 | 0.69 | 0.54 | 0.15 | 25/90 |
| ES | 135 | 1.12 | 0.77 | 0.68 | 0.09 | 38/135 |
| Diverse | 135 | 0.97 | 0.78 | 0.68 | 0.10 | 42/135 |
| Source | Original | Matched | Original | Matched | Original | Matched |
|---|---|---|---|---|---|---|
| Syn | 2.40 | 0.87 | 4.61 | 2.52 | 5.97 | 2.76 |
| Syn-B | 0.54 | 0.41 | 5.47 | 4.74 | 2.13 | 1.76 |
| ES | 0.68 | 0.54 | 3.03 | 2.38 | 5.11 | 3.90 |
| Diverse | 0.68 | 0.52 | 2.66 | 2.04 | 5.84 | 4.27 |
| Source | cells | Pair ind. | Mean | Mean | |
|---|---|---|---|---|---|
| Syn | 792 | +1.00 | 0.08 | 0.96 | 98.6% |
| Syn-B | 792 | +0.38 | 0.09 | 0.39 | 85.0% |
| ES | 2,374 | +0.52 | 0.21 | 0.67 | 86.3% |
| Diverse | 7,010 | +0.66 | 0.50 | 1.57 | 76.8% |
| Source | Tier | Cells | |||
|---|---|---|---|---|---|
| Syn | all tiers | 792 | 0.32 [0.28, 0.35] | 0.32 [0.27, 0.37] | 0.20 [0.16, 0.23] |
| Syn | mild | 792 | 0.32 [0.28, 0.35] | 0.32 [0.27, 0.37] | 0.20 [0.16, 0.23] |
| Syn-B | all tiers | 792 | 0.33 [0.29, 0.37] | 0.31 [0.25, 0.36] | 0.20 [0.17, 0.23] |
| Syn-B | mild | 792 | 0.33 [0.29, 0.37] | 0.31 [0.25, 0.36] | 0.20 [0.17, 0.23] |
| ES | all tiers | 2376 | 0.55 [0.53, 0.58] | 0.55 [0.50, 0.59] | 0.44 [0.41, 0.46] |
| ES | mild | 1936 | 0.47 [0.44, 0.50] | 0.47 [0.41, 0.51] | 0.34 [0.31, 0.36] |
| Source | Tier | width, classwise | width, global | Observed | Pinned (%) |
|---|---|---|---|---|---|
| Syn | all tiers | 13.78 | 23.31 | 13.87 | 48.4 |
| Syn | mild | 13.78 | 23.31 | 13.87 | 48.4 |
| Syn-B | all tiers | 6.44 | 11.32 | 9.33 | 83.9 |
| Syn-B | mild | 6.44 | 11.32 | 9.33 | 83.9 |
| ES | all tiers | 5.16 | 10.39 | 26.04 | 87.1 |
| ES | mild | 5.90 | 11.15 | 16.02 | 85.2 |
| Source | Cells | Shared | ||
|---|---|---|---|---|
| Syn | 90 | 82 | 77 | 28 |
| Syn-B | 90 | 60 | 75 | 6 |
| ES | 135 | 92 | 47 | 2 |
| Diverse | 135 | 78 | 42 | 6 |
| Total | 450 | 312 | 241 | 42 |
| Source | Cells | Feasible mean | Minimax | Better | Worse | Tied |
|---|---|---|---|---|---|---|
| Syn | 90 | 0.004778 | 0.004222 | 1 | 1 | 88 |
| Syn-B | 90 | 0.005556 | 0.005556 | 0 | 0 | 90 |
| ES | 135 | 0.012593 | 0.004444 | 3 | 0 | 132 |
| Diverse | 135 | 0.003704 | 0.005185 | 0 | 2 | 133 |
| (pp) | Better | Worse | Tied | |
|---|---|---|---|---|
| 0.0 | 0.000000 | 0 | 0 | 450 |
| 0.1 | 0.000044 | 1 | 0 | 449 |
| 0.2 | -0.000542 | 0 | 3 | 447 |
| 0.3 | -0.002400 | 2 | 5 | 443 |
| 0.4 | -0.003111 | 4 | 7 | 439 |
| 0.5 | 0.002111 | 4 | 3 | 443 |
| Target | Observed | |||
|---|---|---|---|---|
| [0.4, 0.8] | [0.4, 0.8] | [0.4, 0.6] | 0.6 | |
| [17.8, 26.2] | [19.0, 26.0] | [19.0, 26.0] | 23.2 | |
| [19.52, 30.02] | [20.58, 29.43] | [20.58, 29.43] | 26.24 |
| Source | Tier | ret | Class- | [95%] | ||||
|---|---|---|---|---|---|---|---|---|
| ES | mild | 0.778 | 0.854 | 0.768 | 0.823 | 0.580 | 0.855 | +0.076 [+0.070, +0.080] |
| ES | severe | 0.593 | 0.647 | 0.630 | 0.614 | 0.546 | 0.647 | +0.054 [+0.044, +0.063] |
| ES | extreme | 0.538 | 0.578 | 0.572 | 0.560 | 0.535 | 0.576 | +0.040 [+0.027, +0.056] |
| Diverse | mild | 0.678 | 0.801 | 0.781 | 0.755 | 0.587 | 0.809 | +0.123 [+0.110, +0.131] |
| Diverse | severe | 0.601 | 0.700 | 0.698 | 0.656 | 0.565 | 0.710 | +0.099 [+0.084, +0.109] |
| Diverse | extreme | 0.504 | 0.512 | 0.512 | 0.529 | 0.478 | 0.517 | +0.008 [-0.006, +0.024] |
| Source | Tier | Cells | |||
|---|---|---|---|---|---|
| ES | mild | 1935 | +0.076 [+0.060,+0.093] | -0.086 [-0.107,-0.065] | +0.077 [+0.060,+0.097] |
| ES | severe | 219 | +0.054 [+0.035,+0.070] | -0.016 [-0.028,-0.005] | +0.054 [+0.034,+0.072] |
| ES | extreme | 214 | +0.040 [+0.020,+0.057] | -0.005 [-0.020,+0.011] | +0.038 [+0.017,+0.056] |
| Diverse | mild | 2244 | +0.123 [+0.107,+0.142] | -0.020 [-0.031,-0.007] | +0.131 [+0.110,+0.153] |
| Diverse | severe | 1099 | +0.099 [+0.084,+0.116] | -0.002 [-0.013,+0.009] | +0.108 [+0.091,+0.128] |
| Diverse | extreme | 3174 | +0.008 [-0.013,+0.039] | +0.001 [-0.015,+0.017] | +0.013 [-0.013,+0.052] |
| Source | Tier | ret | random | ret | random | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ES | mild | 0.406 | 0.489 | 0.395 | 0.487 | 0.096 | 0.578 | 0.672 | 0.568 | 0.648 | 0.199 |
| ES | severe | 0.126 | 0.133 | 0.134 | 0.135 | 0.100 | 0.241 | 0.254 | 0.252 | 0.248 | 0.201 |
| ES | extreme | 0.104 | 0.106 | 0.106 | 0.106 | 0.100 | 0.208 | 0.210 | 0.209 | 0.208 | 0.200 |
| Diverse | mild | 0.208 | 0.260 | 0.271 | 0.271 | 0.099 | 0.357 | 0.449 | 0.452 | 0.441 | 0.200 |
| Diverse | severe | 0.130 | 0.144 | 0.150 | 0.147 | 0.100 | 0.248 | 0.274 | 0.280 | 0.270 | 0.199 |
| Source | Tier | [95%] | [95%] | [95%] | ||
|---|---|---|---|---|---|---|
| ES | mild | 0.578 | 0.672 | +0.094 [+0.078,+0.106] | +0.104 [+0.079,+0.140] | +0.024 [+0.013,+0.042] |
| ES | severe | 0.241 | 0.254 | +0.013 [+0.011,+0.017] | +0.002 [-0.001,+0.007] | +0.007 [+0.001,+0.016] |
| ES | extreme | 0.208 | 0.210 | +0.002 [+0.002,+0.003] | +0.001 [-0.000,+0.002] | +0.001 [+0.000,+0.004] |
| Diverse | mild | 0.357 | 0.449 | +0.092 [+0.075,+0.109] | -0.003 [-0.009,+0.007] | +0.008 [-0.006,+0.036] |
| Diverse | severe | 0.248 | 0.274 | +0.026 [+0.021,+0.032] | -0.006 [-0.008,-0.005] | +0.004 [-0.002,+0.016] |
| Diverse | extreme | 0.203 | 0.204 | +0.001 [+0.001,+0.002] | -0.000 [-0.001,-0.000] | +0.000 [-0.000,+0.001] |
| Source | Tier | Budget | [classes] | [classes] | [classes] |
|---|---|---|---|---|---|
| ES | mild | 0.1 | [+0.064,+0.111] | [+0.077,+0.116] | [-0.009,+0.013] |
| ES | mild | 0.2 | [+0.075,+0.114] | [+0.076,+0.125] | [+0.008,+0.038] |
| ES | mild | 0.3 | [+0.079,+0.120] | [+0.057,+0.115] | [+0.015,+0.048] |
| ES | severe | 0.1 | [+0.004,+0.011] | [-0.004,+0.004] | [-0.005,+0.001] |
| ES | severe | 0.2 | [+0.006,+0.020] | [-0.003,+0.006] | [+0.001,+0.011] |
| ES | severe | 0.3 | [+0.008,+0.024] | [-0.001,+0.009] | [+0.008,+0.021] |
| Source | Tier | attainable ceiling | / ceiling | ||||
|---|---|---|---|---|---|---|---|
| ES | mild | 0.167 | 0.748 | 0.883 | 0.941 | 0.676 | 0.760 |
| ES | severe | 0.675 | 0.160 | 0.319 | 0.478 | 0.864 | 0.830 |
| ES | extreme | 0.908 | 0.112 | 0.223 | 0.335 | 0.956 | 0.947 |
| Diverse | mild | 0.360 | 0.373 | 0.643 | 0.815 | 0.777 | 0.739 |
| Diverse | severe | 0.652 | 0.170 | 0.339 | 0.500 | 0.891 | 0.853 |
| Source | Tier | [groups] | |||||
|---|---|---|---|---|---|---|---|
| ES | mild | 0.778 | 0.854 | 0.832 | 0.854 | 0.851 | +0.054 [+0.049,+0.058] |
| ES | severe | 0.593 | 0.647 | 0.628 | 0.649 | 0.644 | +0.035 [+0.027,+0.044] |
| ES | extreme | 0.538 | 0.578 | 0.565 | 0.580 | 0.575 | +0.028 [+0.016,+0.045] |
| Diverse | mild | 0.678 | 0.801 | 0.757 | 0.805 | 0.794 | +0.079 [+0.069,+0.086] |
| Diverse | severe | 0.601 | 0.700 | 0.662 | 0.704 | 0.695 | +0.061 [+0.052,+0.069] |
| Diverse | extreme | 0.504 | 0.512 | 0.506 | 0.511 | 0.512 | +0.003 [-0.008,+0.016] |
| Source | Tier | ret | AUROC | Cov. | ||||
|---|---|---|---|---|---|---|---|---|
| ES | mild | 0.778 | 0.856 | 0.779 | 0.831 | 0.856 | +0.078 [+0.072,+0.083] | +0.111 [+0.097,+0.129] |
| ES | severe | 0.572 | 0.641 | 0.646 | 0.608 | 0.644 | +0.069 [+0.059,+0.075] | +0.012 [+0.010,+0.015] |
| ES | extreme | 0.504 | 0.560 | 0.577 | 0.553 | 0.563 | +0.056 [+0.034,+0.072] | +0.003 [+0.002,+0.003] |
| Diverse | mild | 0.673 | 0.801 | 0.790 | 0.754 | 0.815 | +0.128 [+0.119,+0.135] | +0.093 [+0.075,+0.120] |
| Diverse | severe | 0.588 | 0.701 | 0.709 | 0.663 | 0.718 | +0.113 [+0.096,+0.128] | +0.026 [+0.020,+0.036] |
| Diverse | extreme | 0.475 | 0.507 | 0.536 | 0.518 | 0.520 | +0.032 [+0.019,+0.049] | +0.001 [+0.001,+0.002] |
| Source | Tier (eval-defined) | Conditions | AUROC [95%] | Cov. [95%] | ||
|---|---|---|---|---|---|---|
| ES | mild | 44 | 0.778 | 0.856 | +0.078 [+0.072,+0.083] | +0.111 [+0.097,+0.129] |
| ES | severe | 5 | 0.572 | 0.641 | +0.069 [+0.059,+0.075] | +0.012 [+0.010,+0.015] |
| ES | extreme | 5 | 0.504 | 0.560 | +0.056 [+0.034,+0.072] | +0.003 [+0.002,+0.003] |
| Diverse | mild | 56 | 0.667 | 0.794 | +0.128 [+0.118,+0.135] | +0.088 [+0.071,+0.114] |
| Diverse | severe | 20 | 0.583 | 0.695 | +0.112 [+0.093,+0.127] | +0.023 [+0.018,+0.032] |
| Diverse | extreme | 86 | 0.475 | 0.507 | +0.032 [+0.019,+0.049] | +0.001 [+0.001,+0.002] |
| Source | Tier | Cells | Derangement | LOO mean | Gap [95%] | Wins (%) | ||
|---|---|---|---|---|---|---|---|---|
| ES | mild | 1936 | 0.775 | 0.853 | 0.553 | 0.548 | +0.299 [+0.275,+0.319] | 99.9 |
| ES | severe | 218 | 0.594 | 0.647 | 0.523 | 0.519 | +0.124 [+0.119,+0.130] | 99.1 |
| ES | extreme | 216 | 0.543 | 0.581 | 0.504 | 0.504 | +0.077 [+0.058,+0.103] | 80.5 |
| Diverse | mild | 2244 | 0.675 | 0.800 | 0.581 | 0.582 | +0.219 [+0.205,+0.230] | 100.0 |
| Diverse | severe | 1099 | 0.599 | 0.699 | 0.537 | 0.536 | +0.162 [+0.148,+0.171] | 99.2 |
| Diverse | extreme | 3485 | 0.498 | 0.508 | 0.461 | 0.453 | +0.047 [+0.038,+0.059] | 56.0 |
| Source | Tier | Class (cc. dev) | Class (all dev) | Class [95%] | Cov. Class [95%] | ||
|---|---|---|---|---|---|---|---|
| ES | mild | 0.778 | 0.854 | 0.597 | 0.590 | +0.257 [+0.238,+0.277] | +0.374 [+0.332,+0.422] |
| ES | severe | 0.593 | 0.646 | 0.557 | 0.543 | +0.089 [+0.080,+0.100] | +0.037 [+0.027,+0.052] |
| ES | extreme | 0.540 | 0.578 | 0.545 | 0.541 | +0.033 [+0.012,+0.050] | +0.006 [+0.005,+0.008] |
| Diverse | mild | 0.678 | 0.801 | 0.604 | 0.596 | +0.197 [+0.185,+0.206] | +0.177 [+0.148,+0.213] |
| Diverse | severe | 0.601 | 0.701 | 0.567 | 0.554 | +0.134 [+0.122,+0.143] | +0.051 [+0.041,+0.063] |
| Diverse | extreme | 0.502 | 0.510 | 0.449 | 0.419 | +0.061 [+0.036,+0.093] | +0.003 [+0.003,+0.004] |
| Source | Stratum | Derangement | Gap [95%] | Retention gap | ||
|---|---|---|---|---|---|---|
| ES | low | 0.683 | 0.771 | 0.517 | +0.255 [+0.235,+0.274] | +0.225 |
| ES | mid | 0.548 | 0.749 | 0.455 | +0.294 [+0.255,+0.332] | +0.100 |
| ES | high | 0.530 | 0.755 | 0.460 | +0.295 [+0.252,+0.325] | +0.075 |
| Diverse | low | 0.581 | 0.676 | 0.541 | +0.136 [+0.110,+0.156] | +0.150 |
| Diverse | mid | 0.527 | 0.633 | 0.490 | +0.144 [+0.117,+0.170] | +0.127 |
| Diverse | high | 0.517 | 0.646 | 0.507 | +0.138 [+0.122,+0.158] | +0.116 |
| Source | Subsets | Coverage [min,max] | [min] | AUROC [min,max] | ||
|---|---|---|---|---|---|---|
| ES | 1 | random | 18 | 0.597 [0.569,0.619] | +0.019 [-0.009] | 0.799 [0.785,0.816] |
| ES | 2 | random | 20 | 0.625 [0.609,0.638] | +0.047 [+0.030] | 0.818 [0.801,0.833] |
| ES | 3 | random | 20 | 0.640 [0.626,0.652] | +0.062 [+0.048] | 0.830 [0.818,0.844] |
| ES | 6 | random | 20 | 0.654 [0.641,0.666] | +0.076 [+0.063] | 0.840 [0.826,0.852] |
| ES | 9 | random | 20 | 0.663 [0.655,0.669] | +0.085 [+0.077] | 0.849 [0.837,0.854] |
| ES | 9 | severity 2 only | 1 | 0.658 | +0.080 | 0.837 |
| Source | Tier | native | ||||
|---|---|---|---|---|---|---|
| ES | mild | 0.1 | +0.057 [+0.046,+0.074] | +0.060 [+0.046,+0.082] | +0.054 [+0.043,+0.069] | +0.002 [-0.002,+0.007] |
| ES | mild | 0.2 | +0.076 [+0.062,+0.086] | +0.073 [+0.061,+0.082] | +0.072 [+0.060,+0.081] | +0.004 [+0.002,+0.007] |
| ES | mild | 0.3 | +0.078 [+0.062,+0.090] | +0.076 [+0.062,+0.088] | +0.076 [+0.061,+0.089] | +0.002 [+0.001,+0.003] |
| ES | severe | 0.1 | +0.007 [+0.006,+0.009] | +0.006 [+0.005,+0.007] | +0.007 [+0.005,+0.009] | +0.000 [-0.001,+0.001] |
| ES | severe | 0.2 | +0.012 [+0.009,+0.016] | +0.010 [+0.006,+0.013] | +0.012 [+0.009,+0.015] | +0.001 [+0.000,+0.002] |
| ES | severe | 0.3 | +0.017 [+0.012,+0.021] | +0.014 [+0.008,+0.019] | +0.017 [+0.011,+0.021] | +0.001 [-0.001,+0.002] |
| Implementation | Models | Pred. agreement (min) | Mean (max) | Clean acc. diff (pp, max) | Scale (range) | Within tolerances |
|---|---|---|---|---|---|---|
| DINOv2 | 1 | 1.0000 | 0.0000 | 0.00 | – | 1/1 |
| DINOv3 | 1 | 1.0000 | 0.0000 | 0.00 | – | 1/1 |
| OpenCLIP | 35 | 1.0000 | 0.0000 | 0.00 | 14.3–117.3 | 35/35 |
| timm | 6 | 1.0000 | 0.0000 | 0.00 | – | 6/6 |
| torchvision | 1 | 1.0000 | 0.0000 | 0.00 | – | 1/1 |
| Source | Tier | [groups] | [classes] | AUROC [groups] | ||||
|---|---|---|---|---|---|---|---|---|
| ES | mild | 0.561 | 0.671 | 0.583 | +0.088 [+0.076,+0.103] | [+0.058,+0.125] | +0.065 [+0.060,+0.069] | +0.085 [+0.073,+0.100] |
| ES | severe | 0.228 | 0.241 | 0.229 | +0.012 [+0.009,+0.015] | [+0.006,+0.018] | +0.065 [+0.055,+0.071] | +0.011 [+0.009,+0.015] |
| ES | extreme | 0.204 | 0.207 | 0.204 | +0.002 [+0.002,+0.003] | [+0.001,+0.004] | +0.057 [+0.037,+0.072] | +0.002 [+0.002,+0.003] |
| Diverse | mild | 0.334 | 0.428 | 0.345 | +0.082 [+0.066,+0.109] | [+0.069,+0.100] | +0.117 [+0.108,+0.123] | +0.080 [+0.064,+0.105] |
| Diverse | severe | 0.239 | 0.265 | 0.242 | +0.024 [+0.018,+0.033] | [+0.019,+0.030] | +0.104 [+0.087,+0.118] | +0.024 [+0.018,+0.033] |
| Diverse | extreme | 0.202 | 0.203 | 0.202 | +0.001 [+0.001,+0.002] | [+0.001,+0.002] | +0.035 [+0.021,+0.049] | +0.001 [+0.001,+0.002] |
| Source | Pool | Subsets | Coverage: mean [min, max] | Gain over : mean [min] | AUROC: mean [min, max] | |
|---|---|---|---|---|---|---|
| ES | 1 | 18 | 0.597 [0.569,0.619] | +0.019 [-0.009] | 0.799 [0.785,0.816] | |
| ES | 2 | 20 | 0.625 [0.609,0.638] | +0.047 [+0.030] | 0.818 [0.802,0.833] | |
| ES | 3 | 20 | 0.640 [0.626,0.652] | +0.062 [+0.048] | 0.830 [0.818,0.844] | |
| ES | 6 | 20 | 0.654 [0.641,0.666] | +0.076 [+0.063] | 0.840 [0.826,0.852] | |
| ES | 9 | 20 | 0.663 [0.655,0.669] | +0.085 [+0.077] | 0.849 [0.837,0.854] | |
| ES | 18 | 1 | 0.672 | +0.094 | 0.854 |
| Natural transforms | Corruptions | |||
| Scoring rule | ES | Diverse | ES | Diverse |
| Failure recall at 20% budget | ||||
| Probability mean | 0.596 | 0.366 | 0.672 | 0.449 |
| score | 0.588 | 0.360 | 0.654 | 0.447 |
| AUROC | ||||
| Probability mean | 0.791 | 0.688 | 0.854 | 0.801 |
| Source | A3 probes | Metric | Group interval | Class interval | |
|---|---|---|---|---|---|
| ES | Natural | Recall@.2 | +8.38 | [+7.15, +9.40] | [+6.72, +10.06] |
| ES | Natural | AUROC | +6.95 | [+6.30, +7.48] | [+5.51, +8.47] |
| ES | Corruption | Recall@.2 | +1.85 | [+0.73, +2.74] | [+0.69, +2.86] |
| ES | Corruption | AUROC | +1.77 | [+1.37, +2.07] | [+1.30, +2.31] |
| Diverse | Natural | Recall@.2 | +8.91 | [+7.41, +10.45] | [+7.06, +10.97] |
| Diverse | Natural | AUROC | +12.09 | [+10.67, +13.04] | [+10.56, +14.00] |
| Source | Budget | Difference | Group interval | ||
|---|---|---|---|---|---|
| ES | 10% | 0.489 | 0.479 | +0.93 | [-0.84, +2.13] |
| ES | 20% | 0.672 | 0.654 | +1.85 | [+0.73, +2.74] |
| ES | 30% | 0.771 | 0.758 | +1.29 | [+0.71, +2.16] |
| Diverse | 10% | 0.260 | 0.268 | -0.80 | [-1.37, -0.44] |
| Diverse | 20% | 0.449 | 0.447 | +0.23 | [-0.40, +1.18] |
| Diverse | 30% | 0.591 | 0.577 | +1.41 | [+0.59, +2.64] |
| Source | Probe set | Evaluations | Recall@.2 | AUROC | vs. TTA |
|---|---|---|---|---|---|
| ES | Clean | 1 | 0.578 | 0.778 | -0.018 |
| ES | Natural TTA | 18 | 0.596 | 0.791 | +0.000 |
| ES | Severity 2 | 9 | 0.658 | 0.837 | +0.062 |
| ES | Severity 4 | 9 | 0.669 | 0.854 | +0.073 |
| ES | All corruptions | 18 | 0.672 | 0.854 | +0.076 |
| Diverse | Clean | 1 | 0.357 | 0.678 | -0.009 |
| Coverage at | AUROC | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Source | Tier | [groups] | [classes] | [groups] | |||||
| ES | mild | 0.578 | 0.672 | 0.596 | +0.076 [+0.063,+0.086] | [+0.059,+0.088] | 0.854 | 0.791 | +0.063 [+0.058,+0.067] |
| ES | severe | 0.241 | 0.254 | 0.242 | +0.013 [+0.010,+0.016] | [+0.007,+0.018] | 0.647 | 0.596 | +0.051 [+0.041,+0.059] |
| ES | extreme | 0.208 | 0.210 | 0.208 | +0.002 [+0.002,+0.003] | [+0.001,+0.004] | 0.578 | 0.540 | +0.038 [+0.026,+0.052] |
| Diverse | mild | 0.357 | 0.449 | 0.366 | +0.083 [+0.066,+0.098] | [+0.066,+0.103] | 0.801 | 0.688 | +0.114 [+0.100,+0.123] |
| Diverse | severe | 0.248 | 0.274 | 0.250 | +0.023 [+0.019,+0.028] | [+0.018,+0.031] | 0.700 | 0.607 | +0.093 [+0.078,+0.104] |
| Source | Score | Recall@.2 | AUROC | score | Group CI | Class CI |
|---|---|---|---|---|---|---|
| ES | 0.672 | 0.854 | – | – | – | |
| ES | 0.578 | 0.778 | +0.094 | [+0.079,+0.106] | [+0.076,+0.115] | |
| ES | TTA | 0.596 | 0.791 | +0.076 | [+0.063,+0.086] | [+0.059,+0.088] |
| ES | Margin | 0.577 | 0.779 | +0.096 | [+0.079,+0.110] | [+0.073,+0.115] |
| ES | Entropy | 0.574 | 0.773 | +0.098 | [+0.079,+0.114] | [+0.081,+0.121] |
| ES | Agreement | 0.423 | 0.648 | +0.249 | [+0.208,+0.284] | [+0.212,+0.283] |
| Source | Group | Cells | ||
|---|---|---|---|---|
| Diverse | SSL | 204 | +0.085 | +0.074 |
| Diverse | DataComp | 204 | +0.107 | +0.093 |
| Diverse | DFN | 102 | +0.115 | +0.091 |
| Diverse | LAION | 612 | +0.093 | +0.091 |
| Diverse | MetaCLIP | 153 | +0.123 | +0.115 |
| Diverse | OpenAI | 255 | +0.072 | +0.064 |
| Sel. risk | Coverage | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Source | Protocol | raw | Platt | raw | Platt | raw | Platt | raw | Platt |
| ES | Q80 | 1.39 | 1.39 | 9.90 | 9.90 | 10.48 | 10.48 | 0.531 | 0.531 |
| ES | A0.9 | 0.65 | 0.26 | 9.48 | 10.81 | 7.92 | 5.44 | 0.484 | 0.476 |
| ES | A0.7 | 2.31 | 1.82 | 7.48 | 6.45 | 14.03 | 13.26 | 0.602 | 0.608 |
| ES | A0.5 | 5.39 | 5.27 | 4.20 | 3.26 | 19.78 | 20.60 | 0.703 | 0.712 |
| Diverse | Q80 | 2.05 | 2.05 | 8.16 | 8.16 | 18.95 | 18.95 | 0.186 | 0.186 |
| Model identifier | Group | Params. / M |
|---|---|---|
| DINOv2-vitg14-lc__meta | SSL | 1144 |
| DINOv3-ViT7B16-lc__meta | SSL | 6716 |
| EVA02-L-14__merged2b_s4b_b131k | other-clip | 428 |
| MAE-ViT-B-ft__meta | SSL | 87 |
| MAE-ViT-L-ft__meta | SSL | 304 |
| PE-Core-L-14-336__meta | pe-core | 671 |
| Set / step | Features |
|---|---|
| Clean accuracy on synthetic base images and real development base images | |
| , synthetic , mean on each side, mean among each side’s correct images, perturbed clipped log loss, and target-matched synthetic risk | |
| plus independent counterparts of | |
| C1 | Add to |
| C2 | Add to C1 |
| C3a | Add to C2 (primary hypothesis) |
| Target source | Target | Baseline | C1 gain | C1 interval | gain |
|---|---|---|---|---|---|
| ES | 2.88 | -0.12 | [-0.52, +0.15] | +0.11 | |
| ES | 3.26 | +0.22 | [+0.12, +0.32] | +0.24 | |
| ES | 0.61 | +0.00 | [-0.09, +0.07] | -0.02 | |
| ES | 2.42 | -0.06 | [-0.15, -0.00] | -0.00 | |
| Diverse | 3.35 | +1.10 | [+0.40, +2.12] | +0.01 | |
| Diverse | 2.60 | +0.38 | [-0.22, +1.29] | +0.02 |
| Features | Source | Target | Original | ||
|---|---|---|---|---|---|
| Syn | ES | -0.12 [-0.52, +0.15] | +0.52 [-0.43, +2.56] | +0.11 [-0.00, +0.27] | |
| Syn | ES | +0.22 [+0.12, +0.32] | +0.05 [-0.07, +0.22] | +0.24 [-0.00, +0.70] | |
| Syn | ES | +0.00 [-0.09, +0.07] | -0.04 [-0.11, +0.00] | -0.02 [-0.05, +0.00] | |
| Syn | ES | -0.06 [-0.15, -0.00] | -0.55 [-1.88, +0.05] | -0.00 [-0.06, +0.04] | |
| Syn | Diverse | +1.10 [+0.40, +2.12] | +0.02 [+0.00, +0.05] | +0.01 [-0.19, +0.13] | |
| Syn | Diverse | +0.38 [-0.22, +1.29] | -0.03 [-0.11, +0.01] | +0.02 [-0.04, +0.07] |
| Features | Source | Target | Baseline | C1 | Gain (%) | Deletion gain (pp) | Positive |
|---|---|---|---|---|---|---|---|
| Syn | ES | 2.88 | 3.00 | -4.1 | [-0.20, +0.19] | 6/10 | |
| Syn | ES | 3.26 | 3.04 | +6.7 | [-0.24, +0.34] | 7/10 | |
| Syn | ES | 0.61 | 0.61 | +0.0 | [-0.01, +0.39] | 6/10 | |
| Syn | ES | 2.42 | 2.48 | -2.5 | [-0.16, +0.06] | 2/10 | |
| Syn | Diverse | 3.35 | 2.25 | +33.0 | [+0.50, +1.55] | 10/10 | |
| Syn | Diverse | 2.60 | 2.22 | +14.5 | [+0.30, +1.67] | 10/10 |
| Target | Source | MAE | MAE | Gain [95% interval] |
|---|---|---|---|---|
| ES | 4.87 | 2.69 | +2.18 [+0.68, +4.40] | |
| ES | 5.32 | 2.68 | +2.65 [+1.19, +5.01] | |
| DIV | 6.62 | 3.03 | +3.58 [+1.54, +6.87] | |
| DIV | 8.07 | 3.96 | +4.11 [+1.59, +8.32] |
| Domain | Matched images | Original acc. | ReaL acc. | (pp) | Original | ReaL |
|---|---|---|---|---|---|---|
| ES | 473 | 66.88 | 63.33 | -3.55 | 3.99 | 8.01 |
| Diverse | 473 | 26.11 | 24.96 | -1.16 | 1.16 | 2.77 |
| Source | Labels | moved (%) | moved (%) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ES | original | 24.17 | 1.55 | 65.33 | 3.99 | 4.96 | 1.40 | 9.93 | – | – |
| ES | ReaL | 22.02 | 1.30 | 62.03 | 8.01 | 6.64 | 1.27 | 9.38 | 18.4 | 16.0 |
| Diverse | original | 64.00 | 0.62 | 25.49 | 1.16 | 8.73 | 2.08 | 8.31 | – | – |
| Diverse | ReaL | 59.65 | 0.56 | 24.40 | 2.77 | 12.63 | 1.95 | 7.87 | 10.2 | 18.6 |
| Domain | High (%) | Low (%) | High events | Low events |
|---|---|---|---|---|
| ES | 61.45 | 27.10 | 16416 | 79704 |
| Diverse | 18.69 | 6.40 | 49248 | 239112 |