VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations
Organizations: University of Waterloo · National Taiwan University · Nanyang Technological University · NVIDIA
Abstract
Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.
Figures & tables
| Training sources | Additional datasets | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RichHF | PAL4VST | AbHuman | HAD | SynthScars | SDG-30K | |||||||
| Method | IoU | IoU | IoU | IoU | IoU | IoU | ||||||
| Qwen3-VL-8B | 0.022 | 0.137 | 0.026 | 0.124 | 0.024 | 0.108 | 0.044 | 0.111 | 0.026 | 0.117 | 0.114 | 0.269 |
| GPT-5.6-terra † | 0.096 | 0.218 | 0.118 | 0.239 | 0.171 | 0.217 | 0.148 | 0.251 | 0.162 | 0.281 | 0.058 | 0.217 |
| GPT-5.6-sol † | 0.156 | 0.274 | 0.141 | 0.248 | 0.217 | 0.282 | 0.211 | 0.298 | 0.186 | 0.290 | 0.078 | 0.258 |
| Gemini-3-Flash † | 0.047 | 0.082 | 0.048 | 0.068 | 0.057 | 0.079 | 0.075 | 0.119 | 0.062 | 0.101 | 0.031 | 0.068 |
| Method | All | RichHF | EvalMuse | TIG | TIE | SRIG | SRIE | MRIG | MRIE |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B judge ‡ | 0.373 | 0.456 | 0.473 | 0.570 | 0.066 | 0.128 | 0.393 | 0.537 | |
| GPT-5.6-terra † | 0.402 | 0.440 | 0.567 | 0.640 | 0.493 | 0.258 | 0.110 | 0.523 | 0.539 |
| GPT-5.6-sol † | 0.437 | 0.470 | 0.655 | 0.623 | 0.590 | 0.275 | 0.253 | 0.539 | 0.684 |
| Gemini-3-Flash † | 0.491 | 0.515 | 0.535 | 0.531 | 0.686 | 0.356 | 0.404 | 0.418 | 0.677 |
| VIEScore2 | 0.601 | 0.692 | 0.803 | 0.589 | 0.429 | 0.375 | 0.403 | 0.509 | 0.417 |
| Model | Grid IoU | PQ SRCC | SC SRCC | |
|---|---|---|---|---|
| Qwen3-VL-8B | 0.070 | 0.273 | 0.181 | 0.352 |
| GPT-5.6-terra † | 0.112 | 0.247 | 0.357 | 0.508 |
| VIEScore2 | 0.324 | 0.506 | 0.558 | 0.564 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Scores | Localization | Raw supervision |
|---|---|---|---|
| ImagenWorld | PQ, SC | single grid | ratings, masks |
| RichHF | PQ, SC | artifact and misalignment grids | ratings, heatmaps |
| PAL4VST | – | single grid | masks |
| EvalMuse | overall | – | alignment ratings |
| COCO | PQ, SC | single grid | clean constraint |
| Source | Train | Validation | Primary eval. | Conditioning images |
|---|---|---|---|---|
| RichHF | 13,920 | 1,518 | 400 | no |
| EvalMuse | 9,000 | 1,000 | 200 | no |
| PAL4VST | 8,177 | 908 | 300 | no |
| ImagenWorld editing/reference | 3,642 | 415 | 250 | yes |
| ImagenWorld TIG | 1,165 | 121 | 50 | no |
| COCO | 2,218 | 272 | 100 | no |
| Primary source | Total | Localization | Overall score | PQ/SC |
|---|---|---|---|---|
| RichHF | 400 | 400 | 400 | 400 |
| PAL4VST | 300 | 300 | 0 | 0 |
| EvalMuse | 200 | 0 | 200 | 0 |
| ImagenWorld (six tasks) | 300 | 300 | 300 | 300 |
| COCO | 100 | 100 | 0 | 0 |
| Total | 1,300 | 1,100 | 900 | 700 |
| Dataset | Evaluation subset | |
|---|---|---|
| PAL4VST | Full test set | 1,405 |
| RichHF | Test split after threshold tuning | 840 |
| EvalMuse | External score subset | 2,000 |
| AbHuman | K validation subset | 1,407 |
| HAD | Validation subset | 709 |
| SynthScars | Official test split ( images) | 951 |
| Method | Spatial output | Threshold | Setting source |
|---|---|---|---|
| RAHF | Two heatmaps | 0.06 | RichHF ( ) |
| ImageDoctor | Two heatmaps | 0.03 | RichHF ( ) |
| PAL | Artifact mask | 0.5 | Released default |
| SegFormer-b0 | Defect probabilities | 0.75 | Validation |
| LEGION | Artifact mask | 0.5 cell fraction | Released rule |
| SDG | Defect boxes | – | Released output |
| Method | Scores | Spatial output | Text feedback | GRPO | Cond. |
|---|---|---|---|---|---|
| VIEScore ( Ku et al., 2024a ) | ✓ | – | Rationale | – | ✓ |
| RAHF ( Liang et al., 2024 ) | ✓ | Heatmaps | – | – | – |
| ImageDoctor ( Guo et al., 2026 ) | ✓ | Heatmap | Reasoning | ✓ | – |
| PAL ( Zhang et al., 2023 ) | – | Mask | – | – | – |
| LEGION ( Kang et al., 2025 ) | – | Mask | Explanation | – | – |
| FGA-BLIP2 ( Han et al., 2026 ) | ✓ | – | – | – | – |
| Model | PQ SRCC | SC SRCC | ||||
|---|---|---|---|---|---|---|
| Qwen3-VL-8B | 0.211 | 0.388 | 0.273 | 0.070 | 0.181 | 0.352 |
| GPT-5.6-terra † | 0.257 | 0.239 | 0.247 | 0.112 | 0.357 | 0.508 |
| VIEScore2 ( ) | 0.415 | 0.545 | 0.471 | 0.296 | 0.580 | 0.532 |
| VIEScore2 | 0.466 | 0.552 | 0.506 | 0.324 | 0.558 | 0.564 |
| SegFormer-b0 (matched loc.) | 0.433 | 0.573 | 0.493 | 0.272 | – | – |
| Task/source | Task/source | ||||
|---|---|---|---|---|---|
| TIG | 0.572 | 0.589 | MRIG | 0.545 | 0.509 |
| TIE | 0.692 | 0.429 | MRIE | 0.496 | 0.417 |
| SRIG | 0.512 | 0.375 | RichHF | 0.465 | 0.692 |
| SRIE | 0.412 | 0.403 | PAL4VST | 0.376 | – |
| EvalMuse | – | 0.803 |
| Method | All | RichHF | EvalMuse | TIG | TIE | SRIG | SRIE | MRIG | MRIE |
|---|---|---|---|---|---|---|---|---|---|
| ImageReward | 0.239 | 0.227 | 0.595 | 0.107 | 0.195 | 0.151 | 0.035 | 0.263 | |
| PickScore | 0.233 | 0.263 | 0.441 | 0.412 | 0.288 | 0.184 | 0.391 | 0.407 | |
| HPSv2 | 0.166 | 0.182 | 0.473 | 0.209 | 0.034 | 0.181 | 0.177 | 0.486 | 0.406 |
| VQAScore | 0.224 | 0.238 | 0.440 | 0.110 | 0.081 | 0.145 | 0.085 | 0.105 | 0.140 |
| FGA-BLIP2 | 0.423 | 0.452 | 0.918 | 0.422 | 0.255 | 0.151 | 0.233 | ||
| ImageDoctor | 0.411 | 0.714 | 0.280 | 0.459 | 0.160 | 0.271 | 0.355 | 0.479 | 0.299 |
| Task | ||||||
|---|---|---|---|---|---|---|
| TIE | 50 | 0.429 | 0.403 | 0.692 | 0.683 | |
| SRIG | 50 | 0.375 | 0.378 | 0.512 | 0.503 | |
| SRIE | 50 | 0.403 | 0.265 | 0.412 | 0.370 | |
| MRIG | 50 | 0.509 | 0.470 | 0.545 | 0.522 | |
| MRIE | 50 | 0.417 | 0.509 | 0.496 | 0.458 | |
| All | 250 | 0.451 | 0.420 | 0.536 | 0.514 |
| EvalMuse slice | VIEScore2 SRCC | FGA-BLIP2 SRCC | |
|---|---|---|---|
| Primary suite | 200 | 0.803 | 0.918 |
| External 2K | 2,000 | 0.773 | 0.910 |
| Method | Input | T2I | Edit |
|---|---|---|---|
| ImageReward | 0.530 | 0.559 | |
| PickScore | 0.574 | 0.574 | |
| HPSv2 | 0.546 | 0.541 | |
| VQAScore | 0.537 | 0.554 | |
| ImageDoctor | 0.546 | 0.525 | |
| Q-Align | 0.513 | 0.523 |
| Variant | Grid IoU | Overall SRCC | |
|---|---|---|---|
| VIEScore2 ( ) | 0.296 | 0.471 | 0.596 |
| VIEScore2 ( ) | 0.320 | 0.481 | 0.601 |
| VIEScore2 ( ) | 0.323 | 0.496 | 0.596 |
| VIEScore2 | 0.324 | 0.506 | 0.601 |
| Run | ||||
|---|---|---|---|---|
| VIEScore2 ( ) (epoch 4) | 0.471 | 0.296 | – | – |
| VIEScore2 ( ) (epoch 5) | 0.472 | 0.297 | ||
| VIEScore2 (seed 42) | 0.506 | 0.324 | ||
| VIEScore2 (seed 2) | 0.483 | 0.320 | ||
| VIEScore2 (seed 3) | 0.493 | 0.325 |
| Policy | |||||
|---|---|---|---|---|---|
| VIEScore2 ( ) | 0.374 | 0.548 | 0.445 | 0.303 | 14 |
| GRPO | 0.452 | 0.450 | 0.451 | 0.300 | 19 |
| GRPO | 0.411 | 0.534 | 0.464 | 0.322 | 26 |
| GRPO | 0.334 | 0.695 | 0.451 | 0.310 | 27 |
| GRPO Tversky | 0.446 | 0.494 | 0.469 | 0.313 | 21 |
| Model | Artifact | Misalign. | Coverage | COCO FP/100 | PAL FP/115 |
|---|---|---|---|---|---|
| VIEScore2 ( ) | 0.488 | 0.165 | 0.330 | 0 | 63 |
| VIEScore2 | 0.493 | 0.198 | 0.297 | 0 | 85 |
| SegFormer-b0 | – | – | – | 34 | 62 |
| Output form | ||||
|---|---|---|---|---|
| Dense bitmap | 0.580 | 0.099 | 0.169 | 0.116 |
| Point list | 0.250 | 0.446 | 0.320 | 0.206 |
| Bounding boxes | 0.354 | 0.671 | 0.463 | 0.280 |
| Sparse cells | 0.298 | 0.596 | 0.397 | 0.246 |
| Representation ceiling (model-free) | Serialized target | Trained model | |||||
| pixel-F1 | pixel-IoU | GT vanished | tokens (mean) | tokens (p95) | F1@ | pixel-IoU p | |
| 4 | 0.219 | 0.178 | 68.4% | 7 | 35 | — | — |
| 8 | 0.477 | 0.374 | 24.1% | 27 | 116 | 0.425 | 0.192 |
| 12 | 0.657 | 0.535 | 9.8% | 66 | 250 | 0.398 | 0.235 |
| 16 | 0.731 | 0.615 | 6.4% | 120 | 471 | 0.415 | 0.231 |
| 24 | 0.813 | 0.712 | 2.8% | 273 | 1,089 | 0.388 | 0.209 |