VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
Organizations: The University of Tokyo
Abstract
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.
Figures & tables
| Benchmark | #Size | #Refs | #VIs | Reference–VI Conflict | Metrics |
| Without visual instructions | |||||
| DreamBooth ( Ruiz et al., 2022 ) | 75 | 1 | – | ✗ | CLIP, DINO |
| OmniContext ( Wu et al., 2025b ) | 400 | 3 | – | ✗ | GPT (3 dim.) |
| DreamOmni2 ( Xia et al., 2025b ) | 319 | 4 | – | ✗ | Gemini, Doubao ( ByteDance, 2025 ) |
| MultiBanana ( Oshima et al., 2026b ) | 3,769 | 8 | – | ✗ | GPT, Gemini (5 dim.) |
| With visual instructions | |||||
| Model | Text Instruction Following | Reference Consistency | Visual Instruction Adherence | Visual Instruction Cleanliness | Scene Coherence | Visual Quality | Avg. |
| GPT-Image-1.5 | 6.79 | 8.12 | 4.57 | 9.39 | 7.84 | 8.88 | 7.60 |
| + multi-step | 6.53 | 7.29 | 4.43 | 9.80 | 8.24 | 8.89 | 7.53 |
| Nano Banana Pro | 6.07 | 8.27 | 5.67 | 6.27 | 7.01 | 8.48 | 6.96 |
| + multi-step | 6.85 | 7.37 | 5.09 | 8.37 | 8.03 | 8.69 | 7.40 |
| GPT-Image-1 | 6.51 | 7.84 | 4.35 | 9.69 | 7.99 | 8.90 | 7.55 |
| Nano Banana | 6.40 | 8.39 | 5.23 | 7.09 | 7.13 | 8.55 | 7.13 |
| Judge | Pearson | Spearman |
| GPT-5 | 0.78 | 0.75 |
| Gemini 2.5 | 0.74 | 0.71 |
| Qwen3-VL | 0.73 | 0.70 |
| Human | 0.80 | 0.78 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Text Instruction Following | Reference Consistency | Visual Instruction Adherence | Visual Instruction Cleanliness | Scene Coherence | Visual Quality | Avg. |
| GPT-Image-1.5 | 8.27 | 9.34 | 7.90 | 9.48 | 8.87 | 9.74 | 8.93 |
| + multi-step | 7.88 | 8.75 | 7.19 | 9.85 | 8.85 | 9.69 | 8.70 |
| Nano Banana Pro | 7.34 | 9.30 | 7.74 | 6.60 | 7.69 | 9.46 | 8.02 |
| + multi-step | 7.58 | 8.51 | 7.08 | 8.42 | 8.43 | 9.55 | 8.26 |
| GPT-Image-1 | 8.01 | 9.27 | 7.40 | 9.81 | 8.95 | 9.72 | 8.86 |
| Nano Banana | 7.62 | 9.48 | 7.76 | 7.46 | 7.93 | 9.52 | 8.29 |
| Model | IQ | IF | SF | Avg. |
| Show-o ( Xie et al., 2024 ) | 0.764 | 0.616 | 0.462 | 0.614 |
| OmniGen ( Xiao et al., 2024 ) | 0.730 | 0.532 | 0.438 | 0.567 |
| ACE ( Han et al., 2025 ) | 0.740 | 0.655 | 0.528 | 0.641 |
| ChatDiT ( Huang et al., 2024a ) | 0.811 | 0.713 | 0.574 | 0.699 |
| Claude + SD 2.1 ( Rombach et al., 2022 ) | 0.812 | 0.726 | 0.572 | 0.703 |
| Claude + SD 3 ( Esser et al., 2024 ) | 0.876 | 0.817 | 0.658 | 0.784 |
| Model | Text Instr. | Ref. Consist. | Vis. Adher. | Vis. Clean. | Scene | Quality | Avg. |
| GPT-Image-1 | 6.39 / 6.63 | 8.08 / 7.59 | 4.16 / 4.54 | 9.74 / 9.63 | 7.62 / 8.37 | 8.95 / 8.86 | 7.49 / 7.60 |
| GPT-Image-1.5 | 6.70 / 6.89 | 8.45 / 7.79 | 4.22 / 4.91 | 9.41 / 9.37 | 7.51 / 8.16 | 8.99 / 8.77 | 7.55 / 7.65 |
| Nano Banana Pro | 6.12 / 6.02 | 8.82 / 7.72 | 5.52 / 5.82 | 6.30 / 6.24 | 7.17 / 6.84 | 8.83 / 8.13 | 7.13 / 6.80 |
| Nano Banana | 6.42 / 6.38 | 8.84 / 7.93 | 4.93 / 5.54 | 7.13 / 7.05 | 7.11 / 7.15 | 8.81 / 8.28 | 7.21 / 7.06 |
| Qwen-Image-2511 | 2.79 / 3.23 | 3.42 / 3.34 | 2.55 / 2.44 | 7.82 / 7.73 | 5.02 / 6.22 | 7.41 / 7.26 | 4.84 / 5.04 |
| DreamOmni2 | 2.62 / 3.09 | 4.03 / 3.74 | 2.25 / 2.39 | 5.55 / 5.59 | 4.53 / 6.13 | 7.97 / 7.75 | 4.49 / 4.78 |
| Metric | Gemini–Human | GPT-5–Human | Qwen3-VL–Human | Human–Human |
| Text Instr. Follow. | 0.69 / 0.67 | 0.72 / 0.69 | 0.64 / 0.69 | 0.63 / 0.63 |
| Reference Consist. | 0.73 / 0.72 | 0.73 / 0.71 | 0.67 / 0.69 | 0.75 / 0.78 |
| Visual Instr. Adher. | 0.53 / 0.60 | 0.76 / 0.72 | 0.61 / 0.63 | 0.75 / 0.74 |
| Visual Instr. Clean. | 0.70 / 0.72 | 0.71 / 0.72 | 0.71 / 0.72 | 0.76 / 0.78 |
| Scene Coherence | 0.48 / 0.45 | 0.53 / 0.55 | 0.51 / 0.50 | 0.58 / 0.59 |
| Visual Quality | 0.49 / 0.57 | 0.57 / 0.63 | 0.49 / 0.59 | 0.57 / 0.60 |
| Score source | AP v1 PLCC | AP v1 SRCC | AP v2 PLCC | AP v2 SRCC | |
| GPT-5 Visual Quality | 168 | 0.566 | 0.506 | 0.680 | 0.598 |
| Gemini Visual Quality | 168 | 0.472 | 0.382 | 0.551 | 0.468 |
| Qwen3-VL-32B Visual Quality | 168 | 0.540 | 0.438 | 0.640 | 0.524 |
| Human Visual Quality | 168 | 0.415 | 0.380 | 0.497 | 0.435 |
| Geometric metric | PLCC | SRCC | |
| Bounding-box IoU | 168 | 0.357 | 0.435 |
| Center-in-region rate | 168 | 0.508 | 0.551 |
| Model | #Tasks | CLIP div. | LPIPS div. | Avg. Score | |
| GPT-Image-1.5 | 1 | 3 | 0.066 | 0.412 | 8.55 |
| GPT-Image-1.5 | 2 | 3 | 0.042 | 0.381 | 8.58 |
| GPT-Image-1.5 | 3 | 3 | 0.079 | 0.438 | 8.49 |
| GPT-Image-1.5 | 4 | 3 | 0.124 | 0.505 | 6.03 |
| Nano Banana Pro | 1 | 3 | 0.078 | 0.481 | 7.74 |
| Nano Banana Pro | 2 | 3 | 0.142 | 0.542 | 7.40 |
| Benchmark | #Size | #Refs | #VIs | Reference–VI Conflict | Metrics |
| Without visual instructions | |||||
| EditBench ( Wang et al., 2023 ) | 240 | 1 | – | ✗ | CLIP ( Radford et al., 2021 ) |
| EditVal ( Basu et al., 2023 ) | 648 | 1 | – | ✗ | CLIP, VLM, manual |
| EmuEdit ( Sheynin et al., 2024 ) | 3,055 | 1 | – | ✗ | L1, CLIP, DINO ( Caron et al., 2021 ) |
| MagicBrush ( Zhang et al., 2023a ) | 1,053 | 1 | – | ✗ | L1, L2, CLIP, DINO |
| AnyEdit ( Yu et al., 2025 ) | 1,250 | 1 | – | ✗ | L1, CLIP, DINO |