With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.
Figures & tables
Figure 1: Overview of SQUARE-Bench. SQUARE-Bench possesses four key characteristics, including (1) diverse visual categories, (2) fine-grained taxonomy, (3) multi-Format tasks, and (4) a dual-answer design. Specifically, Answer 1 (Visual GT) is used to evaluate the LMM’s ability to interpret actual visual content, while Answer 2 (Intended GT) serves as a baseline to evaluate the T2I model’s ability to generate images that match the intended prompt.
Dataset
Visual Format
Taxonomy Focus
Annotator
Evaluation Dimensions
Dual-Answer Mechanism
Semantics
Quality
Authenticity
Responsibility
Q-Bench +
Image
LMM Capability
Expert
✗
✓
✗
✗
✗
A-Bench
Image
LMM Capability
Expert
✓
✓
✗
✗
✗
FakeBench
Image
Question Type
LMM + Expert
✗
✗
✓
✗
✗
LOKI
Mixed
Visual Format
LMM + Expert
✗
✗
✓
✗
✗
FakeClue
Image
Image Category
Multi-LMMs
✗
✗
✓
✗
✗
Table 1: Comparison of SQUARE-Bench with existing LMM Benchmarks.
Figure 2: Distribution of fine-grained dimensions and image counts across the four aspects of SQUARE-Bench: (a) Semantics, (b) Quality, (c) Authenticity, and (d) Responsibility.
Figure 3: Distributions of AIGI scores used for data curation: (a) Technical Quality scores sourced from AIGIQA-20K ( Li et al., 2024b ) ; (b) Aesthetic Quality scores predicted by Q-Align ( Wu et al., 2023 ) ; and (c) the predicted RichHF ( Qian et al., 2025 ) score distribution specifically for the Authenticity aspect.
Figure 4: Sampled SQUARE-Bench examples from four aspects.
Table 2: Benchmark results on the SQUARE-Bench semantics aspect.
Table 3: Benchmark results on the SQUARE-Bench quality aspect.
Table 4: Benchmark results on the SQUARE-Bench authenticity aspect.
Figure 5: A Quick Look at the SQUARE-Bench outcomes. (a) showcases a comparative analysis of the overall accuracy between human performance, selected LMMs and random guess . (b) displays a radar chart detailing the performance distribution across the four fundamental dimensions.
Table 5: Benchmark results on the SQUARE-Bench responsibility aspect. - indicates that the model does not support the multi-image inputs required for these sub-dimensions.
Figure 6: Overview of the proposed SQUARE-Bench-guided iterative editing framework. The framework operates as an assessment-instruction-editing loop for post-generation correction: a dimension-specific LMM evaluates the current image and generates an editing instruction, while a fixed image editor applies the correction for subsequent reassessment. Only the guidance LMMs’ LoRA adapters are trained on SQUARE-Bench; Qwen3-VL-8B handles semantics, while Ovis2.5-9B handles quality, authenticity, and responsibility.
Metric
Orig.
Q
Q+G
S
S+G
Semantic
CLIP ↑
0.2791
0.2889
0.2883
0.2842
0.2819
BLIP ↑
0.5209
0.5430
0.5445
0.5362
0.5343
LMM4LMM ↑
0.5030
0.5646
0.5703
0.5357
0.5428
RichHF ↑
0.6202
0.6155
0.6267
0.6248
0.6168
Qwen3-32B ↑
5.8453
7.8885
8.1619
6.8813
7.3094
Table 6: Single-pass versus LMM-guided iterative editing. Results compare original images, single-pass editors, and their guided variants across four evaluation aspects. Arrows indicate the preferred direction; bold denotes the best mean; shaded columns denote guided results.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Source Dataset
Prompt
Caption
Real Image
AIGIs
Sampled Size
Sampled Size
Sampled Size
Sampled Size
AIGIs Evaluation
PRISM-Bench ( Fang et al., 2025 )
400
0
0
0
EvalMi-50K ( Wang et al., 2025b )
1305
0
0
0
WISE ( Niu et al., 2025 )
1000
0
0
0
AIGIQA-20k ( Li et al., 2024b )
0
0
0
3000
Synthetic Data Detection
DFbench ( Wang et al., 2025a )
0
500
500
0
Appendix
Table 7: Overview of 16 Diverse Source Datasets in The SQUARE-Bench
Figure 7: Overview of AIGIs from semantics dimension.
Figure 8: Overview of AIGIs from quality dimension.
Figure 9: Overview of AIGIs from authenticity dimension.
Figure 10: Overview of AIGIs from responsibility dimension.
Figure 11: User Interface demonstrating the Fine-Grained Dimension Alignment process.
Figure 12: Illustration of the manual QA authoring interface. Experts are shown the assigned sub-dimension and record an instance-specific question, candidate options, and the corresponding Visual and Intended Ground Truth answers.
Figure 13: Dataset Statistics of SQUARE-Bench. (a) Distribution of distinct question types across evaluation dimensions. This illustrates not only the diversity of inquiry formats but also the adaptive alignment between question types and visual attributes, ensuring that the interrogation method is tailored to the specific dimension rather than applying rigid templates. (b) Word cloud visualization sampled from the entire question corpus, showcasing the semantic richness and the comprehensive coverage of visual concepts across the benchmark.
Figure 14: Illustration of the interface for the user-study.
Figure 15: Recorded trajectories from our LMM-guided editing experiments with Qwen-Image-Edit-2511 (top) and Step1X-Edit-v1p2 (bottom). Each row shows the original image followed by five consecutive editing outputs.
The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks' own prompt sets. When compared to existing evaluation metrics such as CLIPScore, VIEScore, and VQAScore, MT2IE's T2I model rankings are more faithful and far more consistent across multiple evaluation seeds when using the same number of prompts. MT2IE can also adapt evaluation to the model being tested: rewriting each prompt based on the model's own measured performance to produce a bespoke per-model benchmark that still recovers the official rankings and keeps the evaluated model in an informative scoring range. We hope that these results will encourage the development of dynamic and interactive evaluation frameworks, and mitigate the deprecation of automatic evaluation benchmarks.
Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible but semantically incorrect generations. In this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited human-annotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides fine-grained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.
Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3
Indian Institute of Science · VCL, IISc · Hugging Face
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.