Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.
Figures & tables
Figure 1: Overview of PortraitAes-Bench. Each portrait is routed to one of nine photographic intents and evaluated by six specialists spanning 22 secondary criteria, yielding structured routing, rating, and intent-conditioned fusion targets.
Figure 2: Construction of PortraitAes-Bench. Expert rubrics define an intent router and six specialist roles; schema checks, Qwen3.5 Flash screening, and expert review precede intent-weighted fusion.
Figure 3: Intent-conditioned overall-score distributions. The standard benchmark spans a broad range, whereas the hard-case set concentrates visually similar portraits for fine-grained ranking.
Figure 4: Illustration of the numerical rewards. Gaussian calibration anchors predictions to rubric-derived scores; same-dimension cross-group ranking penalizes ordinal inversions. Images and rollout values illustrate the mechanism rather than report an evaluation case.
Model
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
Qwen3.5-VL-9B
0.833
0.822
0.724
0.742
0.741
0.698
0.807
0.789
0.706
0.714
0.698
0.740
0.857
0.874
InternVL3-8B
0.393
0.453
0.518
0.505
0.581
0.623
0.530
0.570
0.448
0.511
0.297
0.296
0.607
0.668
PortraitCraft-4B †
0.698
0.801
0.647
0.760
0.738
0.773
0.578
0.698
0.613
0.754
0.535
0.655
0.652
0.785
Qwen3.5 Flash
0.791
0.837
0.620
0.700
0.807
0.798
0.795
0.818
0.590
0.705
0.667
0.723
0.804
0.763
ArtiMuse †
0.741
0.785
0.744
0.796
0.556
0.602
0.688
0.717
0.725
0.756
0.712
0.703
0.785
0.804
Table 1: Standard PortraitAes-Bench results (PLCC/SRCC; higher is better). † denotes diagnostic correlations from a holistic-only scorer.
Model
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
Qwen3.5-VL-9B
0.515
0.513
0.346
0.281
0.558
0.506
0.602
0.574
0.385
0.474
0.391
0.413
0.649
0.629
InternVL3-8B
0.030
0.036
0.156
0.112
0.402
0.476
0.228
0.273
-0.010
-0.033
0.121
0.184
0.210
0.215
PortraitCraft-4B †
0.406
0.403
0.361
0.331
0.479
0.463
0.151
0.145
0.342
0.328
0.288
0.297
0.406
0.400
Qwen3.5 Flash
0.522
0.522
0.240
0.214
0.578
0.543
0.624
0.601
0.304
0.411
0.407
0.400
0.615
0.635
ArtiMuse †
0.430
0.421
0.407
0.398
0.336
0.346
0.368
0.350
0.410
0.418
0.387
0.383
0.501
0.495
Table 2: PortraitAes-Bench hard-case results (PLCC/SRCC; higher is better). † denotes diagnostic correlations from a holistic-only scorer.
Method
Training component
PLCC/SRCC
SFT
Gaussian
Cross-BPR
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
Qwen3.5-VL-9B
0.833/0.822
0.724/0.742
0.741/0.698
0.807/0.789
0.706/0.714
0.698/0.740
0.857/0.874
PortraitAes-SFT
✓
0.856/0.896
0.852/0.885
0.873 / 0.868
0.904/0.905
0.855 /0.877
0.810/0.830
0.909/0.917
PortraitAes-GRPO (Gaussian only)
✓
✓
0.863 / 0.901
0.869 / 0.897
0.872/ 0.868
0.907 / 0.907
0.854/ 0.883
0.820 / 0.837
0.915 / 0.924
PortraitAes-GRPO
✓
✓
✓
0.875/0.912
0.878/0.903
0.889/0.891
0.914/0.913
0.879/0.901
0.847/0.863
0.924/0.934
Table 3: Component ablation on the standard PortraitAes-Bench. Each result is reported as PLCC/SRCC; higher is better. Checkmarks denote enabled training components.
Method
Training component
PLCC/SRCC
SFT
Gaussian
Cross-BPR
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
Qwen3.5-VL-9B
0.515/0.513
0.346/0.281
0.558/0.506
0.602/0.574
0.385/0.474
0.391/0.413
0.649/0.629
PortraitAes-SFT
✓
0.562/ 0.668
0.558/0.524
0.737 / 0.723
0.795 / 0.753
0.570/ 0.598
0.549/0.559
0.780/ 0.758
PortraitAes-GRPO (Gaussian only)
✓
✓
0.617 /0.624
0.597 / 0.540
0.713/0.694
0.768/0.731
0.661 /0.590
0.610 / 0.587
0.789 /0.755
PortraitAes-GRPO
✓
✓
✓
0.679 / 0.650
0.646 / 0.614
0.768 / 0.750
0.789 / 0.752
0.689 / 0.627
0.650 / 0.629
0.829 / 0.795
Table 4: Component ablation on the PortraitAes-Bench hard-case set. Each result is reported as PLCC/SRCC; higher is better. Checkmarks denote enabled training components.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Resource
Scale
Port.
Supervision
Dim.
Text
AVA ( Murray et al., 2012 )
255K
no
ratings
global
limited
AADB ( Kong et al., 2016 )
10K
no
score/attr./rank
attr.
no
ArtiMuse ( Cao et al., 2026 )
10K
no
score/attr./text
8 attr.
yes
UniPercept ( Cao et al., 2025 )
800K
no
public VR/VQA
perception
yes
Qwen-Image-Bench ( Li et al., 2026 )
1K
no
judge scores
5D/56F
yes
MPAD/MPAAD ( Huang et al., 2022 )
18K/3K
yes
overall/attr.
3 attr.
no
Appendix
Table 5: Comparison with representative aesthetics and portrait-related resources. PortraitAes-Bench emphasizes intent-conditioned, multi-dimensional portrait understanding rather than a single global score.
Figure 5: Specialist-score distributions in PortraitAes-Bench. The six portrait-aesthetic dimensions show distinct score profiles rather than a single shared difficulty pattern, supporting dimension-wise evaluation of structure, composition, image quality, tone, theme/style, and emotion.
Intent
Struct.
Comp.
Quality
Tone
Theme
Emotion
I Functional
0.20
0.20
0.30
0.15
0.10
0.05
II Commercial
0.10
0.15
0.35
0.25
0.10
0.05
III Narrative
0.15
0.15
0.15
0.20
0.10
0.25
IV Aesthetic
0.10
0.15
0.15
0.30
0.20
0.10
V Social
0.05
0.15
0.15
0.20
0.30
0.15
VI Artistic
0.10
0.15
0.05
0.20
0.20
0.30
Appendix
Table 6: Intent-conditioned fusion weights used by the router.
Hyperparameter
Value
Hyperparameter
Value
Completions per prompt K
8
Training epochs
1
Learning rate
2×10−6
Warmup ratio
0.05
KL coefficient
0.02
Clip radius
0.12
Per-device batch size
4
Gradient accumulation
1
Calibration/ranking weights
1.4 / 0.8
Semantic/format weights
0.05 / 0.02
Gaussian bandwidth σ
0.5
Primary/sub-score weights
0.8 / 0.2
Appendix
Table 7: Hyperparameters for rank-aligned GRPO.
Figure 6: Structured predictions for twelve portraits across six intent categories, grouped by category and ordered by decreasing overall score. Colored bars show the six specialist scores.
Figure 7: Dimension-level aesthetic explanations for an aesthetic portrait with an overall score of 3.97.
Figure 8: Dimension-level aesthetic explanations for a snapshot with an overall score of 2.61.
Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt only one, leaving their complementarity unmeasured under controlled conditions. We introduce PPaint, a matched dual-protocol benchmark in which 15 domain experts, 5 per category, annotate 150 Chinese paintings under both protocols across five aesthetic dimensions, collecting 45,900 pairwise expert judgments through a locally dense preference design alongside the matched ratings. The matched design reveals complementary strengths: preferences yield more consistent ordinal rankings, while ratings anchor the absolute score scale. Fusing both signals via two independent preference-to-score methods yields a fused expert ground truth on which the two constructions converge to nearly identical scores. The same preference-to-score principle extends to label-free VLM training. PSDistill converts VLM pairwise judgments into calibrated pseudo-scores via an Elo reference pool, and trains the same VLM with confidence-weighted ranking optimization to produce a single-pass aesthetic scorer. Trained on a single painting category, the distilled Qwen3-VL-8B improves mean SRCC from 0.504 to 0.709 across all three categories, outperforming all open-source baselines including the dedicated aesthetic model ArtiMuse and matching closed-source Gemini-3.1-Pro within 0.04 SRCC at single-pass inference cost, with cross-domain transfer further validated on APDDv2. We will release the full PPaint dataset and training code.
Yuanpei Zhao, Jie Lin, Chao Zhang +5
1Sichuan University · 2NetEase Fuxi AI Lab · †Work done during an internship at NetEase Fuxi AI Lab. +1
Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predicting a scalar score for a single image. We first ask whether such scores faithfully capture comparative preference: in a controlled study with eight expert annotators, score-derived rankings align poorly with the same annotators' direct comparisons, while direct ranking yields substantially higher inter-annotator agreement on best- and worst-image labels. Motivated by this finding, we introduce the Visual Aesthetic Benchmark (VAB), which casts aesthetic evaluation as comparative selection over candidate sets with matched subject matter. VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, with labels derived from the consensus of 10 independent expert judges per task. Evaluating 20 frontier MLLMs and six dedicated visual-quality reward models, we find that the strongest system identifies both the best and the worst image correctly across three random permutations of the candidate order in only 26.5% of tasks, far below the 68.9% achieved by human experts. Fine-tuning a 35B-parameter model on 2,000 expert examples brings its accuracy close to that of a 397B-parameter open-weight model, suggesting that the comparative signal in VAB is transferable. Together, these results expose a clear and measurable gap between current multimodal models and expert aesthetic judgment, and VAB provides the first set-based, expert-grounded testbed on which that gap can be tracked and closed.
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.