Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.
Figures & tables
Figure 1: Overview of PortraitAes-Bench. Each portrait is routed to one of nine photographic intents and evaluated by six specialists spanning 22 secondary criteria, yielding structured routing, rating, and intent-conditioned fusion targets.
Figure 2: Construction of PortraitAes-Bench. Expert rubrics define an intent router and six specialist roles; schema checks, Qwen3.5 Flash screening, and expert review precede intent-weighted fusion.
Figure 3: Intent-conditioned overall-score distributions. The standard benchmark spans a broad range, whereas the hard-case set concentrates visually similar portraits for fine-grained ranking.
Figure 4: Illustration of the numerical rewards. Gaussian calibration anchors predictions to rubric-derived scores; same-dimension cross-group ranking penalizes ordinal inversions. Images and rollout values illustrate the mechanism rather than report an evaluation case.
Model
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
Qwen3.5-VL-9B
0.833
0.822
0.724
0.742
0.741
0.698
0.807
0.789
0.706
0.714
0.698
0.740
0.857
0.874
InternVL3-8B
0.393
0.453
0.518
0.505
0.581
0.623
0.530
0.570
0.448
0.511
0.297
0.296
0.607
0.668
PortraitCraft-4B †
0.698
0.801
0.647
0.760
0.738
0.773
0.578
0.698
0.613
0.754
0.535
0.655
0.652
0.785
Qwen3.5 Flash
0.791
0.837
0.620
0.700
0.807
0.798
0.795
0.818
0.590
0.705
0.667
0.723
0.804
0.763
ArtiMuse †
0.741
0.785
0.744
0.796
0.556
0.602
0.688
0.717
0.725
0.756
0.712
0.703
0.785
0.804
Table 1: Standard PortraitAes-Bench results (PLCC/SRCC; higher is better). † denotes diagnostic correlations from a holistic-only scorer.
Model
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
Qwen3.5-VL-9B
0.515
0.513
0.346
0.281
0.558
0.506
0.602
0.574
0.385
0.474
0.391
0.413
0.649
0.629
InternVL3-8B
0.030
0.036
0.156
0.112
0.402
0.476
0.228
0.273
-0.010
-0.033
0.121
0.184
0.210
0.215
PortraitCraft-4B †
0.406
0.403
0.361
0.331
0.479
0.463
0.151
0.145
0.342
0.328
0.288
0.297
0.406
0.400
Qwen3.5 Flash
0.522
0.522
0.240
0.214
0.578
0.543
0.624
0.601
0.304
0.411
0.407
0.400
0.615
0.635
ArtiMuse †
0.430
0.421
0.407
0.398
0.336
0.346
0.368
0.350
0.410
0.418
0.387
0.383
0.501
0.495
Table 2: PortraitAes-Bench hard-case results (PLCC/SRCC; higher is better). † denotes diagnostic correlations from a holistic-only scorer.
Method
Training component
PLCC/SRCC
SFT
Gaussian
Cross-BPR
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
Qwen3.5-VL-9B
0.833/0.822
0.724/0.742
0.741/0.698
0.807/0.789
0.706/0.714
0.698/0.740
0.857/0.874
PortraitAes-SFT
✓
0.856/0.896
0.852/0.885
0.873 / 0.868
0.904/0.905
0.855 /0.877
0.810/0.830
0.909/0.917
PortraitAes-GRPO (Gaussian only)
✓
✓
0.863 / 0.901
0.869 / 0.897
0.872/ 0.868
0.907 / 0.907
0.854/ 0.883
0.820 / 0.837
0.915 / 0.924
PortraitAes-GRPO
✓
✓
✓
0.875/0.912
0.878/0.903
0.889/0.891
0.914/0.913
0.879/0.901
0.847/0.863
0.924/0.934
Table 3: Component ablation on the standard PortraitAes-Bench. Each result is reported as PLCC/SRCC; higher is better. Checkmarks denote enabled training components.
Method
Training component
PLCC/SRCC
SFT
Gaussian
Cross-BPR
Structure
Composition
Quality
Tone
Theme
Emotion
Overall
Qwen3.5-VL-9B
0.515/0.513
0.346/0.281
0.558/0.506
0.602/0.574
0.385/0.474
0.391/0.413
0.649/0.629
PortraitAes-SFT
✓
0.562/ 0.668
0.558/0.524
0.737 / 0.723
0.795 / 0.753
0.570/ 0.598
0.549/0.559
0.780/ 0.758
PortraitAes-GRPO (Gaussian only)
✓
✓
0.617 /0.624
0.597 / 0.540
0.713/0.694
0.768/0.731
0.661 /0.590
0.610 / 0.587
0.789 /0.755
PortraitAes-GRPO
✓
✓
✓
0.679 / 0.650
0.646 / 0.614
0.768 / 0.750
0.789 / 0.752
0.689 / 0.627
0.650 / 0.629
0.829 / 0.795
Table 4: Component ablation on the PortraitAes-Bench hard-case set. Each result is reported as PLCC/SRCC; higher is better. Checkmarks denote enabled training components.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Resource
Scale
Port.
Supervision
Dim.
Text
AVA ( Murray et al., 2012 )
255K
no
ratings
global
limited
AADB ( Kong et al., 2016 )
10K
no
score/attr./rank
attr.
no
ArtiMuse ( Cao et al., 2026 )
10K
no
score/attr./text
8 attr.
yes
UniPercept ( Cao et al., 2025 )
800K
no
public VR/VQA
perception
yes
Qwen-Image-Bench ( Li et al., 2026 )
1K
no
judge scores
5D/56F
yes
MPAD/MPAAD ( Huang et al., 2022 )
18K/3K
yes
overall/attr.
3 attr.
no
Appendix
Table 5: Comparison with representative aesthetics and portrait-related resources. PortraitAes-Bench emphasizes intent-conditioned, multi-dimensional portrait understanding rather than a single global score.
Figure 5: Specialist-score distributions in PortraitAes-Bench. The six portrait-aesthetic dimensions show distinct score profiles rather than a single shared difficulty pattern, supporting dimension-wise evaluation of structure, composition, image quality, tone, theme/style, and emotion.
Intent
Struct.
Comp.
Quality
Tone
Theme
Emotion
I Functional
0.20
0.20
0.30
0.15
0.10
0.05
II Commercial
0.10
0.15
0.35
0.25
0.10
0.05
III Narrative
0.15
0.15
0.15
0.20
0.10
0.25
IV Aesthetic
0.10
0.15
0.15
0.30
0.20
0.10
V Social
0.05
0.15
0.15
0.20
0.30
0.15
VI Artistic
0.10
0.15
0.05
0.20
0.20
0.30
Appendix
Table 6: Intent-conditioned fusion weights used by the router.
Hyperparameter
Value
Hyperparameter
Value
Completions per prompt K
8
Training epochs
1
Learning rate
2×10−6
Warmup ratio
0.05
KL coefficient
0.02
Clip radius
0.12
Per-device batch size
4
Gradient accumulation
1
Calibration/ranking weights
1.4 / 0.8
Semantic/format weights
0.05 / 0.02
Gaussian bandwidth σ
0.5
Primary/sub-score weights
0.8 / 0.2
Appendix
Table 7: Hyperparameters for rank-aligned GRPO.
Figure 6: Structured predictions for twelve portraits across six intent categories, grouped by category and ordered by decreasing overall score. Colored bars show the six specialist scores.
Figure 7: Dimension-level aesthetic explanations for an aesthetic portrait with an overall score of 3.97.
Figure 8: Dimension-level aesthetic explanations for a snapshot with an overall score of 2.61.