Style transfer lacks a reliable evaluation standard: ground truth is inherently ill-defined, and existing automatic metrics often fail to reflect human preference. This paper introduces ASTRA (Assessment of Style TRansfer Algorithms), an approach for automatic evaluation of style transfer algorithms; it contains two components, ASTRA-Data and ASTRA-Score. ASTRA-Data consists of a benchmark image set of content and style references, a collection of style transfer results generated on the benchmark set, and user study data capturing human judgements through a two-stage pairwise comparison protocol. From these annotations, we derive ranking-based ground truth for content preservation, style fidelity, and overall preference. Based on ASTRA-Data, we construct ASTRA-Score, a learnt evaluator that predicts preference-aligned scores from content-style-stylization image triplets, enabling automatic and scalable evaluation of new models applied to the benchmark set. Experimental results demonstrate that ASTRA-Score achieves substantially higher correlation with human rankings compared to prior metrics. Overall, ASTRA establishes a robust mechanism for standardised evaluation of style transfer methods.
Figures & tables
Figure 1 : Examples of selected images produced by different algorithms.
Figure 2 : Examples of the two evaluation interfaces used in our user study. ( Top ) In Stage 1, participants compare two stylised results generated from the same content-style pair by different methods, together with the corresponding content and style references. ( Bottom ) In Stage 2, participants compare two stylised results that come from different content-style pairs.
Method
Stage 1
Stage 2
Total
AAMS [ YRX∗19 ]
2886
2712
5598
AdaAttN [ LLH∗21 ]
2983
2854
5837
AdaIN [ HB17 ]
3026
2845
5871
ArtFlow [ AHS∗21 ]
2960
2855
5815
ChatGPT [ AAA∗23 ]
2836
2771
5607
Gatys [ GEB16 ]
3052
2859
5911
Table 1 : Number of appearances of each method in Stage 1 and Stage 2 pairwise comparisons.
Figure 3 : Top-5 (top row) and bottom-5 (bottom row) stylised results selected from the top-10 and bottom-10 samples ranked by their mean Rank Centrality (RC) scores for overall preference. The style references for each column, arranged from left to right and top to bottom, are: Kandinsky, Carr, Picasso, Ptolemaic Egyptian, Chinese landscape, Australian Aboriginal, Australian Aboriginal, Rembrandt, Hokusai, and Picasso.
Figure 4 : Scatter plots of (a) overall effectiveness vs. content preservation; (b) overall effectiveness vs. style fidelity.
Figure 5 : Trade-off between style similarity and content preservation measured from Stage 1 pairwise comparisons. The dashed curve indicates the Pareto frontier of the two objectives, which highlights the optimal trade-off boundary between the two criteria.
Figure 6 : Content–style preference trade-off under different style references.
Figure 7 : Human preference rankings derived from Rank Centrality. Each point shows the mean score and the error bar denotes bootstrap uncertainty. “Overall preference” refers to the subjects’ holistic assessment of the effectiveness of the style transfer.
Figure 8 : Overview of the proposed evaluation model. A frozen VGG11 extracts content features ( fcg,fc ) and style features ( fsg,fs ) from the generated image Ig , content reference Ic , and style reference Is . Three independent heads are employed: the Content Head uses ϕ(fcg,fc) , the Style Head uses ψ(fsg,fs) , and the Overall Head combines both relation features to predict an overall quality score.
Metric
Spearman
Kendall
Content
CFSD [ CHH24 ]
0.281 ± 0.212
0.196 ± 0.151
NMI [ SHH98 ]
0.415 ± 0.141
0.288 ± 0.100
SIFID [ SDM19 ]
0.429 ± 0.170
0.300 ± 0.129
GMSD [ XZMB13 ]
0.456 ± 0.176
0.323 ± 0.133
FSIM [ ZZMZ11 ]
0.515 ± 0.160
0.368 ± 0.118
Table 2 : Correlation between automatic metrics and study-derived rankings across the evaluated methods. Baseline metrics are evaluated using leave-one-method-out (LOMO) evaluation and reported as mean ± standard deviation.
Metric
Spearman
Kendall
Content
CFSD [ CHH24 ]
0.440 ± 0.055
0.310 ± 0.044
Qwen [ BBC∗23 ]
0.455 ± 0.184
0.357 ± 0.145
NMI [ SHH98 ]
0.492 ± 0.043
0.345 ± 0.033
GMSD [ XZMB13 ]
0.498 ± 0.103
0.361 ± 0.081
FSIM [ ZZMZ11 ]
0.544 ± 0.086
0.395 ± 0.074
Table 3 : Correlation between automatic metrics and study-derived rankings across the evaluated methods. Baseline metrics are evaluated using leave-one-content-out (LOCO) evaluation and reported as mean ± standard deviation.
Metric
Spearman
Kendall
Content
CFSD [ CHH24 ]
0.250 ± 0.131
0.178 ± 0.091
GMSD [ XZMB13 ]
0.391 ± 0.179
0.280 ± 0.133
Qwen [ BBC∗23 ]
0.417 ± 0.191
0.325 ± 0.143
FSIM [ ZZMZ11 ]
0.435 ± 0.168
0.312 ± 0.125
SIFID [ SDM19 ]
0.467 ± 0.152
0.334 ± 0.114
Table 4 : Correlation between automatic metrics and study-derived rankings across the evaluated methods. Baseline metrics are evaluated using leave-one-style-out (LOSO) evaluation and reported as mean ± standard deviation.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Top-3 and bottom-3 stylised results for each method ranked by overall human preference scores. In each row, the left three images show the top-3 results and the right three show the bottom-3 results. Methods are listed in alphabetical order: AAMS, AdaAttN, AdaIN, ArtFlow, ChatGPT, NST-Ghiasi, Gatys, LLIP, OmniStyle, SANET, StyTR-2, and USO.
Figure 10 : Illustration of content consistency: the same scene rendered in different styles.
Figure 11 : Illustration of style consistency: different contents rendered in the same style.
Figure 12 : Example of style transfer: content from one image rendered in the style of another.
Figure 13 : Ranking stability under simulated within-group comparisons (Stage 1). The Spearman correlation between the recovered ranking and the assumed ground-truth ranking increases as the number of comparisons grows. The curve rapidly converges and exceeds ρ=0.99 at approximately 1.2×104 comparisons, indicating that the estimated method ranking becomes highly stable beyond this point. Each point represents the mean correlation over multiple simulation runs.
Figure 14 : Examples of the two illustrative examples used in our user study.
Figure 15 : Global ranking recovery under simulated cross-group comparisons (Stage 2). Additional comparisons between items belonging to different content-style pairs gradually improve the accuracy of the recovered global ranking. The Spearman correlation steadily increases as more cross-group comparisons are introduced, demonstrating that sufficient Stage 2 comparisons enable reliable global ranking estimation.
Style
AAMS
AdaAttN
AdaIN
ArtFlow
ChatGPT
Gatys
LLIP
OmniStyle
SANET
StyTR-2
USO
NST-Ghiasi
Australian Aboriginal
236
288
230
276
177
238
255
228
268
213
276
241
Bierstadt
229
249
280
225
248
281
201
241
276
263
248
227
Carr
225
243
266
258
254
274
223
310
230
217
237
195
Hokusai
215
268
321
272
221
256
232
263
209
221
214
232
Gris
246
224
240
231
322
257
271
210
190
268
204
257
Picasso
236
251
291
225
268
251
212
213
290
253
230
208
Appendix
Table 5: Exposure counts between reference styles and style transfer methods in the Stage 1 user study. Each entry reports the number of pairwise comparisons in which a given method appears under the corresponding style reference.
Content
AAMS
AdaAttN
AdaIN
ArtFlow
ChatGPT
Gatys
LLIP
OmniStyle
SANET
StyTR-2
USO
NST-Ghiasi
angel
516
493
544
469
500
491
480
453
512
522
425
437
athletes
509
522
479
474
432
428
515
496
486
543
548
402
barn
389
520
549
498
497
494
437
473
525
479
510
489
berries
507
506
459
562
462
523
443
516
458
496
431
491
daisy
506
494
476
489
468
549
496
521
469
485
489
456
mac
459
448
519
468
477
567
498
474
453
509
516
460
Appendix
Table 6: Exposure counts between content images and style transfer methods in the Stage 1 user study. Each entry indicates the number of pairwise comparisons in which a given method appears for the corresponding content image.
Style
AAMS
AdaAttN
AdaIN
ArtFlow
ChatGPT
Gatys
LLIP
OmniStyle
SANET
StyTR-2
USO
NST-Ghiasi
Australian Aboriginal
214
254
264
246
232
237
254
228
229
202
250
269
Bierstadt
249
231
245
231
245
255
241
234
251
209
265
247
Carr
241
219
229
219
223
239
223
243
235
228
225
251
Hokusai
238
248
233
223
210
219
239
237
236
252
217
250
Gris
215
245
240
251
243
242
234
255
259
221
239
269
Picasso
209
237
241
228
232
237
252
223
222
248
257
241
Appendix
Table 7: Exposure counts between style references and methods in Stage 2 comparisons.
Content
AAMS
AdaAttN
AdaIN
ArtFlow
ChatGPT
Gatys
LLIP
OmniStyle
SANET
StyTR-2
USO
NST-Ghiasi
Angel
442
477
490
528
468
505
473
471
454
465
480
490
Athletes
456
493
461
441
503
502
494
504
444
452
482
520
Barn
473
508
478
500
486
470
510
454
497
452
480
466
Berries
446
437
492
470
423
445
472
474
438
460
459
519
Daisy
441
476
503
475
419
459
466
504
491
456
515
481
Mac
446
459
436
442
475
461
470
484
493
457
471
496
Appendix
Table 8: Exposure counts between content images and methods in Stage 2 comparisons.
Figure 16 : Content-style preference across different style references. Each subplot corresponds to a specific style, arranged from left to right and top to bottom as: Mondrian, Carr, Chinese Landscape, Kandinsky, Bierstadt, Rembrandt, Picasso, Australian Aboriginal, Hokusai, Ptolemaic Egyptian, Gris, and Veronese. Each point represents a stylisation method positioned by its content and style win rates. The red curve indicates the Pareto frontier of the two objectives, which highlights the optimal trade-off boundary between the two criteria. The relationship between content preservation and style similarity varies significantly across styles, highlighting strong style-dependent behaviour.
Figure 17 : Content-style preference across different content images. Each subplot corresponds to a specific content image, arranged from left to right and top to bottom as: Angel, Athletes, Barn, Berries, Daisy, and Mac. Each point represents a stylisation method positioned by its content and style win rates. The red curve indicates the Pareto frontier of the two objectives, which highlights the optimal trade-off boundary between the two criteria. Compared to the per-style analysis, the distributions remain largely consistent across different contents, indicating that the trade-off is relatively stable with respect to the input image.
Content
Style
Overall
Backbone
Spearman
Kendall
Spearman
Kendall
Spearman
Kendall
ResNet18
0.658 ± 0.209
0.489 ± 0.164
0.674 ± 0.076
0.497 ± 0.062
0.580 ± 0.156
0.418 ± 0.118
ResNet34
0.628 ± 0.198
0.464 ± 0.157
0.672 ± 0.077
0.492 ± 0.063
0.566 ± 0.179
0.410 ± 0.135
ResNet50
0.656 ± 0.173
0.488 ± 0.135
0.707 ± 0.088
0.521 ± 0.074
0.608 ± 0.156
0.442 ± 0.123
VGG19
0.671 ± 0.138
0.491 ± 0.115
0.659 ± 0.082
0.480 ± 0.067
0.622 ± 0.119
0.453 ± 0.095
ViT-B/16
0.691 ± 0.152
0.513 ± 0.127
0.648 ± 0.136
0.475 ± 0.115
0.505 ± 0.164
0.360 ± 0.128
Appendix
Table 9 : Backbone ablation for ASTRA-Score under leave-one-method-out (LOMO) evaluation. Results are reported as the mean ± standard deviation of Spearman’s ρ and Kendall’s τ across held-out stylisation methods. The best result for each criterion and correlation metric is shown in bold.
Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-quality large-scale triplet datasets with stylized ground truth forces existing methods to rely on self-supervised or proxy objectives, which limits model capability. To mitigate this gap, we design a data construction pipeline to build TST100K, a large-scale dataset of 100,000 content-reference-stylized triplets. At the core of this pipeline, we train a tone style scorer to ensure strict stylistic consistency for each triplet. In addition, existing methods typically extract content and reference features independently and then fuse them in a decoder, which may cause semantic loss and lead to inappropriate color transfer and degraded visual aesthetics. Instead, we propose ICTone, a diffusion-based framework that performs tone transfer in an in-context manner by jointly conditioning on both images, leveraging the semantic priors of generative models for semantic-aware transfer. Reward feedback learning using the tone style scorer is further incorporated to improve stylistic fidelity and visual quality. Experiments demonstrate the effectiveness of TST100K, and ICTone achieves state-of-the-art performance on both quantitative metrics and human evaluations.
Yuhai Deng, Huimin She, Wei Shen +4
VCIP, School of Computer Science, Nankai University · OPPO AI Center, OPPO Inc. · Nankai International Advanced Research Institute (Shenzhen Futian)
Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at https://github.com/0606zt/CLeaR.
We introduce Color Disentangled Style Transfer (CDST), a novel and efficient two-stream style transfer training paradigm which completely isolates color from style and forces the style stream to be color-blinded. With one same model, CDST unlocks universal style transfer capabilities in a tuning-free manner during inference. Especially, the characteristics-preserved style transfer with style and content references is solved in the tuning-free way for the first time. CDST significantly improves the style similarity by multi-feature image embeddings compression and preserves strong editing capability via our new CDST style definition inspired by Diffusion UNet disentanglement law. By conducting thorough qualitative and quantitative experiments and human evaluations, we demonstrate that CDST achieves state-of-the-art results on various style transfer tasks.