Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.
Figures & tables
Figure 1 : Overall comparison on Ecom-Bench visual quality and Ecom-CTR ranking. Hatched bars denote closed-source models.
Figure 2 : Showcases of KwaiMind across multiple image-editing tasks.
Figure 3 : Overall comparison on general image editing benchmarks: ImgEdit, GEdit, and the English and Chinese splits of REDEdit. Hatched bars denote closed-source models.
Table 1 : Sample states maintained by Coordinator Agent. Each state determines the next work-order destination. States ending in Escalated indicate that the corresponding loop budget has been exhausted and human review is required.
Figure 4 : Overview of the end-to-end data pipeline. Coordinator Agent routes samples through filtering, generation, captioning, and post-filtering with bounded feedback loops and human review.
Figure 5 : Composition of the curated SFT corpus. The upper half shows the e-commerce subset and the lower half shows the general editing subset; surrounding examples illustrate representative source–target pairs from different tasks.
Data source
Number of samples
ScaleEdit-12M ( Chen et al., 2026 )
12.0M editing pairs
X2Edit ( Ma et al., 2026 )
3.7M editing pairs
AnyEdit ( Yu et al., 2025 )
2.5M editing pairs
KwaiData
8.0M editing pairs
Pair-based source total
26.2M pairs
Maintained training corpus
∼ 1.8M pairs
Table 2 : Data sources processed by the Data Agent and the resulting maintained training corpus.
Figure 6 : Overview of the proposed post-training framework.
Model
PairAcc
Qwen3-VL
47.9
GPT-4o
50.2
Qwen3-VL-Trained
51.2
VAM
52.2
LLaVA
52.8
CG4CTR
53.1
Table 3 : Pairwise ranking accuracy on the public CreativeRanking benchmark. Our vision-only variant (vision encoder plus MLP head, with the text branch and fusion module removed) is compared against general-purpose and task-specific baselines.
Task
Description
Virtual Try-On †
Dress a model image with a specified garment
Clothing Detail
Generate a zoomed-in detail view of garment texture
Clothing Display
Clothe a virtual mannequin model with a target garment
Universal Wearing †
Apply any wearable product to a model image
Pose Change
Alter the model’s pose while preserving garment appearance
Background Replace
Swap the product backdrop with a specified scene
Table 4 : The 11 tasks of Ecom-Bench. † denotes tasks with multi-image reference input (model image + product image). 100 test samples are used per task.
Task
D1
D2
D3
D4
Virtual Try-On
E5 Fit & Silhouette
E3 Texture
G2 Seamless
G3 Physical
Clothing Detail
E1 Product Id.
E3 Texture
G5 Quality
G6 Aesthetics
Clothing Display
E3 Texture
E4 Model
G3 Physical
G6 Aesthetics
Universal Wearing
E1 Product Id.
E4 Model
G2 Seamless
G3 Physical
Pose Change
E4 Model
E5 Fit & Silhouette
G1 Comply
G3 Physical
Background Replace
G1 Comply
E1 Product Id.
E8 Complete
G3 Physical
Table 5 : Evaluation dimension assignments for the 11 Ecom-Bench tasks. The G and E prefixes denote general and e-commerce-specific dimensions, respectively.
Panel A: ImgEdit (Gemini judge)
Model
Add
Adjust
Remove
Replace
Back.
Style
Extract
Action
Compose
Overall
Flux-2.0
4.21
3.61
4.07
4.36
3.78
4.56
3.13
3.23
2.26
3.69
Joy-Edit
4.34
4.06
4.25
4.25
4.00
4.94
4.40
3.10
3.44
4.09
Joy-Edit-Plus
4.41
4.24
3.62
3.67
3.53
4.41
2.22
3.39
2.60
3.57
LongCat
4.28
4.07
4.29
4.44
4.05
4.98
4.38
3.00
3.47
4.10
FireRed-Edit
4.36
4.11
4.40
4.59
3.94
4.96
4.84
2.89
2.90
4.11
Table 6 : Results on ImgEdit and GEdit under a common Gemini 3.1 Pro Preview judging protocol. All task columns and the overall score are reported. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined.
Panel A: REDEdit English
Model
Add
Adjust
Back.
Beauty
Color
Compose
Extract
Portrait
Low
Motion
Remove
Replace
Style
Text
View
Overall
Flux-2.0
3.96
3.32
4.09
3.02
3.28
2.96
1.31
3.79
3.66
4.30
3.43
4.12
4.41
2.75
2.95
3.42
Joy-Edit
3.95
3.32
3.90
2.29
3.58
3.01
2.46
3.54
3.11
3.86
3.64
4.10
4.75
3.28
1.67
3.36
Joy-Edit-Plus
2.90
1.69
2.43
1.42
1.93
2.22
1.07
3.42
2.20
3.38
1.83
2.10
3.18
2.25
1.75
2.25
LongCat
4.02
3.25
3.91
2.31
3.55
2.97
2.32
3.49
2.98
3.91
3.62
4.20
4.69
3.48
1.69
3.36
FireRed-Edit
4.37
3.75
4.27
2.82
3.92
3.51
2.38
3.60
2.89
4.33
4.10
4.33
4.77
3.66
2.39
3.67
Table 7 : Results on the English and Chinese splits of REDEdit. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined.
Ecom-Bench visual score and CTR
Model
VTO
Display
Wear
Detail
Pose
Back.
Outpaint
Extract
Text
TLR
Sell
CTR
Overall
Flux-2.0
4.09
3.75
2.23
3.04
3.79
3.36
4.57
3.70
2.68
3.46
3.65
293
3.48
Joy-Edit
–
3.61
–
2.66
4.04
3.21
3.46
3.97
3.78
2.36
4.06
–
3.46
Joy-Edit-Plus
4.13
3.49
2.15
2.22
3.74
2.93
3.60
2.69
2.76
3.00
3.91
406
3.15
LongCat
–
2.63
–
2.47
4.23
3.14
3.55
2.77
3.31
2.61
3.38
–
3.12
FireRed-Edit
4.27
2.78
1.81
1.82
4.33
3.27
4.23
3.63
3.57
3.33
3.79
362
3.35
Table 8 : Ecom-Bench visual scores and aggregate CTR ranking scores. VTO, Display, Wear, TLR, and Sell denote Virtual Try-On, Clothing Display, Universal Wearing, Tagline Removal, and Selling Point Display, respectively. For each of the 1,100 test cases, outputs are ranked by predicted CTR; rank-1, rank-2, and rank-3 appearances receive equal weight, and CTR reports their sum. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined. Missing results are denoted by “–”.