Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.
Figures & tables
Figure 1 : Overall comparison on Ecom-Bench visual quality and Ecom-CTR ranking. Hatched bars denote closed-source models.
Figure 2 : Showcases of KwaiMind across multiple image-editing tasks.
Figure 3 : Overall comparison on general image editing benchmarks: ImgEdit, GEdit, and the English and Chinese splits of REDEdit. Hatched bars denote closed-source models.
Table 1 : Sample states maintained by Coordinator Agent. Each state determines the next work-order destination. States ending in Escalated indicate that the corresponding loop budget has been exhausted and human review is required.
Figure 4 : Overview of the end-to-end data pipeline. Coordinator Agent routes samples through filtering, generation, captioning, and post-filtering with bounded feedback loops and human review.
Figure 5 : Composition of the curated SFT corpus. The upper half shows the e-commerce subset and the lower half shows the general editing subset; surrounding examples illustrate representative source–target pairs from different tasks.
Data source
Number of samples
ScaleEdit-12M ( Chen et al., 2026 )
12.0M editing pairs
X2Edit ( Ma et al., 2026 )
3.7M editing pairs
AnyEdit ( Yu et al., 2025 )
2.5M editing pairs
KwaiData
8.0M editing pairs
Pair-based source total
26.2M pairs
Maintained training corpus
∼ 1.8M pairs
Table 2 : Data sources processed by the Data Agent and the resulting maintained training corpus.
Figure 6 : Overview of the proposed post-training framework.
Model
PairAcc
Qwen3-VL
47.9
GPT-4o
50.2
Qwen3-VL-Trained
51.2
VAM
52.2
LLaVA
52.8
CG4CTR
53.1
Table 3 : Pairwise ranking accuracy on the public CreativeRanking benchmark. Our vision-only variant (vision encoder plus MLP head, with the text branch and fusion module removed) is compared against general-purpose and task-specific baselines.
Task
Description
Virtual Try-On †
Dress a model image with a specified garment
Clothing Detail
Generate a zoomed-in detail view of garment texture
Clothing Display
Clothe a virtual mannequin model with a target garment
Universal Wearing †
Apply any wearable product to a model image
Pose Change
Alter the model’s pose while preserving garment appearance
Background Replace
Swap the product backdrop with a specified scene
Table 4 : The 11 tasks of Ecom-Bench. † denotes tasks with multi-image reference input (model image + product image). 100 test samples are used per task.
Task
D1
D2
D3
D4
Virtual Try-On
E5 Fit & Silhouette
E3 Texture
G2 Seamless
G3 Physical
Clothing Detail
E1 Product Id.
E3 Texture
G5 Quality
G6 Aesthetics
Clothing Display
E3 Texture
E4 Model
G3 Physical
G6 Aesthetics
Universal Wearing
E1 Product Id.
E4 Model
G2 Seamless
G3 Physical
Pose Change
E4 Model
E5 Fit & Silhouette
G1 Comply
G3 Physical
Background Replace
G1 Comply
E1 Product Id.
E8 Complete
G3 Physical
Table 5 : Evaluation dimension assignments for the 11 Ecom-Bench tasks. The G and E prefixes denote general and e-commerce-specific dimensions, respectively.
Panel A: ImgEdit (Gemini judge)
Model
Add
Adjust
Remove
Replace
Back.
Style
Extract
Action
Compose
Overall
Flux-2.0
4.21
3.61
4.07
4.36
3.78
4.56
3.13
3.23
2.26
3.69
Joy-Edit
4.34
4.06
4.25
4.25
4.00
4.94
4.40
3.10
3.44
4.09
Joy-Edit-Plus
4.41
4.24
3.62
3.67
3.53
4.41
2.22
3.39
2.60
3.57
LongCat
4.28
4.07
4.29
4.44
4.05
4.98
4.38
3.00
3.47
4.10
FireRed-Edit
4.36
4.11
4.40
4.59
3.94
4.96
4.84
2.89
2.90
4.11
Table 6 : Results on ImgEdit and GEdit under a common Gemini 3.1 Pro Preview judging protocol. All task columns and the overall score are reported. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined.
Panel A: REDEdit English
Model
Add
Adjust
Back.
Beauty
Color
Compose
Extract
Portrait
Low
Motion
Remove
Replace
Style
Text
View
Overall
Flux-2.0
3.96
3.32
4.09
3.02
3.28
2.96
1.31
3.79
3.66
4.30
3.43
4.12
4.41
2.75
2.95
3.42
Joy-Edit
3.95
3.32
3.90
2.29
3.58
3.01
2.46
3.54
3.11
3.86
3.64
4.10
4.75
3.28
1.67
3.36
Joy-Edit-Plus
2.90
1.69
2.43
1.42
1.93
2.22
1.07
3.42
2.20
3.38
1.83
2.10
3.18
2.25
1.75
2.25
LongCat
4.02
3.25
3.91
2.31
3.55
2.97
2.32
3.49
2.98
3.91
3.62
4.20
4.69
3.48
1.69
3.36
FireRed-Edit
4.37
3.75
4.27
2.82
3.92
3.51
2.38
3.60
2.89
4.33
4.10
4.33
4.77
3.66
2.39
3.67
Table 7 : Results on the English and Chinese splits of REDEdit. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined.
Ecom-Bench visual score and CTR
Model
VTO
Display
Wear
Detail
Pose
Back.
Outpaint
Extract
Text
TLR
Sell
CTR
Overall
Flux-2.0
4.09
3.75
2.23
3.04
3.79
3.36
4.57
3.70
2.68
3.46
3.65
293
3.48
Joy-Edit
–
3.61
–
2.66
4.04
3.21
3.46
3.97
3.78
2.36
4.06
–
3.46
Joy-Edit-Plus
4.13
3.49
2.15
2.22
3.74
2.93
3.60
2.69
2.76
3.00
3.91
406
3.15
LongCat
–
2.63
–
2.47
4.23
3.14
3.55
2.77
3.31
2.61
3.38
–
3.12
FireRed-Edit
4.27
2.78
1.81
1.82
4.33
3.27
4.23
3.63
3.57
3.33
3.79
362
3.35
Table 8 : Ecom-Bench visual scores and aggregate CTR ranking scores. VTO, Display, Wear, TLR, and Sell denote Virtual Try-On, Clothing Display, Universal Wearing, Tagline Removal, and Selling Point Display, respectively. For each of the 1,100 test cases, outputs are ranked by predicted CTR; rank-1, rank-2, and rank-3 appearances receive equal weight, and CTR reports their sum. Within the open-source and closed-source groups separately, the best result is shown in bold and the second-best is underlined. Missing results are denoted by “–”.
Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.
Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-scale preference annotation and additional model training. This creates a data-efficiency gap: humans can often infer the target evaluation criteria from only a few examples, while models are usually trained on hundreds of thousands of comparisons. We present RewardHarness, a self-evolving agentic reward framework that reframes reward modeling as context evolution rather than weight optimization. Instead of learning from large-scale annotations, RewardHarness aligns with human preferences by iteratively evolving a library of tools and skills from as few as 100 preference demonstrations. Given a source image, candidate edited images, and an editing instruction, an Orchestrator selects the most relevant subset of tools and skills from the maintained library, and a frozen Sub-Agent uses them to construct a reasoning chain that produces a preference judgment. By comparing predicted judgments with ground-truth preferences and analyzing successes and failures in the reasoning process, the Orchestrator automatically refines its library of tools and skills without additional human annotation. Using only 0.05% of the EditReward preference data, RewardHarness achieves 47.4% average accuracy on image-editing evaluation benchmarks, surpassing GPT-5 by 5.3 points. When used as a reward signal for GRPO fine-tuning, RL-tuned models achieve 3.52 on ImgEdit-Bench. Project page: https://rewardharness.com.
Yuxuan Zhang, Penghui Du, Bo Li +11
University of British Columbia · Kolors Team, Kuaishou Technology · University of Waterloo +4
Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions. However, in product-centric scenarios where preserving product features, branding, and textual elements are critical, current open and closed source models often struggle to maintain this fine-grained object identity. This issue is further compounded by the lack of datasets for instruction-based product image editing with text fidelity constraints, leaving it largely treated as an implicit capability of instruction-based image editing models. In this work, we introduce the ProductConsistency dataset which is designed to improve product-centric image editing. Our approach includes a supervised fine-tuning (SFT) dataset of 87k samples for product editing, a reinforcement learning (RL) dataset with 869 unique product images, and a new benchmark dataset, the ProductConsistency Benchmark, to allow rigorous and standardized evaluation of editing models. To guide RL training, we propose a Cyclic Consistency reward that enforces semantic preservation of product identity by using caption similarity between the original product description and captions generated from the edited image. We fine-tune both Qwen-Image-Edit-2511 and Flux.1-Kontext-dev using our dataset and demonstrate consistent improvements over baseline models in OCR and Perceptual metrics, and MLLM-based evaluations as well, indicating stronger product consistency, text rendering, and overall visual quality; with the Qwen-Image-Edit-2511 model achieving a 5x reduction in the character error rate. The code and pipeline is available at https://anonymous.4open.science/r/ProductConsistency-6FCC/README.md