Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight, or an interpolation factor. This strength relates poorly to perceptual change. The image can partially revert as the slider moves, long stretches of the range produce no visible difference, and short intervals transform the image abruptly. Remapping the strength could fix this uneven pace, but only if the trajectory is monotone, which current methods do not enforce. We therefore distinguish the slider from the strength, and require perceptual distance from the input to grow linearly with the slider value. We introduce UniSlider, a lightweight LoRA trained on a few-step editing backbone so that its strength approximates this ideal slider. Few-step sampling lets us impose this objective in pixel space without intermediate ground truth, and the backbone's output is preserved at full strength. However, a low-rank adapter cannot make the strength fully uniform. Our slider is thus an inference-time remapping of the strength, obtained by adaptive sampling. Since training optmizes to make the trajectory monotone, this remapping closes the remaining gap without extra training or parameters. On a new benchmark of 300 continuous edits evaluating uniformity, monotonicity, edit fidelity, and identity preservation, UniSlider outperforms all prior methods and is preferred in a user study.
Figures & tables
Figure 1: UniSlider produces perceptually uniform edit strengths. Each row shows one image edited at increasing strength under the instruction above it. Equidistant steps along the slider produce equal amounts of perceptual change; the full interval is thus utilized smoothly and predictably.
Figure 2: Training closes most of the gap to a uniform slider; adaptive sampling closes the rest. All three rows edit the same image with the same instruction. Left: frames at equal slider values for SliderEdit, our loss alone ( Vanilla ), and our full UniSlider method with strength reparametrization. Right: perceptual distance from the input to every slider frame, normalized by the total input–edit distance. The diagonal marks a perfectly uniform slider. We use DreamSim ( Fu et al., 2023 ) as a perceptual metric.
Figure 3: UniSlider training scheme. We sample s , scale the trainable LoRA against the frozen backbone ( Eq. 3 ), and run the full T -step inference process to get xs ; xedit is cached once at s=1 . The loss ( Eq. 4 ) pushes xs to the s -implied perceptual position along the path from xsrc to xedit .
Figure 4: Uniformity degrades with the number of training examples, and is recovered by adaptive sampling. Normalized DreamSim distance from the input to the output, for adapter rank r and N overfitting examples. The diagonal is the ideal uniform slider. First six columns: the strength is the slider ( s=u ). Rightmost column (full method): one adapter trained on the full dataset, evaluated on N=32 unseen examples, with adaptive sampling mapping u to u=g(s) .
Edit Fidelity
Identity
Endpoint
IQA
Monotonicity
Uniformity
User Study
Method
CLIP dir ↑
VQA edit ↓
VQA id ↓
RMSE ↓
MUSIQ ↑
VVQA↓
VLPIPS↓
CV-DSim ↓
★↑
≡↑
K. Kontext
0.078
0.430
0.161
0.021
65.34
0.047
0.116
1.153
40.3%
46.8%
Diff. Steering
0.107
0.710
0.167
0.151
64.59
0.094
0.390
1.045
61.4%
43.9%
FlowSlider
0.090
0.322
0.378
0.047
65.84
0.067
0.165
0.675
14.1%
27.0%
VeloEdit
0.079
0.470
0.151
0.023
53.18
0.072
0.078
0.996
57.7%
61.7%
SliderEdit
0.087
0.440
0.103
0.048
63.70
0.047
0.251
1.199
42.2%
35.6%
Table 1: Comparison against baselines. All metrics are averaged over the 300 benchmark examples. Shading marks the best and second-best method per column. ★ marks the method preferred overall in the user study and ≡ the one preferred for uniformity.
Figure 5: UniSlider results. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B.
Figure 6: Composing a masked edit with a global one. A 4×4 grid over two instructions applied to the same input. Strength increases left to right and top to bottom. The second axis is obtained by re-running UniSlider on the output of the first, no composed or masked examples appear in training.
Table 8
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Backbone and DreamSim failure case. The backbone’s default edit changes the colors of the ball, and this change is carried along the whole slider. DreamSim places the second frame on the diagonal because of a small color change on the right of the ball, although the ball is not yet inflated.
Figure 8: Adaptive sampling vs. dense reference. Left: five of the eleven frames obtained with adaptive sampling (top) and with the dense reference, which inverts a PCHIP spline fitted to 51 measurements (bottom). Right: normalized DreamSim distance from the input to each frame against the slider value u . Without remapping (dashed), adaptive sampling (blue) and dense fit (red).
Figure 9: Training data. Nine of the 300 (image, instruction) pairs. These two things are the entire supervision, the target at s=1 is the frozen backbone’s default edit, so no edited image and no intermediate is required. Instructions carry no degree or mechanism (no “slightly”, no “by 30%”), since strength is carried entirely by s ; they are correspondingly short, a median of eight words.
Edit Fidelity
Identity
Endpoint
IQA
Monotonicity
Uniformity
Method
CLIP dir ↑
VQA edit ↓
VQA id ↓
RMSE ↓
MUSIQ ↑
VVQA↓
VLPIPS↓
CV-DSim ↓
Global
K. Kontext
0.097
0.414
0.076
0.020
62.86
0.042
0.061
0.985
Diff. Steering
0.124
0.660
0.025
0.182
61.91
0.128
0.348
0.986
FlowSlider
0.095
0.242
0.143
0.047
64.44
0.045
0.167
0.673
VeloEdit
0.106
0.470
0.053
0.021
54.19
0.070
0.022
0.782
SliderEdit
0.106
0.412
0.037
0.081
62.78
0.068
0.132
1.000
Appendix
Table 4: Benchmark results by edit type. The same metrics as Tab. 1 , split over the three edit categories ( 100 examples each, 300 total). K marks a method running on the FLUX.2-klein-4B backbone, i.e. the same base as ours, separating backbone advantage from method.
Edit Fid.
IQA
Monotonicity
Uniformity
Fidelity × Preservation
Method
QEval
CLIPIQA ↑
VDreamSim↓
VQeval↓
CV-TPIPS ↓
VQA edit×(1− VQA id)↑
KontinuousKontext
0.660
0.488
0.0007
0.167
0.438
0.357
DiffusionSteering
0.759
0.489
0.0008
0.196
0.468
0.592
FlowSlider
0.607
0.503
0.0004
0.175
0.420
0.196
VeloEdit
0.682
0.402
0.0005
0.168
0.427
0.407
SliderEdit
0.671
0.469
0.0009
0.147
0.433
0.386
Appendix
Table 5: Additional metrics. Averaged over the 300 benchmark examples. Averaged over all benchmark examples.
Figure 10: Benchmark 2AFC user study. Left: a screenshot of the study interface, in which annotators drag a slider back and forth to compare methods A and B at matching strengths. Right: Thurstone Case V values with confidence intervales from the pairwise judgments, for the two questions asked: which method is better overall, and which is more uniform. A total of 25 observers participated in the study making a total of 881 examples.
Figure 11: Perceptual metric 2AFC user study. Left: two examples of the rendered images at different strengths. Right: PLCC and SRCC for all evaluated metrics. A total of 15 observers participated in the study making a total of 538 annotations.
Figure 12: Benchmark comparisons with per-frame metrics. The first two examples compare UniSlider against Firefly Tune (top) and against FlowSlider and DiffusionSteering (second), annotating each frame with its VQAedit and VQAid values. The last two compare UniSlider against Firefly Tune and SliderEdit, plotting instead the perceptual distance from the input at each frame, so that a flat segment marks a dead zone and a steep one a jump. Note how UniSlider achieves better uniformity across all different types of edits.
Figure 13: Additional UniSlider results with changes in visual style. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 14: Additional UniSlider results with attribute edits. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 15: Additional UniSlider results with structural edits. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 16: Additional UniSlider results with structural edits. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 17: Additional UniSlider with masking. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 18: Additional 2D UniSlider results. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Figure 19: Additional 2D UniSlider results. From left to right input image, all the intermediate strength levels and the default output of Flux-Klein-4B. All the input images are from the DIV2K ( Agustsson & Timofte, 2017 ) dataset.
Edit Fidelity
Identity
Endpoint
IQA
Monotonicity
Uniformity
Efficiency
CLIP dir ↑
VQA edit ↓
VQA id ↓
RMSE ↓
MUSIQ ↑
VVQA↓
VLPIPS↓
CV-DSim ↓
Time(s)
Param
Qwen-Edit-Dist
0.102
0.693
0.158
0.059
67.21
0.055
0.031
0.641
32.4
29B
FireRed-1.0-Dist
0.109
0.802
0.114
0.183
67.93
0.089
0.052
1.038
68.7
29B
Flux-Klein-4B
0.112
0.799
0.125
0.018
67.09
0.046
0.016
0.652
7.2
8B
Flux-Klein-9B
0.114
0.820
0.141
0.018
68.19
0.051
0.018
0.748
14.8
17B
r=4
0.112
0.799
0.125
0.020
66.94
0.052
0.018
0.661
7.2
1M
Appendix
Table 6: Editing backbone and LoRA rank ablations. Top: UniSlider trained on four editing backbones. Bottom: adapter rank swept on Flux-Klein-4B, our default setting ( r=32 , shown in both blocks). Time is the wall-clock cost of generating an 11 -frame slider on an RTX Pro 6000; Params counts the full model in the top block and the trainable adapter in the bottom one.
Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: https://github.com/Showwwwwwwww/ARRO
Diffusion models are a leading paradigm for data generation, but training-free editing typically re-runs the full denoising trajectory for every edit strength, making iterative refinement expensive. To address this issue, we instead edit near the data manifold, where small local updates can replace repeated re-synthesis. To enable this, we estimate a local manifold tangent space directly from perturbed samples and prove that this sample-based estimator closely approximates the true tangent. Building on this guarantee, we devise a Jacobian-free algorithm that constructs a tangent frame via small perturbations to the initial noise and alternates small tangent moves with diffusion-based projections. Updates within this frame follow principled on-manifold directions while suppressing off-manifold drift, enabling fine-grained edits without full re-diffusion or additional training. Edit strength is controlled by the number of steps for rapid, continuous adjustments that preserve fidelity and plug into existing samplers. Empirically, the resulting tangent directions yield smooth, semantic unsupervised traversals and effective CLIP-guided optimization, demonstrating practical interactive continuous editing.
Yiming Zhang, Sitong Liu, Ke Li +3
University of California San Diego, La Jolla, CA, USA · University of Washington, Seattle, WA, USA · Xidian University, Xi’an, China
Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a single scalar reward. When multiple criteria matter, the prevailing practice of ``early scalarization'' collapses rewards into a fixed weighted sum. This commits the model to a single trade-off point at training time, providing no inference-time control over inherently conflicting goals -- such as prompt adherence versus source fidelity in image editing. We introduce ParetoSlider, a multi-objective RL (MORL) framework that trains a single diffusion model to approximate the entire Pareto front. By training the model with continuously varying preference weights as a conditioning signal, we enable users to navigate optimal trade-offs at inference time without retraining or maintaining multiple checkpoints. We evaluate ParetoSlider across three state-of-the-art flow-matching backbones: SD3.5, FluxKontext, and LTX-2. Our single preference-conditioned model matches or exceeds the performance of baselines trained separately for fixed reward trade-offs, while uniquely providing fine-grained control over competing generative goals.
Shelly Golan, Michael Finkelson, Ariel Bereslavsky +2