As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state of the generation and often causes unintended global changes. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on the noisy states of flow models, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Despite using no explicit spatial masks or object priors, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free.
Figures & tables
Figure 2 : Comparison of safety steering methods on Ring-a-Bell. Left: qualitative comparison using FLUX1. Right: trade-off between NSFW suppression and semantic fidelity on FLUX1 (top) and SD3.5 (bottom).
Figure 3 : The Steering Fields formulation enables a natural extension to image editing tasks. Evaluated on the PieBench dataset, Steering Fields achieves state-of-the-art semantic adherence performance, measured by CLIP and VQAScore, as well as a state-of-the-art HPSv2 score (see Tab. 2 ).
Method
Ring-a-Bell
P4D
COCO-1k (retain)
NudeNet ↓
VQAScore det↓
NudeNet ↓
VQAScore det↓
CLIP ↑
VQAScore ↑
FID ↓
FLUX
70.67 ±1.70
0.79 ±.007
55.67 ±1.85
0.64 ±.014
0.31 ±.0003
0.89 ±.0010
0.00 ±0.00
FLUX + UCE
72.25 ±1.83
0.74 ±.006
56.13 ±1.64
0.55 ±.010
0.31 ±.0009
0.89 ±.0008
38.31 ±.27
FLUX + ESD
41.93 ±2.74
0.68 ±.013
30.97 ±2.40
0.57 ±.019
0.30 ±.0005
0.86 ±.0019
47.62 ±.56
FLUX + EA
62.98 ±1.95
0.74 ±.010
38.92 ±2.36
0.51 ±.020
0.31 ±.0003
0.88 ±.0049
28.54 ±0.33
FLUX + Steering Fields
37.32 ±1.96
0.52 ±.012
24.43 ±1.46
0.33 ±.013
0.31 ±.0003
0.89 ±.0011
34.16 ±0.39
Table 1 : Steering towards safe content (T2I) and semantic retention on COCO (1k). NudeNet and VQAScore detect measure NSFW suppression on Ring-a-Bell and P4D. CLIP, VQAScore, and FID measure retention on COCO-1k, where values close to the base model are desirable.
Figure 4 : Comparison against baselines illustrating that inversion-based approaches trade editing fidelity for source preservation, whereas Steering Fields achieves both.
Method
CLIP-txt ↑
CLIP-img ↑
CLIP-dir ↑
VQAScore ↑
HPSv2 ↑
FLUX Img2Img
0.255 ±.0003
0.812 ±.0011
0.093 ±.0006
0.673 ±.0686
0.257 ±.0007
RF-Inversion
0.259 ±.0007
0.806 ±.0182
0.108 ±.0358
0.696 ±.0198
0.278 ±.0011
StableFlow
0.242 ±.0001
0.919 ±.0001
0.075 ±.0001
0.592 ±.0763
0.252 ±.0001
FlowEdit
0.260 ±.0001
0.865 ±.0003
0.116 ±.0004
0.702 ±.0533
0.287 ±.0002
Steering Fields (ours)
0.271 ±.0002
0.809 ±.0007
0.127 ±.0006
0.770 ±.0520
0.284 ±.0004
Table 2 : Image editing on PieBench. Semantic metrics averaged over 64 seeds. CLIP-txt, CLIP-dir, and VQAScore measure target prompt alignment; CLIP-img measures source preservation. Steering Fields leads on all semantic metrics despite not being optimised for editing.
Figure 5 : Examples of concept blending using Steering Fields. By setting λ=0 in Eq. 6 , it is possible to blend semantically distant concepts, like dogs with spaghetti (left) or clouds (right).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
zs←(1−σs)z1+σsz0
Appendix
Algorithm 1 Steering Fields
Method
VQA-Nudity ↓
CLIP ↑
SD3
0.74
0.32
SD3 + Ours
0.55
0.31
Appendix
Table 3 : Evaluation on nudity suppression and text-image alignment.
Method
T2I RiskyPrompt ↓
T2I Safety Violence ↓
CLIP ↑
FID ↓
Baseline
0.80
0.65
0.31
–
UCE
0.75
0.56
0.29
35.49
ESD
0.63
0.40
0.28
46.71
EraseAnything
0.61
0.37
0.28
34.92
Ours
0.55
0.34
0.28
48.77
Appendix
Table 4 : Comparison with concept erasure baselines. Lower is better for T2I RiskyPrompt, T2I Safety Violence, and FID; higher is better for CLIP.
Layer
λ
NudeNet ↓
CLIP ↑
FID ↓
VQA Score ↑
16
2.00
65.03
0.31
35.66
0.88
16
3.00
64.08
0.31
35.67
0.88
16
4.00
62.97
0.31
36.26
0.88
16
5.00
57.12
0.31
37.21
0.88
16
6.00
48.26
0.31
38.79
0.88
16
7.00
41.77
0.30
40.57
0.87
Appendix
Table 5 : Comparison across layers from which the steering direction is extracted, and steering strengths. Steering Fields consistently shows significantly better performance in suppression of sensitive contents, and retain performances on unrelated prompts. FID is also consistently lower compared to all possible configurations of activation steering.
Figure 6 : Additional Examples for concept editing via Steering Fields. Each pair shows the source generation (left) and the edited output (right). Semantically distant concepts are merged while preserving the compositional structure and visual coherence of the source, without inversion or finetuning.
Figure 7 : Concept blending via Steering Fields. Each pair shows the source generation (left) and the blended output (right). Semantically distant concepts are merged while preserving the compositional structure and visual coherence of the source, without inversion or fine-tuning.
Component
Specification
GPU
NVIDIA A100-SXM-64GB
GPU Memory
65536 MiB
CPU
Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz
Appendix
Table 6 : Computational resources used for the experiments.
Benchmark
Avg. Words
COCO
10.38
Ring-a-Bell
19.10
P4D
12.95
T2ISafetyViolence
11.90
T2IRiskyPrompt
31.21
Appendix
Table 7 : Average number of words per prompt across benchmarks.
Figure 8 : Qualitative examples over the Violence category. Our method suppresses unsafe content while remaining semantically and geometrically close to the original image.
Figure 9 : Qualitative examples on Ring-a-Bell for Flux vanilla, UCE, ESD, EraseAnything and Steering Fields (ours) for T2I steering. Our method suppresses unsafe content while remaining semantically and geometrically close to the original image.
Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate along the editing trajectory and become visible in the final result. To overcome this, we propose Constrained Edit Fields (CEF), which assigns each spatial location a continuous edit responsibility that quantifies its relevance to the desired edit. CEF estimates edit responsibility directly from the source image when the relevant content is present. For edits whose target content is absent from the source, CEF first generates an unconstrained proposal to reveal its realized spatial support and then estimates responsibility from that proposal. At each editing step, CEF decomposes the base edit field into prompt-induced and trajectory-induced components, enabling edit responsibility to preserve instruction-relevant updates while suppressing unintended trajectory-induced changes. Evaluated on all 700 PIE-Bench examples, CEF achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment. On Stable Diffusion 3.5 Medium, it reduces these metrics over the previous best results by 10.2%, 21.2%, and 48.0%, respectively.
Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity, we directly edit the velocity field via a safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score-based filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we propose VESFlow+, a stronger variant of VESFlow that not only edits the velocity toward the safe direction, but also pushes it away from the unsafe direction. Experimental results show that VESFlow+ removes the target concept, reducing the attack success rate by NudeNet to 6.3% on Ring-A-Bell and 6.8% on MMA-Diffusion on the 4-step MeanFlow model, while preserving fidelity on benign prompts.
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture. In this work, we ask whether safety can be represented as a portable latent direction, learned once and reused across heterogeneous generators. We introduce the first framework for cross-model safety steering, in which a safety direction is estimated in a source LLM from paired safe-unsafe prompts, transported to a target generator through a lightweight alignment fitted on benign data alone, and applied at inference time. Crucially, our pipeline never accesses unsafe data on the target side, isolating whether safety can be transferred through shared representation geometry. Beyond a single global direction, we also identify a multi-vector extension that captures category-specific safety behaviors, enabling more selective control. We evaluate our approach in text-to-image and text-to-video generation across diverse source-target model pairs. Across models, transferred safety directions achieve ASR reduction and CLIP-Score/FID trade-offs comparable to directions learned natively on the target model using unsafe data, while requiring no target-side unsafe data. This indicates that safety improvements do not come at the expense of generation quality. Our results point to a modular view of safety: safety-relevant behavior is not purely model-local, but can be controlled through latent directions that persist across models. This suggests a new path toward lightweight, reusable safety mechanisms that do not require target-side unsafe data.
Tobia Poppi, Silvia Cappelletti, Sara Sarto +5
University of Modena and Reggio Emilia · University of Pisa · Amazon Prime Video