Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.
Figures & tables
conditioning
reconstruction MAE ↓
raw change
colour share
translation ↑
Sana DC-AE latent (32 × 32 × 32)
0.042
0.583
43 %
0.331
FLUX VAE latent (16 × 128 × 128)
0.061
0.404
41 %
0.239
DINOv2-reg (768 × 32 × 32)
0.113
0.516
−11 %
0.572
Table 1: The identical adapter and recipe on three conditionings. Raw change is how much the adapter changes its input; translation is that change with a per-channel affine colour map factored out, so a global re-grade does not count; colour share is the part of the raw change that the affine map explains. A negative share means the change is not a global colour shift at all.
Figure 1: Training is one-sided by default. Top row: the frozen encoder maps the frame to a feature map, G adapts it, the frozen decoder renders it, and the target discriminators judge the render against real frames, their gradient reaching G through the decoder. Bottom row: F maps target features back and learns only from the cycle and identity terms. Dashed: the symmetric variant (Section 4.3 ). At inference only the top row runs, with the decoder at two steps.
condition / method
KID ↓
FID ↓
DINO ↓
MAE ↓
step
ACDC fog → clear
do nothing
0.1066
166.73
0.0000
0.1761
—
decoder only (identity)
0.1096
169.89
0.0140
0.1834
—
RFA
0.0262
116.68
0.0213
0.1766
5000
CycleGAN-Turbo
0.0478
127.19
0.0639
0.1605
25000
CycleGAN
0.0333
130.73
0.0462
0.1451
400 ep
Table 2: Every method at the end of its own recipe, or as released (last column). CycleGAN-Sprint is CycleGAN-Turbo rebuilt on our Sana-Sprint backbone (Section 7 ). Fifty crops per condition, 25 for haze; haze and night are unpaired, so no MAE. DINO is structure distance. Bold best, underline second best.
Figure 2: The weather benchmark (Table 2 ). Last column: the paired clear frame where ACDC has one, else a target frame. On rain the wet sheen survives every pixel-fed method but Cosmos-Transfer’s edge control, which redraws the scene, and only the RFA removes the rain while keeping the scene. CycleGAN-Turbo is our training with its released recipe per condition, except haze (the released module of Henein et al. (2026) ) and night (the released night-to-day model).
PreSIL → Mapillary
KID ↓
FID ↓
DINO ↓
checkpoint
do nothing
0.0470
135.64
0.0000
—
REGEN, GTA → Vistas
0.0346
123.94
0.0081
released
HyPER-GAN, GTA → Vistas
0.0331
125.35
0.0080
released
REGEN, GTA → Cityscapes (other target)
0.0237
116.73
0.0110
released
RFA (ours)
0.0123
107.55
0.0318
5000
RFA (ours, DINOv3-S, 0.6 B decoder)
0.0248
116.85
0.0261
5000
Table 3: Sim-to-real on 100 matched PreSIL centre crops at 1024 2 . KID and FID against 2 000 Mapillary crops; DINO is structure distance to the source crop (protocol details in Appendix B.1 ). DINOv3-S row: the DINOv3-S configuration (Appendix D.2 ). Bold best, underline second best.
Figure 3: Sim-to-real on one matched crop (Section 6.2 ); more in Figure 6 .
Figure 4: Portability. PiD (pixel diffusion, four steps) and RAE (one deterministic pass), two decoders of DINOv2 feature maps from other groups, decode the unadapted PreSIL features (“raw”) and the same features after the sim-to-real RFA. Neither decoder saw the RFA, and the RFA saw neither. Three more frames in Figure 7 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
term
weight
role
cycle
30.0
F(G(x)) ≈ x and G(F(y)) ≈ y
identity
10.0
G(y) ≈ y for y already in target; the direct read on drift
structure
20.0
patch self-similarity; holds scene geometry
image GAN
1.0
RGB PatchGAN on the decoded image
vision-aided
0.2
frozen DINOv2 features with a small trained head ( Kumari et al., 2022 )
R1 penalty
0.001
gradient penalty on real images
Appendix
Table 4: Loss weights.
backbone
Sana-Sprint 1.6 B, 1024 px, bf16, frozen
decoder
ControlNet on DINOv2-B/reg (PiD normalisation, 448 px, 32×32 map), Mapillary,
7 blocks, checkpoint 15 000, frozen ; DINOv3-S configuration: 0.6 B at 1024 px
optimiser
lr 2e-4 ( G and D ), constant with warmup, 500 warmup steps
schedule
5 000 steps, batch 4, gradient checkpointing
discriminator
vision-aided DINOv2, R1 0.001
augment
horizontal/vertical flip, ±15∘ rotation with a black-corner-free crop
Appendix
Table 5: Training configuration of every RFA.
CycleGAN-Turbo
CycleGAN-Sprint
RFA
trainable weights
470 M (LoRA r = 128)
148 M (LoRA r = 128)
2.9 M (G + F)
steps
25 000
25 000
5 000 ( ≈ 1 800 to converge)
per step
1.86 s
2.40 s
1.79 s
per condition
13.6 h
16.7 h
2.4 h
one-time prerequisite
none
none
decoder, 8–19 h
all four conditions
54.4 h
66.7 h
17.6–28.6 h
Appendix
Table 6: Training cost on a single server-grade GPU. CycleGAN-Turbo as released (SD-Turbo, 512 px) and CycleGAN-Sprint (Sana-Sprint, rank 128, 512 px, batch 1; Appendix D.3 ), its 3.7 M CLIP discriminator heads excluded as in the original; the RFA at 1024 px, batch 4. Per-condition times are wall-clock. Per-step cost is of the same order everywhere; the RFA’s saving is the step count, and it pays once for the decoder (8 to 19 h), which every condition and both directions share.
adapter move
distance to another frame
condition
n
adverse input
clear input
paired clear
same weather
clear
ACDC fog
50
0.141 ± 0.012
0.027
1.18
1.01
1.22
ACDC snow
50
0.136 ± 0.010
0.026
1.16
1.05
1.23
ACDC rain
50
0.132 ± 0.029
0.027
1.06
1.14
1.21
HIVIS haze
25
0.131 ± 0.049
0.059
—
1.21
1.24
BDD night
50
0.128 ± 0.023
0.036
—
1.16
1.22
Appendix
Table 7: The residue in numbers, the adapters and test crops of Table 2 : ∥G(f)−f∥/∥f∥ per frame (mean and standard deviation) on adverse input and on clear input, and, in the same units, the distance from the adverse input to the paired clear photograph of the scene (paired conditions) and to an unrelated frame of the same test set, of the same weather and clear.
Figure 5: The weather benchmark on two further frames per condition, none of them the frames of Figure 2 ; columns and references as there.
Figure 6: Sim-to-real on three further matched crops. The pixel-fed GANs keep identity and change appearance; the RFA changes what the features leave free.
Figure 7: Portability on three further PreSIL frames: PiD and RAE decoding raw features and the same features after the sim-to-real RFA.
method
FID ↓
DINO-Struct ↓
RFA (ours)
50.1
2.30
RFA (ours), no patch discriminator
50.7
2.34
CycleGAN
74.9
3.22
CUT
43.9
6.58
CycleGAN-Turbo, retrained ∗
49.0
2.65
∗ paper result ( Parmar et al., 2024 ) : 41.0 / 2.1.
Appendix
Table 8: Horse → zebra on CycleGAN-Turbo’s protocol: FID of the 120 translated test horses against the 140 test zebras at 256 px and DINO-Struct × 100 between input and output, lower is better. Bold best.
Figure 8: Horse → zebra, four test horses: the released CycleGAN and CUT models, CycleGAN-Turbo trained with the authors’ recipe. The two RFA runs compare the default recipe with training without the patch discriminator.
Figure 9: The 0.6 B DINOv3-S ControlNet on one held-out frame: the 20-step teacher it was trained against and the 2-step student it runs on. The student reconstructs better, 19.5 against 18.0 dB on the 50 frames of Table 9 scored at the 1024-px output, as does the 1.6 B pair (17.8 against 17.0); Table 9 scores the same renders at 512 px.
decoder
features
PSNR ↑
gradNCC ↑
RAE, one pass (ViT-XL)
DINOv2-B/reg, 32 2
20.14
0.557
Sana ControlNet, two steps (RFA decoder)
DINOv2-B/reg, 32 2
18.21
0.519
Sana ControlNet, two steps (DINOv3-S configuration)
DINOv3-S at 1024 px, 32 2
20.10
0.650
DC-AE round trip (ceiling, no features)
pixels
28.83
0.923
Appendix
Table 9: Reconstruction from the conditioning alone: 50 held-out Mapillary frames, the same DINO feature map decoded by each decoder, scored against the input at 512 px. RAE was trained on ImageNet, never on driving frames; our ControlNets on Mapillary and run at their shipped two steps. The DC-AE row is the autoencoder round trip of the ground truth, the ceiling for any render through it. Bold best of the three decoders.
adapter trained through
native
KID / FID, 512 px
KID / FID, native
DINO ↓
our Sana ControlNet (DINOv2-B/reg)
1024
0.0117 / 107.6
0.0123 / 107.6
0.0318
RAE (DINOv2-B/reg)
512
0.0198 / 114.7
0.0198 / 114.7
0.0265
RAEv2, DINOv3-S, last layer
256
0.0472 / 130.4
0.0336 / 118.2
0.0313
RAEv2, DINOv3-L, 23-layer sum
256
0.0277 / 116.8
0.0201 / 108.9
0.0192
Appendix
Table 10: Sim-to-real adapters trained through four frozen decoders on the 100 PreSIL crops of Table 3 : the paper recipe of Table 5 in every row, only the decoder changes. KID and FID at a common 512 px against the 2 000 Mapillary crops at that size, then at each decoder’s native output size against references at that size (comparable only within a row’s size), and structure distance to the source crop.
Figure 10: One adapter recipe, four frozen decoders (Table 10 ). The scene survives in all four; what differs is fine content and how much of the layout moves. Decoders render at 1024, 512 and 256 px, printed at one size.
Figure 11: The frames behind Table 9 .
Figure 12: Decoder bias, the frames of Figure 2 . Decoder only is the frozen decoder on the unadapted features. The weather, and the night, come back in every decoder-only frame, so the removal in the RFA row is the adapter’s.
method
fog
snow
rain
do nothing
0.00
0.00
0.00
decoder only (identity)
0.00
0.00
0.02
RFA
0.68
1.00
1.00
RFA, DINOv3-S configuration
0.82
1.00
0.96
CycleGAN-Turbo
0.84
1.00
0.46
CycleGAN-Sprint
0.02
1.00
0.36
Appendix
Table 11: Share of test outputs a classifier trained on the ACDC training crops calls clear, per method and condition, 50 frames each.
BDD100K snow
BDD100K rain
PreSIL → Mapillary
method
mAP
mAP 50
mAP 75
mAP
mAP 50
mAP 75
mAP
mAP 50
mAP 75
do nothing
0.276
0.494
0.227
0.254
0.442
0.252
1.000
1.000
1.000
RFA
0.148
0.262
0.146
0.146
0.268
0.142
0.153
0.331
0.158
CycleGAN-Turbo
0.275
0.474
0.320
0.224
0.427
0.200
—
—
—
CycleGAN
0.270
0.467
0.241
0.221
0.408
0.203
—
—
—
CUT
0.229
0.410
0.177
0.212
0.374
0.218
—
—
—
Appendix
Table 12: Downstream detection with a fixed COCO Faster R-CNN at 512 px. Weather: BDD100K snow and rain frames with official boxes, 300 each, ACDC-trained models applied unchanged. Sim-to-real: the 100 PreSIL crops of Table 3 , ground truth replaced by the detector’s own confident boxes on the untranslated render, so doing nothing scores 1.0 by construction. Higher is better.
KID
FID
condition
n
RFA
baseline
difference
RFA
baseline
difference
ACDC fog
50
0.027 ± 0.006
0.047 ± 0.007
− 0.020 ± 0.008
116.4 ± 3.3
127.2 ± 3.7
− 10.7 ± 4.6
ACDC snow
50
0.021 ± 0.005
0.024 ± 0.004
− 0.002 ± 0.005
119.6 ± 3.5
117.7 ± 3.4
+ 1.9 ± 3.1
ACDC rain
50
0.016 ± 0.004
0.016 ± 0.004
− 0.001 ± 0.005
120.6 ± 3.9
128.7 ± 7.1
− 8.0 ± 5.4
HIVIS haze
25
0.038 ± 0.011
0.037 ± 0.012
0.000 ± 0.006
165.3 ± 10.6
174.7 ± 15.5
− 9.3 ± 9.7
BDD night
50
0.013 ± 0.004
0.023 ± 0.006
− 0.008 ± 0.005
128.0 ± 2.6
130.4 ± 3.2
− 2.4 ± 2.6
Appendix
Table 13: Leave-one-out jackknife standard errors of KID and FID for the RFA and its main baseline (CycleGAN-Turbo on the weather conditions, REGEN and HyPER-GAN on sim-to-real), and of the paired difference RFA minus baseline on the same crops. Separate scoring pass from Table 2 .
We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike prior methods that train separate adapters for each conditioning modality, UNITY jointly learns shared semantics across multiple conditioning types and subsequently specializes without modifying the underlying architecture. The proposed two stage training paradigm consists of a Universal Stage that captures cross modal representations across all conditioning modalities using half of the total training steps, followed by a Specialization Stage that refines modality specific features using the remaining training budget. At the core of UNITY are the Morphable Attention Flow (MAF) Network and Morph Wrapper modules, which enable channel aware and spatially adaptive feature alignment through learnable flow fields and attention based fusion. This constant complexity formulation supports flexible operation under both single and composite conditioning settings while significantly reducing inference latency and memory consumption. Extensive experiments across multiple datasets demonstrate that UNITY achieves state of the art image fidelity while maintaining superior memory efficiency. Code: https://github.com/arya-domain/UNITY
Aryan Das, Koushik Biswas, Moloud Abdar +1
VIT Bhopal, India · IIIT Delhi, India · The University of Queensland, Australia +1
Pre-trained image restoration models often fail on out-of-distribution (OOD) real-world degradations. Adapting to these domains is challenging as real-world data lacks paired ground truth, and unsupervised methods often require unstable architectural changes. We propose Generative Manifold Distillation (GMD), which reframes domain adaptation as geometric manifold alignment. GMD operates in a strictly unpaired setting, requiring only low-quality (LQ) target observations. By leveraging the flow-matching dynamics of a frozen text-to-image foundation model, GMD projects off-manifold restorations onto the natural image manifold to generate high-quality pseudo-targets. To ensure stability, a quality-gated manifold filter rejects off-manifold samples, while source-anchored trajectory regularization prevents error accumulation. Ultimately, GMD distills a powerful generative prior into an efficient restoration network. Experiments demonstrate that GMD seamlessly adapts to new distributions using only LQ inputs, drastically improving perceptual quality with zero architectural modifications or added inference latency.
Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function evaluations. While convenient, this ignores the heterogeneity of natural images: some regions are easy to denoise, whereas others benefit from more refinement or additional context. Motivated by this, we explore patch-level noise scales for image synthesis. We find that naively varying timesteps across image tokens performs poorly, as it exposes the model to overly informative training states that do not occur at inference. We therefore introduce a timestep sampler that explicitly controls the maximum patch-level information available during training, and show that moving from global to patch-level timesteps already improves image generation over standard baselines. By further augmenting the model with a lightweight per-patch difficulty head, we enable adaptive samplers that allocate compute dynamically where it is most needed. Combined with noise levels varying over both space and diffusion time, this yields Patch Forcing (PF), a framework that advances easier regions earlier so they can provide context for harder ones. PF achieves superior results on class-conditional ImageNet, remains orthogonal to representation alignment and guidance methods, and scales to text-to-image synthesis. Our results suggest that patch-level denoising schedules provide a promising foundation for adaptive image generation.