While representation alignment with self-supervised models has been shown to improve diffusion model training, its potential for enhancing inference-time conditioning remains largely unexplored. We introduce Representation-Aligned Guidance (REPA-G), a framework that leverages these aligned representations, with rich semantic properties, to enable test-time conditioning from features, in generation. By optimizing a similarity objective (the potential) at inference, we steer the denoising process toward a conditioned representation extracted from a pre-trained feature extractor. Our method provides versatile control at multiple levels of granularity, ranging from patch level matching via single patches to broad semantic guidance using global image feature tokens. We further extend this to multi-concept composition, allowing for the faithful combination of distinct concepts. REPA-G operates entirely at inference time with no additional training required, offering a flexible and precise alternative to often ambiguous text prompts or coarse class labels. Our approach achieves high-quality, diverse generations on ImageNet and COCO. Code is available at https://github.com/valeoai/REPA-G
Figures & tables
Figure 1: Comparing class-label, text-prompt, and REPA-G conditioning. All models are trained on ImageNet. (Top) We average extracted features ( ) from a anchor image to generate a generic “rabbit” image. (Bottom) We combine a masked feature map with a specific “lava” patch ( ) to synthesize a rabbit on a volcano. While text prompts require lengthy descriptions and often lack precision, our feature-based conditioning offers better compositional control and provides more precise generation.
Figure 2: Impact of representation alignment on feature space. We perform k -means clustering ( k=1,000 ) on four feature spaces across ImageNet (see App. C.1 ). For a reference image (with red frame), we visualize others from its assigned cluster. Without alignment, SiT fails to form semantic groupings, making its latent space unsuitable for conditioning. In contrast, the aligned model successfully replicates the teacher’s semantic structure both before and after the projection layer, resulting in semantically consistent clusters.
Figure 3
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT
-
45.02
9.02
22.33
0.50
0.63
45.02
9.02
22.33
0.50
0.63
45.02
9.02
22.33
0.50
0.63
FSiT
71.35
72.08
14.73
0.33
0.54
45.85
19.23
29.36
0.40
0.62
65.99
22.52
16.69
0.36
0.52
REPA
-
26.58
6.85
41.31
0.56
0.70
26.58
6.85
41.31
0.56
0.70
26.58
6.85
41.31
0.56
0.70
FDINO
7.23
8.86
199.69
0.65
0.69
14.17
10.59
135.08
0.58
0.64
29.06
16.79
83.85
0.46
0.69
FSiT
2.09
6.17
260.97
0.74
0.70
2.67
4.86
222.71
0.74
0.69
6.26
6.14
159.14
0.68
0.70
Table 1: Distribution-level comparison of REPA-G generations on ImageNet [ 33 ] . We compare the standard SiT backbone [ 6 ] (no representation alignment) against REPA [ 7 ] and REPA-E [ 8 ] variants. For each method block, the first row denotes results for unconditional generation, serving as a baseline, and FSiT and FDINO for REPA-G.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
DINOv2
JEPA
CLIP
PSNR
DINOv2
JEPA
CLIP
PSNR
DINOv2
JEPA
CLIP
PSNR
SiT
FSiT
0.27
0.36
0.43
15.35
0.44
0.48
0.46
17.98
0.12
0.35
0.83
7.74
REPA
FDINO
0.75
0.61
0.58
11.02
0.77
0.60
0.57
10.63
0.76
0.69
0.92
8.00
FSiT
0.85
0.78
0.71
20.46
0.87
0.79
0.71
20.72
0.84
0.84
0.95
10.82
REPA-E
FDINO
0.83
0.69
0.64
15.23
0.85
0.68
0.63
14.36
0.91
0.83
0.96
10.97
FSiT
0.86
0.76
0.71
18.68
0.88
0.78
0.71
19.07
0.91
0.88
0.96
11.96
Table 2: Instance-level comparison of REPA-G generations on ImageNet [ 33 ] . Setup is the same as in Table 1 . Alignment is measured with DINOv2 [ 12 ] , JEPA [ 36 ] , and CLIP [ 37 ] feature spaces, supplemented by PSNR for pixel-level fidelity. We highlight ours conditioning methods with FSiT and FDINO .
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT
-
45.13
30.15
22.33
0.42
0.54
45.13
30.15
22.33
0.42
0.54
45.13
30.15
22.33
0.42
0.54
FSiT
77.26
108.79
9.86
0.25
0.37
51.88
42.36
15.76
0.27
0.52
66.09
46.88
16.67
0.28
0.45
REPA
-
37.85
29.11
41.31
0.45
0.58
37.85
29.11
41.31
0.45
0.58
37.85
29.11
41.31
0.45
0.58
FDINO
12.96
29.05
31.11
0.55
0.57
20.51
31.04
26.15
0.45
0.53
32.04
39.05
20.59
0.35
0.53
FSiT
6.18
25.35
33.08
0.65
0.62
6.63
24.37
32.05
0.65
0.61
9.46
25.89
29.43
0.60
0.60
Table 3: Distribution-level zero-shot evaluation on the COCO [ 34 ] dataset. The setup is the same as in Table 1 . Within each block, the first row reports unconditional generation results as a baseline, while FSiT and FDINO denote our method.
Figure 5: Qualitative comparison on ImageNet [ 33 ] using REPA-E [ 8 ] with DINOv2 features. (Left) We show single-source conditioning across three levels of granularity: full, masked, and averaged feature maps given anchor image and input mask . (Right) We illustrate multi-source composition. Within each group, the anchor object and target background are shown at a small scale alongside the enlarged generation. We use full features from the anchor and a single feature patch sampled from the target’s background.
Table 8Table 9
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Representation alignment with diffusion in 2D space. Ground-truth samples are shown in grey, with corresponding samples generated via diffusion in red. The conditional results in the second and third columns are guided by target feature vectors [−1,0] and [0,−1] , respectively. The last column shows the feature space mapping, mapping each data points to its corresponding 2D feature.
Figure 7: Impact of representation alignment on SiT’s feature space. We compare clusters across ImageNet [ 33 ] using four feature spaces. For each space, we extract features for the entire ImageNet dataset and perform k -means clustering into 1,000 discrete clusters. Then, we identify the specific cluster associated with a given reference image (indicated by the red frame) and randomly sample other images from that same cluster. Without alignment, the standard SiT model fails to form semantic clusters; features that are close in latent space represent unrelated concepts, which makes these internal representations unsuitable for conditioning. The aligned model successfully mimics the feature space of the visual teacher backbone both before and after the projection layer. This results in semantically consistent clusters. SiT (Align) refers to REPA [ 17 ] features extracted before the projection layer, whereas SiT (Align + Proj) denotes features extracted after the projection layer.
λ
Full
Masked
Average
FID ( ↓ )
Prec. ( ↑ )
Rec. ( ↑ )
Align. ( ↑ )
FID ( ↓ )
Prec. ( ↑ )
Rec. ( ↑ )
Align. ( ↑ )
FID ( ↓ )
Prec. ( ↑ )
Rec. ( ↑ )
Align. ( ↑ )
5000
3.3
0.72
0.67
0.6
3.7
0.72
0.66
0.7
3.2
0.76
0.64
0.9
10000
2.8
0.75
0.67
0.7
2.8
0.74
0.66
0.8
2.8
0.76
0.65
0.9
20000
1.6
0.76
0.69
0.8
2.4
0.75
0.66
0.8
2.7
0.75
0.67
0.9
50000
1.4
0.76
0.69
0.8
2.3
0.75
0.67
0.9
3.3
0.73
0.68
0.9
100000
1.5
0.75
0.70
0.9
2.6
0.74
0.68
0.9
4.5
0.71
0.69
0.9
Appendix
Table 8: Guidance scale ( λ ) robustness. FID and DINOv2 Alignment scores for single-source conditioning under various values of the guidance scale λ .
Figure 8: Sensitivity plots illustrating the relationship between λ , temperature, and output quality.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT
-
45.02
9.02
22.33
0.50
0.63
45.02
9.02
22.33
0.50
0.63
45.02
9.02
22.33
0.50
0.63
FSiT
71.35
72.08
14.73
0.33
0.54
51.88
42.36
15.76
0.27
0.52
65.99
22.52
16.69
0.36
0.52
REPA
-
26.58
6.85
41.31
0.56
0.70
26.58
6.85
41.31
0.56
0.70
26.58
6.85
41.31
0.56
0.70
FDINO
7.23
8.86
199.69
0.65
0.69
14.17
10.59
135.08
0.58
0.64
29.06
16.79
83.85
0.46
0.69
FSiT
2.09
6.17
260.97
0.74
0.70
2.67
4.86
222.71
0.74
0.69
6.26
6.14
159.14
0.68
0.70
Appendix
Table 9: Distribution-level comparison of unconditional REPA-G generations on ImageNet [ 33 ] . We compare the standard SiT backbone [ 6 ] (trained without representation alignment) against REPA [ 7 ] and REPA-E [ 8 ] variants. All models are trained on ImageNet. For each method block, the first row denotes metrics for unconditional generation, serving as a baseline for the REPA-G results. FSiT∗ denotes features extracted after the projection layer in REPA-based models. Gray results are from Table 1 in the main text for ease of comparison.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Feat .
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
SiT
FSiT
0.27
0.36
0.43
0.93
0.78
15.35
0.44
0.48
0.46
0.95
0.82
17.98
0.12
0.35
0.83
0.99
0.95
7.74
REPA
FDINO
0.75
0.61
0.58
0.94
0.86
11.02
0.77
0.60
0.57
0.94
0.84
10.63
0.76
0.69
0.92
0.99
0.98
8.00
FSiT
0.85
0.78
0.71
0.98
0.95
20.46
0.87
0.79
0.71
0.98
0.94
20.72
0.84
0.84
0.95
0.99
0.990
10.82
FSiT∗
0.72
0.63
0.59
0.94
0.87
11.32
0.74
0.63
0.58
0.94
0.86
11.02
0.77
0.72
0.92
0.99
0.98
8.01
REPA-E
FDINO
0.83
0.69
0.64
0.96
0.90
15.23
0.85
0.68
0.63
0.95
0.88
14.36
0.91
0.83
0.96
0.99
0.99
10.97
Appendix
Table 10: Instance-level comparison of unconditional REPA-G generations on ImageNet [ 33 ] . Setup is the same as in Table 1 . Alignment is measured with DINOv2 [ 12 ] , JEPA [ 36 ] , CLIP [ 37 ] , MAE [ 56 ] and MoCov3 [ 57 ] feature spaces, supplemented by PSNR for pixel-level fidelity. FSiT∗ denotes features extracted after the projection layer in REPA-based models. Gray results are from Table 1 in the main text for ease of comparison.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
DINOv2
JEPA
CLIP
MAE
MoCo
PSNR
SiT
FSiT
0.25
0.35
0.4
0.93
0.76
15.18
0.4
0.45
0.44
0.95
0.82
18.56
0.11
0.38
0.81
0.99
0.95
7.66
REPA
FDINO
0.737
0.603
0.564
0.935
0.847
10.718
0.746
0.561
0.553
0.933
0.835
10.846
0.736
0.679
0.900
0.992
0.979
7.724
FSiT
0.84
0.78
0.7
0.98
0.95
20.61
0.84
0.76
0.69
0.98
0.94
21.38
0.8
0.82
0.92
0.99
0.99
10.47
FSiT∗
0.697
0.627
0.564
0.939
0.866
11.007
0.705
0.593
0.556
0.938
0.857
11.275
0.732
0.701
0.902
0.993
0.982
7.727
REPA-E
FDINO
0.821
0.677
0.626
0.951
0.892
14.776
0.838
0.642
0.616
0.949
0.882
14.715
0.9
0.821
0.943
0.996
0.989
10.663
Appendix
Table 11: Zero-shot evaluation on COCO [ 34 ] dataset using instance-level metrics. Alignment is measured with DINOv2 [ 12 ] , JEPA [ 36 ] , CLIP [ 37 ] , MAE [ 56 ] and MoCov3 [ 57 ] feature spaces, supplemented by PSNR for pixel-level fidelity. FSiT∗ denotes features extracted after the projection layer in REPA-based models.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT
-
10.17
8.53
124.93
0.67
0.67
10.17
8.53
124.93
0.67
0.67
10.17
8.53
124.93
0.67
0.67
FSiT
71.05
49.21
17.42
0.29
0.58
62.65
13.98
19.82
0.34
0.65
130.14
53.16
2.08
0.50
0.04
REPA
-
5.92
5.92
157.10
0.70
0.68
5.92
5.92
157.10
0.70
0.68
5.92
5.92
157.10
0.70
0.68
FDINO
6.24
8.43
197.64
0.69
0.64
13.68
10.21
127.72
0.61
0.60
27.94
15.32
84.04
0.49
0.66
FSiT
2.14
6.71
254.99
0.77
0.65
3.34
5.68
199.55
0.77
0.63
11.24
8.59
122.54
0.63
0.67
Appendix
Table 12: Distribution-level comparison of class-conditional REPA-G generations on ImageNet [ 33 ] . We compare the standard SiT backbone [ 6 ] (trained without representation alignment) against REPA [ 7 ] and REPA-E [ 8 ] variants. All models are trained on ImageNet. For each method block, the first row denotes metrics for unconditional generation, serving as a baseline for the REPA-G results. FSiT∗ denotes features extracted after the projection layer in REPA-based models.
Full Feature Map
Masked Feature Map
Average Feature Map
Model
Cond.
DINO
JEPA
CLIP
MAE
MoCo
PSNR
DINO
JEPA
CLIP
MAE
MoCo
PSNR
DINO
JEPA
CLIP
MAE
MoCo
PSNR
SiT
FSiT
0.3
0.39
0.43
0.94
0.78
14.79
0.4
0.44
0.44
0.95
0.81
17.66
0.13
0.37
0.83
0.99
0.95
8.64
REPA
FDINO
0.737
0.593
0.574
0.935
0.848
10.467
0.757
0.583
0.564
0.931
0.832
10.303
0.753
0.678
0.917
0.99
0.977
8.047
FSiT
0.84
0.77
0.7
0.98
0.94
19.6
0.86
0.77
0.7
0.98
0.94
20
0.79
0.8
0.93
0.99
0.99
10.36
FSiT∗
0.706
0.613
0.578
0.938
0.862
10.714
0.727
0.608
0.570
0.935
0.849
10.595
0.758
0.696
0.922
0.99
0.980
8.024
REPA-E
FDINO
0.821
0.675
0.634
0.952
0.892
14.785
0.839
0.656
0.620
0.947
0.873
13.835
0.882
0.788
0.949
0.99
0.985
10.340
Appendix
Table 13: Distribution-level comparison of class-conditional REPA-G generations on ImageNet [ 33 ] . We compare the standard SiT backbone [ 6 ] (trained without representation alignment) against REPA [ 7 ] and REPA-E [ 8 ] variants. All models are trained on ImageNet. For each method block, the first row denotes metrics for unconditional generation, serving as a baseline for the REPA-G results. FSiT∗ denotes features extracted after the projection layer in REPA-based models.
Figure 9: Effect of guidance strength on generation diversity. Impact of λ on LPIPS distance, DINOv2 cosine distance, and DINOv2-based Vendi score for samples generated from ImageNet and COCO anchors.
Figure 10: Qualitative results of generation diversity. ImageNet examples are shown in the top row and COCO examples in the bottom row. For each example, the anchor image is shown on the left, followed by three samples generated with average feature conditioning and three with full feature map conditioning. We use a small guidance scale of λ=2,000 .
Method
Features
FID
sFID
IS
Prec.
Rec.
DINOv2
JEPA
CLIP
PSNR
REPA-G
Full
1.4
4.1
264.7
0.76
0.69
0.83
0.69
0.64
15.2
Masked
2.3
4.8
212.2
0.75
0.67
0.85
0.68
0.63
14.4
Average
3.2
5.2
188.9
0.73
0.67
0.91
0.83
0.96
11.0
TFG
Full
74.6
53.6
47.4
0.26
0.40
0.91
0.43
0.54
10.0
Masked
85.8
47.2
36.0
0.24
0.38
0.94
0.45
0.55
10.0
Average
81.1
37.0
35.8
0.23
0.41
0.99
0.51
0.91
8.5
Appendix
Table 14: TFG performance. Evaluation of TFG with different guidance granularities.
ImageNet
COCO
Condition
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
DINOv2 ↑
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
DINOv2 ↑
-
15.07
4.46
55.14
0.65
0.69
-
36.23
29.98
55.13
0.51
0.59
-
IP-Adapter
7.52
10.65
145.36
0.71
0.63
0.62
9.90
32.23
30.45
0.62
0.56
0.61
REPA-G
1.45
4.07
264.69
0.76
0.69
0.83
4.61
23.27
35.97
0.67
0.64
0.82
Appendix
Table 15: Comparison with trained adapter-based conditioning. We compare full-feature-map conditional generation using IP-Adapter and REPA-G on ImageNet and COCO, with REPA-E as the shared base model.
Figure 11: Visual comparison with the training-free baseline TFG [ 31 ] and the trained IP-Adapter [ 27 ] . While the images generated with TFG respect the conditioning to the anchors, perceptual artifacts appear and degrade image quality. On the other hand, IP-adapter manages to generate images faithful to the anchor but requires additional training and parameters.
Figure 12: Qualitative results for resolution scalability. Using the fine-tuned 512 × 512 model, REPA-G successfully stochastically reconstruct the conditioning image at a 512 × 512 resolution. The top row shows the anchor images, and the bottom row shows the corresponding results for REPA-G (512x512).
FID
sFID
IS
Prec.
Rec.
DINOv2 Align.
Uncond
33.1
10.1
39.1
0.61
0.6
-
Full
5.4
6.0
178.8
0.75
0.67
0.61
Masked
6.9
5.9
149.1
0.73
0.65
0.68
Average
8.9
5.6
139.9
0.73
0.66
0.85
Appendix
Table 16: Resolution scalability. We evaluate our method’s ability to scale to higher resolutions (512 × 512) and compare the performance across unconditional generation with different granularity settings for our feature conditioning.
Image Resolution
Method
Throughput (img/s)
Peak GPU memory (GB)
256
Uncond.
0.50
8
Ours (Single source)
0.37
10
Ours (Multi source)
0.37
10
TFG
0.21
52
CFG
0.26
8
IP-Adapter
0.19
9
Appendix
Table 17: Inference throughput. Feature conditioning requires backpropagating gradients through the initial blocks of the transformer. While this results in a slight reduction in throughput, the computational overhead remains minimal. The baseline is the standard class-conditioned REPA-E model [ 8 ] , whereas our approach incorporates feature guidance into this framework.
Figure 14: Qualitative results on ImageNet dataset [ 33 ] . We present randomly selected samples from various conditioning modes. The first row shows the ground-truth (anchor) images, followed by their unsupervised foreground masks. Subsequent rows show generations using the full feature map, the masked feature map (using the corresponding mask), and the averaged feature map. All visualizations are from the following configuration: an unconditioned model guided by features from the diffusion model before the projection layer. The following block displays the corresponding results for standart SiT [ 6 ] model trained without representation alignment and REPA-E [ 8 ] . We observe that SiT models trained without representation alignment struggle to produce coherent results under feature guidance.
Figure 15: Qualitative results on COCO dataset [ 34 ] . We present randomly selected samples from various conditioning modes. The first row shows the ground-truth (anchor) images, followed by their unsupervised foreground masks. Subsequent rows show generations using the full feature map, the masked feature map (using the corresponding mask), and the averaged feature map. All visualizations are from the following configuration: an unconditioned model guided by features from the diffusion model before the projection layer. The following block displays the corresponding results for standart SiT [ 6 ] model trained without representation alignment and REPA-E [ 8 ] . We observe that SiT models trained without representation alignment struggle to produce coherent results under feature guidance.
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet 256×256 study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Sihan Xu, Ji Xie, Zilin Wang +2
University of Michigan · Carnegie Mellon University
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3× faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Ning Zhu, An Chen, Mengfei Zhao +4
Glasgow College, University of Electronic Science and Technology of China · School of Mathematical Sciences, University of Electronic Science and Technology of China
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.