Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we construct FreqBlind, a 6,000-image benchmark spanning contours, real words and non-words across three conditioning strengths. The evaluated transformation-based methods show limited recognition of contour patterns and weakly conditioned hidden content. To address this limitation, we propose ControlTrace to recover the grayscale control field used during generation. An 8.4M-parameter U-Net predicts this field from the carrier image, and a VLM then identifies its content. With Qwen2.5-VL-7B-Instruct, ControlTrace achieves 60.2% open-ended contour recognition accuracy across the three conditioning strengths, exceeding the best of the three evaluated prior methods by 26.9 percentage points. On an A100 GPU, the complete pipeline adds only 7.4 ms (5.3%) to direct VLM inference. Recovered fields have lower pixel errors and higher structural similarity than the evaluated transformation views. Across four evaluated VLMs, ControlTrace retains its overall contour recognition advantage. Recognition remains stable under the tested JPEG compression, Gaussian noise and downsampling. These results support control-field recovery for hidden-content recognition in the evaluated setting.
Figures & tables
Figure 1: Qualitative examples on the same carriers. SemVink shows its transformed view; SMSP and AVR show all four and seven views, respectively. ControlTrace predicts the control field, shown alongside the ground truth.
Figure 2: ControlTrace workflow and training losses. Reconstruction and gradient losses supervise the recovered control field; the auxiliary edge loss is used only during training. At inference, a VLM reads the recovered field with a task prompt.
Contour ↑
Word ↑
Method
All
1.0
1.5
2.0
All
1.0
1.5
2.0
Carrier, read directly
6.30
1.1
5.5
12.3
2.57
0.1
0.7
6.9
SemVink † (training-free)
19.43
2.8
18.1
37.4
53.60
9.1
68.8
82.9
SMSP † (training-free)
33.33
12.7
40.1
47.2
63.37
17.6
80.2
92.3
AVR seven-view QA †
19.37
2.9
18.5
36.7
23.03
0.6
17.1
51.4
AVR any-view coverage
45.10
19.2
49.6
66.5
43.00
5.8
44.1
79.1
Table 1: Open-ended recognition on FreqBlind (%), with 1,000 carriers per domain–strength cell. Trained methods average three seeds; † denotes our implementations. AVR QA uses fixed plurality voting; AVR any-view coverage is a reference-scored seven-answer diagnostic. Bold numbers mark the highest single-answer accuracy.
Contour
Word
Method
MSE ↓
MAE ↓
SSIM ↑
MSE ↓
MAE ↓
SSIM ↑
Carrier
0.1049
0.2732
0.1259
0.1210
0.2966
0.0621
SemVink
0.1038
0.2859
0.1276
0.1168
0.3042
0.0631
SMSP (view mean)
0.1172
0.2960
0.1193
0.1401
0.3247
0.0583
AVR (view mean)
0.2579
0.3832
0.1710
0.2678
0.3935
0.1346
Gaussian blur ( σ=16 )
0.1252
0.3234
0.1096
0.1278
0.3254
0.0479
Table 2: Control-field fidelity on FreqBlind ( 3,000 carriers per domain). Trained methods average three seeds; SMSP and AVR average four and seven view scores per carrier after fixed geometric mapping. Bold marks the best point estimate in each column.
Change
Δ accuracy (pp)
Uniform training allocation across strengths
−6.60
Remove gradient loss
−2.20
Remove auxiliary edge head
−0.27
Increase U-Net depth to five levels
−1.13
Increase U-Net depth to six levels
−1.07
Table 3: Component ablations on weak contours ( s=1.0 ), averaged over three training seeds. Values are changes in recognition accuracy, in percentage points, relative to the corresponding reference configuration. Training-data and input settings are given in Appendix C.3 .
Figure 3: Contour recognition across readers and generators. (a) Four readers on 300 weak-contour carriers each. (b) Five QR-based generator configurations, each with 900 carriers across strengths; the asterisk marks the training architecture and control encoder. ControlTrace uses one checkpoint. AVR any denotes reference-scored coverage across seven answers.
Figure 4: ControlTrace under image perturbations, using 300 weak-contour carriers and one checkpoint. (a) Accuracy changes under redistribution and band-limited noise. (b) Accuracy under PGD targeting the recovery network.
Image set
n
SemVink
SMSP
AVR
ControlTrace (Ours)
Positive images: target accuracy ↑
Hidden contours
1,000
0.70
3.90
0.10
36.30
Negative images: false-positive rate ↓
Matched, control off
1,000
2.30
0.40
0.00
3.80
Dense-texture scenes
1,000
1.30
2.00
0.40
5.20
COCO photographs
10,000
10.54
9.04
1.45
5.28
Table 4: Target recognition and false-positive reporting (%) under the same none-capable prompt. Arrows indicate the preferred direction. Positives are contours at s=1.0 ; ControlTrace uses one checkpoint. AVR selects one answer by plurality voting.
Method
Time (ms) ↓
Input tokens ↓
Output tokens ↓
Total tokens ↓
Direct VLM
138.6
384
2.97
386.97
SemVink
83.5
85
2.72
87.72
SMSP
373.8
1400
2.70
1402.70
AVR
1125.8
2688
20.01
2708.01
ControlTrace (Ours)
146.0
384
2.68
386.68
Table 5: Mean latency and VLM token use per carrier over 300 images and two repetitions. Latency includes preprocessing and answer generation. Input counts include visual tokens; total tokens sum inputs and outputs. AVR sums all seven calls.
Grayscale; Gaussian radius 4 ; normalization; 256 -bin Otsu threshold; 3×3 opening and closing, once each.
Inverse-mask
255−M , using the cleaned mask M above.
Foreground
Original RGB foreground selected by M over RGB (128,128,128) .
Appendix
Table 6: Fixed parameters of our AVR seven-view implementation, in evaluation order. Values refer to a longest image side of 512 pixels. Gaussian and box values are Pillow radius arguments. Each output retains the carrier’s dimensions before the reader’s usual preprocessing.
Contours ↑
Words ↑
Method
Reported
Alternative
Reported
Alternative
Carrier, read directly
6.30
6.40
2.57
2.40
SemVink
19.43
19.77
53.60
51.33
SMSP
33.33
34.17
63.37
62.90
AVR
19.37
19.77
23.03
22.37
ControlTrace (Ours)
60.10
62.00
61.23
59.90
Appendix
Table 7: Recognition accuracy (%) under the reported scoring rule and an alternative requiring the final target token to appear as a complete response token. Results aggregate the three conditioning strengths. ControlTrace uses seed 42 ; AVR uses the same plurality-selected answer under both rules. Higher values indicate better recognition.
Figure 5: Single-answer contour recognition across conditioning strengths. Each strength contains 1,000 carriers; ControlTrace averages three training seeds. Error bars show 95% item-bootstrap intervals from 20,000 resamples of 50 targets. AVR uses its fixed plurality answer. The shaded region marks the additional strengths s=1.2 – 1.4 .
SemVink quoted question
FreqBlind task question
Input / method
Objects ↑
Text ↑
Objects ↑
Text ↑
Carrier, read directly
3.6
3.6
—
—
Downscale to 32 px
21.4
28.6
21.4
25.0
Downscale to 40 px
28.6
35.7
30.4
28.6
Downscale to 64 px
28.6
50.0
21.4
50.0
Downscale to 85 px
26.8
57.1
25.0
53.6
Appendix
Table 8: HC-Bench baseline recognition accuracy (%) under two prompt regimes. The shared automatic substring/token rule includes word-spacing aliases. Each regime uses the same 56 object and 28 Latin-text images. Downscaling rows give the output width; a dash marks an unevaluated setting. Higher values indicate better recognition.
Figure 6: Real-word and non-word recognition averaged across conditioning strengths. Paired endpoints compare 25 real words with exact anagrams using the same letter multiset, font, and rendering procedure ( 1,500 carriers per target type). Unmatched benchmark non-words are shown separately. Difference intervals are 95% paired bootstraps over the 25 word pairs; ControlTrace averages three training seeds. AVR uses its plurality-selected answer; seven-answer coverage is retained in the source data.
Setting
Value
Optimizer
AdamW
Learning rate
2×10−4
Weight decay
10−4
Training batch size
24 ( 4 controls ×6 variants)
Validation batch size
16
Training duration
30 epochs
Appendix
Table 9: Implementation settings for the reference ControlTrace model. Training batches group six carrier variants from each of four control images.
Contour
Word
Method
s
MSE ↓
MAE ↓
SSIM ↑
MSE ↓
MAE ↓
SSIM ↑
Carrier
1.0
0.1381
0.3201
0.0855
0.1491
0.3319
0.0453
1.5
0.0990
0.2651
0.1322
0.1161
0.2897
0.0650
2.0
0.0777
0.2343
0.1600
0.0979
0.2683
0.0760
SemVink
1.0
0.1327
0.3270
0.1115
0.1421
0.3361
0.0546
1.5
0.0980
0.2786
0.1325
0.1119
0.2977
0.0659
Appendix
Table 10: Control-field fidelity by conditioning strength, with 1,000 carriers per domain–strength cell. The view and seed aggregation rules match Table 2 ; values are means.
Contour
Word
Method
View
MSE ↓
MAE ↓
SSIM ↑
MSE ↓
MAE ↓
SSIM ↑
SMSP
Original
0.1049
0.2732
0.1259
0.1210
0.2966
0.0621
Low-pass, 100 px
0.1557
0.3485
0.1048
0.1920
0.3872
0.0451
Low-pass, 200 px
0.1068
0.2867
0.1184
0.1289
0.3158
0.0607
Low-pass, 400 px
0.1015
0.2757
0.1279
0.1186
0.2994
0.0654
AVR
Original
0.1049
0.2732
0.1259
0.1210
0.2966
0.0621
Appendix
Table 11: Fidelity of individual SMSP and AVR views, averaged over 3,000 carriers per domain. These scores retain variation within each view bank; no view is selected using the control field.
Supervision target
Backbone
s=1.0↑
All strengths ↑
Scene at s=2.0
U-Net
13.27
29.13
Control field
NAFNet
38.13
56.51
Control field
U-Net (Ours)
45.40
60.21
Appendix
Table 12: Contour recognition for supervision-target and backbone controls. Accuracy (%) is averaged over three training seeds, with 1,000 carriers per conditioning strength. The All strengths column aggregates the three evaluated strengths. Higher values indicate better recognition.
Accuracy (%) ↑
Domain
s
Seed 42
Seed 43
Seed 44
Mean
SD (pp)
Contours
1.0
45.10
46.80
44.30
45.40
1.28
1.5
65.10
66.90
64.80
65.60
1.14
2.0
70.10
68.80
70.00
69.63
0.72
All
60.10
60.83
59.70
60.21
0.57
Words
1.0
27.40
29.60
27.80
28.27
1.17
Appendix
Table 13: Training-seed variability of ControlTrace (Ours) on FreqBlind. Each domain contains 50 target items and 1,000 carriers per strength. The seed scores and Mean column report recognition accuracy (%); SD is the sample standard deviation across the three training seeds, in percentage points. The All rows aggregate the three strengths within each seed.
Figure 7: Parameter sensitivity on weak contours. Gaussian blur and downsampling vary within the tested grids; hollow markers use 300 carriers and filled markers use 1,000 . The panels use separate vertical scales. Lines connect tested settings and do not imply an exhaustive search.
Contour ↑
Word ↑
Operation / method
All
s=1.0
All
s=1.0
C − W
Sobel (edge operator)
8.03
1.00
4.67
0.00
+3.37
Otsu (pointwise map) ( Otsu, 1979 )
12.17
2.00
18.23
0.10
−6.07
blur + histogram equalization ( Qu et al., 2025 )
29.53
10.10
56.30
13.10
−26.77
Gaussian blur σ=8
34.83
11.00
59.83
13.80
−25.00
Gaussian blur σ=16
34.33
15.00
39.17
7.50
−4.83
Appendix
Table 14: Selected transformation families. Aggregate cells use 3,000 carriers per domain; weak-condition cells use 1,000 . The final column is the contour-minus-word difference in percentage points. Method comparisons appear in Table 1 ; Table 15 lists all 27 operations.
Family and grid
Setting
Contour s=1.0↑
n
low-pass (16 operations)
Gaussian blur, σ∈{4,8,16,24,32,48}
σ=16
15.00
1,000
σ=8
11.00
1,000
σ=24
5.30
1,000
σ=4
2.33
300
σ=32
2.10
1,000
Appendix
Table 15: All 27 tested operations on contours at s=1.0 . The n column gives the sample size for the reported result; rows with different sample sizes should not be compared at subpercentage-point precision. Operations are grouped into four families. A dash denotes no parameter sweep.
Reader
s
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Qwen2.5-VL-7B
1.0
15.33
3.67
43.33
1.5
41.00
18.33
65.00
2.0
48.33
36.67
69.67
InternVL3-8B
1.0
12.67
3.00
27.67
1.5
30.67
13.00
48.33
2.0
38.33
22.33
49.00
Appendix
Table 16: Contour recognition accuracy (%) across readers, with 300 carriers per strength and the same recovery checkpoint (training seed 42 ). AVR uses plurality voting.
Reader
s
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Qwen2.5-VL-7B
1.0
14.33
1.00
28.33
1.5
79.67
16.33
76.67
2.0
93.67
49.00
80.00
InternVL3-8B
1.0
27.00
1.33
21.67
1.5
85.67
18.67
76.67
2.0
94.33
60.00
80.33
Appendix
Table 17: Word recognition accuracy (%) across readers, with 300 carriers per strength and the same recovery checkpoint (training seed 42 ). AVR uses plurality voting.
Reader
Contours ↑
Words ↑
Qwen2.5-VL-7B
78.0
80.0
InternVL3-8B
56.0
94.0
Qwen2.5-VL-32B
78.0
92.0
LLaVA-OneVision-7B
90.0
94.0
Appendix
Table 18: Direct recognition of clean control fields (%) associated with the s=1.0 subset of the reader evaluation.
Configuration
s
Direct ↑
SemVink ↑
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
SD1.5 / QR Code Monster
1.0
1.33
2.33
12.67
4.00
51.00
1.5
10.33
17.67
44.33
21.00
68.33
2.0
19.00
33.33
55.67
42.67
72.00
All
10.22
17.78
37.56
22.56
63.78
SDXL / QR Code Monster
1.0
2.00
3.00
12.33
2.00
27.67
1.5
6.33
16.33
39.00
14.33
48.00
Appendix
Table 19: Contour recognition accuracy (%) across generators, with 300 carriers per strength and one recovery checkpoint. The All rows aggregate the 900 carriers across strengths. SD1.5 and QR Code Monster also generate the training pairs; Canny is an auxiliary edge-control reference. AVR uses plurality voting.
Configuration
s
Direct ↑
SemVink ↑
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
SD1.5 / QR Code Monster
1.0
0.00
14.67
31.00
1.33
40.67
1.5
1.67
69.33
78.33
20.67
73.33
2.0
14.33
85.00
92.00
58.67
81.33
SDXL / QR Code Monster
1.0
0.00
18.33
29.67
1.67
25.33
1.5
6.33
72.33
82.00
28.00
65.33
2.0
40.67
80.00
96.00
72.00
72.67
Appendix
Table 20: Word recognition accuracy (%) across generators, with 300 carriers per strength and one recovery checkpoint. Canny is an auxiliary edge-control reference. AVR uses plurality voting.
Configuration
s
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Unperturbed
1.0
15.33
3.67
43.33
1.5
41.00
18.33
65.00
2.0
48.33
36.67
69.67
JPEG ( q=75 )
1.0
16.00
5.00
44.00
1.5
39.67
18.67
66.33
2.0
49.00
36.67
69.33
Appendix
Table 21: Contour recognition accuracy (%) under common image perturbations, with 300 carriers per strength and one recovery checkpoint. AVR uses plurality voting.
Configuration
s
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Unperturbed
1.0
14.33
1.00
28.33
1.5
79.67
16.33
76.67
2.0
93.67
49.00
80.00
JPEG ( q=75 )
1.0
15.33
0.33
29.67
1.5
78.67
15.33
77.00
2.0
94.00
48.67
79.67
Appendix
Table 22: Word recognition accuracy (%) under common image perturbations, with 300 carriers per strength and one recovery checkpoint. AVR uses plurality voting.
Configuration
s
Blur ↑
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Band noise ( 0.01 – 0.08 )
1.0
13.33
15.67
3.33
42.00
1.5
41.00
41.67
20.33
66.00
2.0
49.33
51.00
35.33
70.00
Band noise ( 0.08 – 0.25 )
1.0
13.33
17.00
5.00
43.67
1.5
39.67
40.00
20.33
66.67
2.0
48.67
51.00
40.00
70.00
Appendix
Table 23: Contour recognition accuracy (%) under perturbations and attacks, with 300 carriers per strength and one recovery checkpoint. All methods read the same perturbed carriers. AVR uses plurality voting; Blur denotes the auxiliary Gaussian-blur reference ( σ=16 ).
Configuration
s
Blur ↑
SMSP ↑
AVR ↑
ControlTrace (Ours) ↑
Band noise ( 0.01 – 0.08 )
1.0
6.33
16.67
0.33
26.00
1.5
46.00
77.33
19.33
78.67
2.0
62.67
92.33
51.67
81.33
Band noise ( 0.08 – 0.25 )
1.0
6.33
15.33
1.00
29.00
1.5
46.33
80.00
19.33
77.00
2.0
64.67
93.67
57.33
79.33
Appendix
Table 24: Word recognition accuracy (%) under perturbations and attacks, with 300 carriers per strength and one recovery checkpoint. All methods read the same perturbed carriers. AVR uses plurality voting; Blur denotes the auxiliary Gaussian-blur reference ( σ=16 ).
Image set
n
SemVink
SMSP
AVR
ControlTrace (Ours)
Positive images: content reporting
Hidden contours
1,000
3.50
9.60
0.10
61.5 [53.1, 69.8]
Negative images: false-positive rate ↓
Matched, control off
1,000
2.30
0.40
0.00
3.8 [2.4, 5.5]
Dense-texture scenes
1,000
1.30
2.00
0.40
5.2 [3.1, 7.6]
COCO photographs
10,000
10.54
9.04
1.45
5.3 [4.8, 5.7]
Appendix
Table 25: Content-reporting rates (%) under the shared none-capable prompt. Positive-image reports need not identify the correct target; reports on negative images are false positives. ControlTrace uses one checkpoint, with 95% intervals in brackets. All four methods produced valid final answers, and both rejection parsers agreed. AVR uses its plurality-selected answer.
Figure 8: Configuration-matched positive and negative carriers, with their recovered fields. Six targets were sampled with seed 20260913 , one scene per target, from 1,000 matched pairs. Each pair shares the seed, prompt, scheduler, steps and guidance, with the control branch on or off. Removing control residuals changes the generation trajectory, so the match is by configuration rather than by pixels.
Figure 10: All eight sampled targets, using the same columns as Figure 1 , which shows the first three. Diamond has missing facets, and Snake is recovered as a different shape.
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to a finite set of targets before victim training. This assumption substantially underestimates the threat. We show that a single poisoning phase can implant a programmable backdoor into a VLM, allowing an attacker to choose previously unseen target-caption semantics at inference time and synthesize corresponding stealthy triggers on demand. Unlike fixed-mapping attacks, the proposed any-to-any caption-control paradigm decouples post-training target selection from poisoning, enabling dynamic control of target captions without retraining the VLM. Our method has two components. First, a heuristic poisoning strategy exposes the model to diverse trigger-caption pairs, encouraging it to learn a general trigger-as-instruction rule rather than memorize a specific backdoor pattern. Second, a feature-space trigger steganography method maps any attacker-specified target caption to a stealthy visual trigger, implemented as either a norm-controlled perturbation or a non-semantic patch. Once inserted into arbitrary images, these triggers cause the poisoned VLM to generate outputs semantically aligned with the chosen target caption, even when the target was unseen during poisoning. Extensive experiments show that our attack achieves high any-to-any caption-control success rates, preserves clean model utility, and remains effective under several classical backdoor defenses.
Tao Lin, Gaojie Jin, Zongxin Liu +2
Key Laboratory of System Software (Chinese Academy of Sciences), Beijing, China · Institute of Software, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China +2
We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image and a global text prompt through two separate encoders, one for vision and one for language, UniVL is trained to bind semantics to spatial locations directly from a single unified visual input, where the textual instruction is rendered onto the spatial mask. This removes the need for a standalone text encoder at inference time. The resulting model supports contextual image generation by following user-specified instructions about what should appear where, while substantially reducing computation. To address this task, we propose a framework in which the UniVL encoder, adapted from an optical-character-recognition-pretrained backbone, reads the unified condition optically and produces a UniVL embedding, fVIL, that fuses visual and semantic intent with spatial locations in a single token sequence. A two-stage pipeline first aligns UniVL with the VAE embedding space and then conditions a pretrained diffusion backbone entirely on UniVL embeddings, eliminating the standalone text encoder, such as T5. Although this reframing uses a deliberately minimal text interface, it yields strong empirical gains. On UniVL-ImgGen, a benchmark of 477K mask-annotated images that we construct for training and evaluation, UniVL improves image quality over text-prompted baselines, reducing FID from 14 to 11 and increasing PSNR from 16 to 20. It also eliminates the text encoder entirely, reducing inference TFLOPs by up to 52% and runtime by up to 44%. Additional ablation studies validate the contributions of the proposed components, paving the way for efficient, spatially grounded image generation with a unified conditioning paradigm.
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45% to 58.51%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
Yuxing Cheng, Yuan Wu, Yi Chang
School of Artificial Intelligence, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China · International Center of Future Science, Jilin University