cs.CVSep 28, 2026

ControlTrace: Recovering Control Fields for Hidden-Content Recognition

Authors: Zijian Liu, Yaoguang Chen, Liwei Liu, Weixi Wu, Hanming Zhang, Jiashui Wang, Na Ruan

Organizations: Shanghai Jiao Tong University · Ant Group

Abstract

Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we construct FreqBlind, a 6,000-image benchmark spanning contours, real words and non-words across three conditioning strengths. The evaluated transformation-based methods show limited recognition of contour patterns and weakly conditioned hidden content. To address this limitation, we propose ControlTrace to recover the grayscale control field used during generation. An 8.4M-parameter U-Net predicts this field from the carrier image, and a VLM then identifies its content. With Qwen2.5-VL-7B-Instruct, ControlTrace achieves 60.2% open-ended contour recognition accuracy across the three conditioning strengths, exceeding the best of the three evaluated prior methods by 26.9 percentage points. On an A100 GPU, the complete pipeline adds only 7.4 ms (5.3%) to direct VLM inference. Recovered fields have lower pixel errors and higher structural similarity than the evaluated transformation views. Across four evaluated VLMs, ControlTrace retains its overall contour recognition advantage. Recognition remains stable under the tested JPEG compression, Gaussian noise and downsampling. These results support control-field recovery for hidden-content recognition in the evaluated setting.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

    Aug 11, 2026Tao Lin, Gaojie Jin, Zongxin Liu +2Poisoning

  2. UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

    May 20, 2026Jiayun Wang, Yu Wang, Weijie Gan +2Image Generation

  3. When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context

    Sep 28, 2026Yuxing Cheng, Yuan Wu, Yi ChangScene Text RecognitionVision-Language Foundation Models