cs.CVOct 8, 2026

Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders

Authors: Tao Yang, Jianying Zhou

Organizations: Singapore University of Technology and Design

Abstract

Adversarial attacks on vision-language models optimize an image toward a text target, then cite the attacked model's similarity score as evidence of success. We ask whether that score - victim-space target alignment (VTS) - predicts recovery of the target by an independent model. We first build a measurement instrument: supervised judges outside the attacked geometry, real-target blend controls, shuffled-target negatives, and a reference level derived from a 50% target-image blend. Two preregistered studies then compare six contrastive encoders under a matched attack at three perturbation budgets. Robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP; all eight contrasts reject at the bootstrap floor. However, no cell reaches the blend-derived reference level. The three best cells fall within its replication band, leaving practical recovery undecided. Within robust encoders, per-sample alignment gain correlates with evidence gain (ρ=0.24−0.51ρ= 0.24-0.51); within vanilla CLIP the correlation is consistent with zero. Across encoders we find no monotone alignment-evidence relation. VTS is therefore informative only within a fixed robust encoder, and we provide a reporting protocol in its place.

Figures & tables

Explore similar work

Jun 4, 2026cs.CV

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-world applications. Test-time defenses for VLMs have recently emerged as a promising and efficient approach to defend against adversarial attacks without requiring costly large-scale retraining. In this work, we uncover a surprising phenomenon: under diverse input transformations, adversarial images in CLIP's feature space consistently shift along a dominant direction, in contrast to the dispersed patterns of clean images. We hypothesize that this dominant shift, termed the Defense Direction, opposes the adversarial shift, pointing features back toward their correct class centers. Building on this insight, we propose Directional Bias-guided Defense (DBD), a test-time framework that estimates the Defense Direction and employs a DB-score-based two-stream reconstruction strategy to recover robust representations. Experiments on 15 datasets demonstrate that DBD not only achieves SOTA adversarial robustness while preserving clean accuracy, but also reveals the counterintuitive result that adversarial accuracy can even surpass clean accuracy. This demonstrates that adversarial perturbations inherently encode directional priors about the true decision boundary.
Jul 29, 2026cs.CR

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven published encoding attacks, counting a behavior as broken if any attack succeeds. That metric separates two mechanisms such defenses conflate. Restoring a view the guard never had improves it on both axes at once: it blocks far more attacks, and, measured on a category-balanced benign set, it blocks fewer benign requests, because restating a request normalizes the borderline phrasing a classifier over-flags. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast survives multiple-comparison correction. What does lower ensemble attack success is re-screening the recovered pre-decode surface, and that step is where the entire benign cost falls. The safety-utility trade-off is therefore not a property of recovery; it is localized to one step. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. The per-attack averages usually reported understate the attacker roughly fourfold, which is why this frontier is easy to miss. Composing across defense families is the one lever that moved the safety axis, beating every configuration we measured, and still landing far outside any deployable refusal budget.
Sep 29, 2026cs.CV

It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at ε≤4/255ε\leq 4/255

Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at ε≤4/255\varepsilon \leq 4/255. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at ε=2/255\varepsilon = 2/255 and complete replacement reaches 38% at ε=4/255\varepsilon = 4/255. On video, complete replacement reaches 35.9% at ε=1/255\varepsilon = 1/255. We also observe a phenomenon of semantic fusion, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.