The well-aligned attribute of CLIP-based models enables its effective application like CLIPscore as a widely adopted image quality assessment metric. However, such a CLIP-based metric is vulnerable for its delicate multimodal alignment. In this work, we propose FoCLIP, a feature-space misalignment framework for fooling CLIP-based image quality metric. Based on the stochastic gradient descent technique, FoCLIP integrates three key components to construct fooling examples: feature alignment as the core module to reduce image-text modality gaps, the score distribution balance module and pixel-guard regularization, which collectively optimize multimodal output equilibrium between CLIPscore performance and image quality. Such a design can be engineered to maximize the CLIPscore predictions across diverse input prompts, despite exhibiting either visual unrecognizability or semantic incongruence with the corresponding adversarial prompts from human perceptual perspectives. Experiments on ten artistic masterpiece prompts and ImageNet subsets demonstrate that optimized images can achieve significant improvement in CLIPscore while preserving high visual fidelity. In addition, we found that grayscale conversion induces significant feature degradation in fooling images, exhibiting noticeable CLIPscore reduction while preserving statistical consistency with original images. Inspired by this phenomenon, we propose a color channel sensitivity-driven tampering detection mechanism that achieves 91% accuracy on standard benchmarks. In conclusion, this work establishes a practical pathway for feature misalignment in CLIP-based multimodal systems and the corresponding defense method.
Figures & tables
Figure 1: Illustration of fooling CLIPscore. As shown, 0.32 is the correct score, but through our FoCLIP method, despite this being visually inconsistent, the CLIPscore is unexpectedly high.
Figure 2: The framework of FoCLIP, a tripartite optimization approach for adversarial CLIPscore manipulation. Built upon stochastic gradient descent (SGD) updates to the image feature vector g(x) , this framework iteratively adjusts pixel values to bridge the modality gap between visual and textual embeddings. The architecture decomposes the adversarial process into three synergistic components: (a) Feature Alignment Loss minimizes the cosine distance between image features and target text prompts to enhance semantic alignment in CLIP’s embedding space. (b) Distribution Balance Loss ensures balanced similarity scores across multiple prompts by penalizing variance, avoiding overfitting to specific concepts. (c) Pixel-Guard Regularization Loss constrains pixel values within a predefined range [boundlower,boundupper] via ReLU limitations, preserving visual fidelity during optimization.
Figure 3: Heatmap of CLIPscore of famous artworks and titles, including CLIPMasterPrints for SGD, LVE and PGD approaches [ 14 ] , and comparing with our methods with 1000 and 50,000 iterations. Our fooling examples showed the best performance.
Figure 4: (a) CLIPscore comparison of fooling images generated by SGD, LVE, PGD and our method across 25 target classes, alongside similarity scores of corresponding ImageNet validation images. (b) Average similarity trends across 25-100 categories show our method outperforms others significantly, with minimal score degradation as category count increases (note: some variance values are imperceptible due to scale in (b)).
Figure 5: To illustrate the relationship between pixel-guard regularization bounds and CLIPscore, we visualize it via a 3D graph. The x- and y-axes represent boundlower∈[−1,0] and boundupper∈[0,1] , while the z-axis indicates CLIPscore. Representative fooling images are displayed at key points.
Figure 6: Comparison of four methods and original images using grayscale conversion on images as the same as Fig. 4 .
Figure 7: (a) A bar chart showing the average CLIPscore of 50 images in each of the 25 categories. The colors represent the original image, the FoCLIP image with the category as the prompt, and the grayscale image transformed after FoCLIP. (b) The distribution offset after comparing the FoCLIP of the original image and the grayscale image converted after FoCLIP.
Figure 8: (a) and (b) are the grayscale sensitivity distribution maps of the absolute threshold and the relative threshold. (c) The scatter plot showing the relationship between the original CLIPscore and the grayscale sensitivity. The lower right corner is the confusion matrix using a hybrid method combining absolute and relative thresholds, demonstrating the counts of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN).
Figure 9: The baseline refers to the FoCLIP method, all scores are average scores. The last three figures show the ablation experiment results of three parts in the FoCLIP method, corresponding to the cases of removing Regloss, removing Varloss, and only retaining CLIPLoss respectively. Since CLIPLoss is the core component of this method, no separate experiment was conducted to remove it.
Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend on pre-trained MCL models. Existing detection-based defenses predominantly rely on the CLIPScore metric, under the assumption that poisoned pairs exhibit lower semantic similarity between the image and the caption. However, we identify two critical flaws remaining in existing methods: (1) the substantial overlap between CLIPScore distributions of benign and poisoned pairs undermines the reliability of this metric, and (2) fixed-threshold detection cannot provide statistical guarantees for ambiguous samples within overlapping regions. To overcome these limitations, we propose integrating conformal prediction (CP), a statistical framework that quantifies uncertainty through nonconformity scores (NCSs), to establish provable confidence bounds for detecting poisoned image-caption pairs. Building on CP, we introduce CASCADE, a novel two-stage Coarse-to-Fine Conformal Backdoor Detection framework. The coarse-grained stage uses cross-modality consistency to identify high-confidence benign and poisoned pairs. In the fine-grained stage, a reference set is constructed from high-confidence poisoned pairs, and instance-level NCSs based on text-space similarity are computed for each sample in the unidentified subset. These NCSs measure conformity to the poisoning distribution and enable precise identification of latent poisoned pairs within the unidentified subset. Extensive experiments on the large-scale CC3M dataset demonstrate that CASCADE achieves an average FPR of 5.79% at 100% TPR and an average AUROC of 0.9867 across diverse attacks, while remaining effective against adaptive attacks.
Yiming Chen, Kemou Li, Haiwei Wu +1
State Key Laboratory of Internet of Things for Smart City, University of Macau · School of Computer Science and Engineering, University of Electronic Science and Technology of China
Adversarial attacks pose a challenge to the reliability of deep learning models, motivating effective detection methods. Existing techniques often rely on attack-specific assumptions, access to adversarial samples, or knowledge of the underlying classifier (white-box). We propose A4D Attack- and Architecture-Agnostic Adversarial Detector, a completely black-box, zero-shot adversarial attack detection framework that utilizes prompt-based similarity scores derived from CLIP. To the best of our knowledge this is the first attempt to utilize CLIP for such a task. The method is based on two key observations: (i) CLIP is sensitive even to small imperceptible non-semantic perturbations; (ii) The shift in CLIP embedding space is not arbitrary and can be used as a robust attack indicator. Experiments across multiple attacks, datasets and classifiers validate that A4D achieves SOTA detection results in the attack-agnostic and classifier-agnostic setting.
Hodaya Krakover, Meir Yossef Levi, Eyal Gofer +1
Technion - Israel Institute of Technology, Haifa, Israel
Detecting face forgeries using CLIP has recently emerged as a promising and increasingly popular research direction. Owing to its rich visual knowledge acquired through large-scale pretraining, most existing methods typically rely on the visual encoder of CLIP, while paying limited attention to the text modality. Given the instructive nature of the text modality, we posit that it can be leveraged to instruct Deepfake detection with meticulous design. Accordingly, we shift the focus from the visual modality to the text modality and propose a new Separable Prompt Learning strategy (SePL) that enables CLIP to serve as an effective face forgery detector. The core idea of SePL is to disentangle forgery-specific and forgery-irrelevant information in images via two types of prompt learning, with the former enhancing detection. To achieve this disentangle, we describe a cross-modality alignment strategy and a set of dedicated objectives. Extensive experiments demonstrate that, with this simple adaptation, our method achieves competitive and even superior performance compared to other methods under both cross-dataset and cross-method evaluation, highlighting its strong generalizability. The codes have been released at https://github.com/OUC-YER/SePL-DeepfakeDetection
Enrui Yang, Yuezun Li
School of Computer Science and Technology, Ocean University of China, Qingdao, China