cs.CVOct 1, 2026

Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models

Authors: Yuliang Cai, Mohammad Rostami, Jesse Thomason

Organizations: Viterbi School of Engineering University of Southern California · Amazon Inc.

Abstract

Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.

Figures & tables

Explore similar work

CardsList
  1. HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

    Jun 22, 2026Hoang-Bao Le, Aiden Durrant, Thai Son Mai +3NegationImage-Text Pairs

  2. What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

    Jul 25, 2026Chen-Yi Lu, Yueh-Shao Chen, Somali ChaterjiNegationRepresentational Collapse

  3. Disparities In Negation Understanding Across Languages In Vision-Language Models

    Apr 21, 2026Charikleia Moraitaki, Sarah Pan, Skyler Pulling +3NegationLinguistics