cs.CVSep 16, 2026

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

Authors: Girish A. KoushikDiptesh KanojiaHelen Treharne

Abstract

When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages 0.7400.740 versus 0.4320.432 native macro-F1, while residual reconstruction reaches 0.4860.486, whereas Gemma improves from 0.5320.532 to 0.7140.714. These differences reflect supervised accessibility rather than a pre-existing, native decision rule, and the most influential token role depends on the task. Under the evaluated score scales, Qwen silent-feature ablation is 246324-63 times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is 1614016-140 times more output-sensitive. Calibration-only routing recovers 93.393.3% of the mean gap, and probe-distilled LoRA improves native predictions, although shared multi-task adaptation causes negative transfer. A case study of Gemma-3-12B on Facebook Hateful Memes finds a distributed rank-32 image-prompt interaction, reaching 0.7560.756 versus 0.6850.685 native macro-F1. Robustness controls show that the signal extends beyond English, is not explained solely by accompanying OCR, and depends on paired visual evidence. Thus, routing, rather than representation alone, is a recurring bottleneck in harmful meme classification.

Explore similar work

CardsList
  1. FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

    May 29, 2026Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand +2Hate SpeechRhetoric