In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been studied separately from detection, which does not reflect real-world conditions. Further, since curating large annotated datasets for meme moderation is prohibitively expensive, we propose a novel framework - PEST - that leverages task-specific generative VLMs and the few-shot adaptability of large VLMs to cater to different types of memes. We believe this is the first work focused on generalizable hateful meme moderation under limited data conditions, and has strong potential for deployment in real-world production scenarios. Warning: Contains potentially toxic contents.
Figures & tables
Figure 1. A: Limitations of the current literature, B: overview of PEST – a black-box framework powered by task-specific small agents and few-shot alignment for combined classification, explanation, and intervention of hateful memes, and C: example outputs of unified tasks. \ps{} framework
dataset
task focus
split / label dist.
total
phase 1: training task-specific agents ( AC , AE , and AI )
MemeCap
caption
train + val
5,828
HatReDAug
explanation
H: 2,982 NH: 5,388
8,370
MemeSense
common sense
train only
299
MemeSense
intervention
train only
299
phase 2: support set ( TR )
Table 1. Statistics of datasets used in our experiments. The upper section details the data used to fine-tune the task-specific small agents ( paligemma-3b-pt-448 ), while the two lower section detail the datasets used for selecting the few-shot examples and evaluation of the larger VLMs on PEST framework. AC : caption agent, AE : explanation agent, AI : intervention agent, TR : support set for few-shot exemplars, TS : evaluation set.
Figure 2. Overview of fine-tuning task-specific agents and using them for silver data generation on the support set TR of FHM and MAMI datasets. Note that in the case of generating explanations, we use agent AE only for MAMI dataset and use HatReDAug explanations for FHM.
agent
dataset
# samples
rgL
ss
bsf1
cap
MemeCap
559
0.3571
0.6673
0.9042
exp
HatReD
246
0.3777
0.6326
0.9079
c-s
MemeSense
134
0.2501
0.7212
0.8999
int
MemeSense
134
0.3438
0.7857
0.9114
Table 2. Performance of the four task-specific fine-tuned PaliGemma -3B agents. Higher values indicate strong meaning preservation despite lexical variation. Here, cap: captions, exp: explanation, c-s: common sense, int: intervention, rgL: Rouge -L, ss: Semantic Similarity , and bsf1: BertScore -F1.
FHM-C
FHM-E
FHM-I
MAMI-C
MAMI-E
MAMI-I
task
model
acc
mf1
rgL
ss
bsf1
rgL
ss
bsf1
sprt
acc
mf1
rgL
ss
bsf1
rgL
ss
bsf1
sprt
PH
72.98
72.24
-
-
-
-
-
-
-
70.31
70.18
-
-
-
-
-
-
-
PC
75.1
74.85
-
-
-
-
-
-
-
73.63
73.42
-
-
-
-
-
-
-
MH
57.6
53.88
-
-
-
-
-
-
-
69.05
68.78
-
-
-
-
-
-
-
FS
66
65.8
-
-
-
-
-
-
-
70.5
70.1
-
-
-
-
-
-
-
UC
74.06
74.05
-
-
-
-
-
-
-
76.72
76.34
-
-
-
-
-
-
-
Table 3. Results comparing our method with the previous works. ‘-’ here represent that the prior works cannot perform those mentioned tasks. C: Classification, E: Explanation, I: Intervention, PH: PromptHate , PC: Pro-Cap , MH: Mod-Hate , FS: Few-Shot , UC: U-CoT+ , MK: M2KE , V-L: Vlm-Lim ( GPT-4o ), LM: LoReHM ( GPT-4o ), HR: HatReD , MS: MemeSense , IVL: Intern-VL3 , PX: Pixtral , GPT: GPT-4o , acc: accuracy, rgL: Rouge -L, ss: Semantic Similarity , bsf1: BertScore -F1, sprt: support. Best results are marked green and second best results are marked l ight green .
type
model
dataset
toxicity score
exp
GPT-4o
FHM
0.17 (0.15)
MAMI
0.18 ( 0.14 )
Intern-VL3
FHM
0.15 (0.17)
MAMI
0.27 ( 0.22 )
Pixtral
FHM
0.2 ( 0.2 )
MAMI
0.29 ( 0.21 )
Table 4. Mean (standard deviation) toxicity of generated explanation and intervention using PerspectiveAPI . exp: explanation and int: intervention.
Figure 3. Token count, type token ratio, and perplexity along with error bars at 95% confidence interval. Here, cp: correct positive (circle), cn: correct negative (square), wp: wrong positive (triangle), wn: wrong negative (diamond), g: GPT-4o (maroon), i: Intern-VL3 (purple), and p: Pixtral (green).
Figure 4. Bar charts presenting the distribution of sentiments for all models and across both datasets. Here, cp : correct positive, cn : correct negative, wp : wrong positive, wn : wrong negative, ex : explanation, in : intervention. Sentiment analysis over different prediction types.
FHM-cls
FHM-exp
FHM-int
MAMI-cls
MAMI-exp
MAMI-int
model
shots
acc
mf1
rgL
ss
bsf1
rgL
ss
bsf1
size
acc
mf1
rgL
ss
bsf1
rgL
ss
bsf1
size
2
78.23
78.23
0.221
0.652
0.886
0.144
0.466
0.872
389
86.35
86.22
0.229
0.583
0.882
0.127
0.435
0.864
454
4
78.54
78.53
0.233
0.671
0.889
0.164
0.509
0.876
393
86.03
85.90
0.230
0 .615
0 .884
0.126
0.480
0.866
452
BL
8
79.96
79.95
0 .245
0.684
0.892
0.207
0.576
0.886
402
88.47
88.42
0.241
0.622
0.885
0.141
0.487
0.869
447
2
7 9.45
7 9.44
0.220
0.654
0.886
0.144
0.467
0.871
403
88.47
88.43
0.219
0.582
0.880
0.139
0.508
0.870
443
4
78.94
78.93
0.237
0 .678
0 .890
0.167
0.495
0.875
400
8 8.99
8 8.97
0.230
0.602
0.882
0 .179
0 .585
0 .879
444
Table 5. Results on other embedding retrievers, i.e. on BLIP and CLIP. IVL: Intern-VL3 , PX: Pixtral , GPT: GPT-4o , BL: BLIP, and CL: CLIP. Best results across each model are marked green and second best results are marked l ight green .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5. Word shift graphs computed using Shannon entropy shifts. Each of the two word shift figures presents 30 ranked keywords based on entropy difference among the classes, with a cumulative entropy shift distribution graph present at the bottom left corner.
setup-1
setup-2
FHM
MAMI
GPT-4o (8)
U-CoT+
2.7×10−5
4.4×10−15
Intern-VL3 (8)
6.1×10−8
1.1×10−20
Pixtral (8)
2.9×10−7
3.8×10−21
GPT-4o (8)
GPT-4o (0)
4.5×10−4
0.01
Intern-VL3 (8)
Intern-VL3 (0)
0.051
7.7×10−5
Pixtral (8)
Pixtral (0)
0.003
0.404
Appendix
Table 6. Obtained p -values of exact two-sided binomial McNemar test on two different setups. The upper block presents the comparison of results on GPT-4o with 8 shots vs (a) U-CoT+, and (b) corresponding 8-shot results of Intern-VL3 and Pixtral . The lower block presents a comparison of employed VLMs in the 8-shot setup with their zero-shot setup. Number in (parenthesis) denotes the number of shots employed in the VLM.
Figure 6. Semantic relation based on semantic similarity for all considered models and datasets. Here, cp : correct positive, wp : wrong positive, g : GPT-4o , i : Intern-VL3 , and p : Pixtral .
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: https://meizhiyuan88666.github.io/prokda.
Bo Xu, Chenyuan Wang, Xinyu Chen +5
School of Software, Dalian University of Technology · School of Computer Science and Technology, Dalian University of Technology · School of Computing Technologies, RMIT University
Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.
Hateful and propagandistic memes exploit the interplay between images and text to convey harmful intent that neither modality reveals alone. Although thinking-based multimodal large language models (MLLMs) have advanced vision-language understanding, their application to meme content moderation remains underexplored. We propose a reinforcement learning-based post-training method that improves classification performance and reference-based explanation quality in thinking-based MLLMs via task-specific rewards and Group Relative Policy Optimization (GRPO). Concretely, we (i) conduct a systematic empirical study of off-the-shelf MLLMs for hateful and propagandistic meme understanding across English and Arabic benchmarks, (ii) extend existing meme datasets with weakly supervised chain-of-thought (CoT) rationales via distillation and multi-LLM fine-grained propaganda annotations, (iii) introduce a GRPO-based objective with thinking-length regularization that jointly optimizes classification accuracy and explanation quality, and (iv) investigate self-supervised GRPO on unlabeled memes using consensus-based pseudo-labels. Experiments on the Hateful Memes and ArMeme benchmarks show that our approach improves over previously reported results on FHM accuracy (up to +2.1%, from 79.9% to 82.0%) and on ArMeme macro-F1 (up to +7.6 points, from 0.536 to 0.612 with explanations; +6.1 compared to the original ArMeme benchmark), while also generating natural-language explanations. On ArMeme, sequence-classification baselines remain stronger in terms of raw accuracy, whereas our approach provides more balanced per-class performance along with explanations. We publicly release our code, data extensions, and evaluation resources.
Mohamed Bayan Kmainasi, Mucahid Kutlu, Ali Ezzat Shahroor +2
Qatar Computing Research Institute, Doha, Qatar · Qatar University, Doha, Qatar · APAVI.AI, France +1