cs.CLJan 8, 2026

PEST: Parameter Efficient Steering of Blackbox VLMs via Agentic Few-shot Alignment for Hateful Meme Moderation

Authors: Naquee Rizwan, Subhankar Swain, Paramananda Bhaskar, Shehryaar Shah Khan, Gagan Aryan, Animesh Mukherjee

Organizations: Indian Institute of Technology (IIT), Kharagpur · Simbian

Abstract

In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them before being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been studied separately from detection, which does not reflect real-world conditions. Further, since curating large annotated datasets for meme moderation is prohibitively expensive, we propose a novel framework - PEST - that leverages task-specific generative VLMs and the few-shot adaptability of large VLMs to cater to different types of memes. We believe this is the first work focused on generalizable hateful meme moderation under limited data conditions, and has strong potential for deployment in real-world production scenarios. Warning: Contains potentially toxic contents.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

    Sep 17, 2026Bo Xu, Chenyuan Wang, Xinyu Chen +5Hate Speech DetectionHate Speech

  2. FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

    May 29, 2026Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand +2MemesHate Speech Detection

  3. Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes

    Jun 13, 2026Mohamed Bayan Kmainasi, Mucahid Kutlu, Ali Ezzat Shahroor +2MemesMultimodal Large Language Models