cs.CVJul 19, 2026

BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos

Authors: Md. Ariful Islam, Md Tanvirul Islam, Md. Maruf Hossain Miru, Md Khalid Syfullah

Organizations: Department of Computer Science and Engineering, Bangladesh Army University of Science and Technology, Saidpur Cantonment, Saidpur, Bangladesh

Abstract

Clickbait, where video titles and thumbnails exaggerate or misrepresent content, reduces user trust, wastes attention, and promotes misinformation on video-sharing platforms. Detecting Bengali clickbait remains challenging because publicly available multimodal datasets are limited. To address this gap, we introduce BanClickThumb, a curated dataset of 7,147 Bengali YouTube thumbnail-title pairs from five content domains, annotated by ten annotators with high agreement (Cohen's Kappa: 0.83-0.93). Using this dataset, we benchmark text-only, image-only, and multimodal approaches. Among unimodal models, BanClickTextFormer (XLM-RoBERTa) achieves 0.82 accuracy, while BanClickImageFormer (SwiftFormer) reaches 0.68. Our proposed multimodal model, BanClickFusionFormer, combines ViT and XLM-RoBERTa through intermediate fusion and achieves the best accuracy of 0.84. Error analysis shows that dense thumbnail text, figurative language, and culturally specific slang remain challenging. Our findings demonstrate the effectiveness of multimodal fusion for Bengali clickbait detection and provide a publicly available benchmark to support future research on low-resource multimodal content analysis.

Explore similar work

CardsList
  1. Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection

    Jul 13, 2026Faria Afrin Tisha, Fariya Tabassum, Hafsa Binte Kibria +2Hate SpeechBangla