Organizations: Columbia University · University of Washington · University of Illinois Urbana-Champaign · Johns Hopkins University · Google · Meta · Queen Mary University of London · New York University
Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public
Figures & tables
Figure 1: A mosaic art of “AVMeme Exam" dotted by video frames from the meme clips.
Figure 2: AVMeme Exam tests multimodal understanding beyond language and surface content. Left: Multimodal LLMs understand content much better than context, culture, and world knowledge, and media clips with spoken language than those without. Right: Open-source LLMs still lose to humans on humans’ own familiar memes, while humans remain robust even on their unseen clips, showing that current LLMs are still limited in contextual and cultural thinking and real-world multimedia experience than real humans.
Figure 3: Top: Historical timeline of the 1,032 audio-visual memes curated in AVMeme Exam, spotlighting famous music rhythm, movie lines, sound effects, and viral Internet memes. Bottom: Pie charts summarize the distributions of question types, sound categories, and languages, highlighting the data diversity. Right: Frequent words in the memes’ names and distributions of clip durations and number of choices. The duration is cut to 30 seconds, which is the maximum input audio length for most models.
Figure 4: AVMeme Exam collection & verification pipeline. Videos and Q&As are human collected and verified (yellow). LLMs (gray) are used for text cleanup and to detect questions easily answered by text without audio given.
Figure 5: Meme examples. More examples are provided in Appendix D .
Table 1: Model performance across question types on meme-full and meme-main .
Table 2: Model performance across sound types and languages for meme-full and meme-main .
Model
Evaluation Settings
“This is a meme." +Q
Meme Name+Q
Keep videos for visual_cheat
Default (minimal hint or cheat)
Qwen2-Audio
35.4
44.6
n.a.
34.4
Audio Flamingo 3
41.5
50.0
n.a.
41.7
Music Flamingo
42.9
51.3
n.a.
42.8
Baichuan-Omni
34.1
46.1
n.a.
33.1
+ visual input
41.3
49.4
45.3 (80.3)
39.6 (40.0)
Table 3: The effect of text or visual hint and cheat in multimodal evaluation. The right-most column is our default setting with minimal answer shortcuts. Results from meme-main .
Table 4: Details about model checkpoints and repositories. We follow the provided inference code with default configurations for each model.
Gemini ( thinking_level )
meme-full
meme-main
A
L
C
E
H
U
W
AVG
A
L
C
E
H
U
W
AVG
2.5 Flash ( minimal )
51.8
85.4
61.6
65.8
68.2
60.0
53.3
62.9
42.2
83.3
55.7
58.6
53.6
56.2
51.0
57.2
+ visual input
56.4
82.3
69.2
78.1
72.2
60.6
63.6
68.3
47.7
79.6
66.1
72.4
60.0
57.1
62.1
63.9
3 Flash ( minimal )
64.3
86.3
75.4
75.3
77.4
67.7
67.4
73.2
55.6
84.5
71.4
72.4
67.1
64.2
66.3
69.1
+ visual input
64.3
88.7
80.6
82.2
84.2
74.4
78.9
79.7
55.6
87.4
78.8
79.3
76.5
70.4
77.9
76.6
3 Pro ( low )
66.1
85.5
74.1
84.9
79.7
70.3
70.0
74.9
57.8
83.5
71.4
81.0
71.8
66.0
68.8
71.2
Appendix
Table 5: Gemini family performance across question types, sound types, and languages.
Figure 6: Human evaluation interface. A . Instructions and rules shown before the study, specifying device requirements, headphone use, no-search/no-collaboration constraints, and quality control policies to ensure reliable human judgments. B . Participant information and background survey collected anonymously and linked only via a participant ID, including demographics, language background, and social media usage. C . Familiarity check presented before showing the Q&A, where participants indicate whether they have previously seen the meme clip. D . Q&A, in which participants watch and listen to the meme clip and answer a multiple-choice question.
Figure 7: The effect of different levels of text hint from the visual. Results from meme-main .
Figure 8: Model performance vs. the original year of the clips. Results from meme-main .
Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes. We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware. Providing minimal cultural context yields consistent gains across all models and languages: mean SBERT similarity improves from 44.6 to 56.4 (+11.8), BLEURT from 37.3 to 42.3 (+5.0), and LLM-as-a-Judge scores from 2.57 to 3.43 out of 5 (+0.86). Fine-grained error analysis reveals that closed-source models fail mainly on entity and reference misidentification, while open-source models are bottlenecked by broader cultural knowledge gaps, with linguistic and phonological failures proving the most context-resistant across both. These results highlight the difficulty of culturally grounded meme understanding and motivate future work on explicit cultural knowledge integration. Our dataset and code are publicly available at TawsifDipto17/MemeCULT-1K.
Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5
1Islamic University of Technology · 2South East University · 3Vector Institute +3
Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.
Weihang Wang, Kainan Tu, Jielei Zhang +9
1Bilibili · 2Fudan University · 3Peking University
Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a cognitively inspired benchmark that evaluates Omni-MLLMs across three stages, perception, understanding, and reasoning, through cross-modal tasks requiring joint audio-visual interpretation. This design enables fine-grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models' primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open-source and closed-source models reveal substantial limitations in current Omni-MLLMs. Based on these findings, we present a four-level AVI taxonomy. Overall, AVI-Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI-Bench/
Yaoting Wang, Ziyi Zhang, Wenming Tu +10
Institute of Big Data, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Huazhong University of Science and Technology · Shanghai Jiaotong University +5