The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception
Authors: Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth André
Organizations: Chair for Human-Centered Artificial Intelligence, University of Augsburg · Institute for Distributed Intelligent Systems, University of the Bundeswehr Munich
Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a 40-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at κw=0.468 on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor (κw=0.268-0.486). Under verification all eleven clear it, each of them significantly better at κw=0.507-0.586. Three also significantly beat EIF sitting at κw=0.551 (Small; 0.534 Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to κw=0.254-0.423. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
Figures & tables
Model
Params
Prompt
Decoding
Gemma-3-4b-it [ 16 ]
4.30 B
ours
greedy
Gemma-3-12b-it [ 16 ]
12.19 B
theirs
greedy
Gemma-4-12B-it [ 17 ]
11.96 B
theirs
greedy
GLM-4.6V-Flash [ 18 ]
10.29 B
theirs
T=0.8
InternVL3_5-8B [ 52 , 9 ]
8.53 B
theirs
greedy
MiMo-VL-7B-RL [ 57 ]
8.31 B
ours
T=0.3
Table 1: The eleven evaluated VLMs, with technical reports. All are open-weight and run locally. Parameter counts are the released safetensors totals. Prompt and decoding name the run behind the generative score we report: theirs is the benchmark’s zero-shot prompt [ 25 ] and ours our own wording of the same request ( Appendix D ), reported as the better of the two per model, the choice favorable to generative elicitation. Every model was also run under the other prompt and under the other decoding policy, and neither choice carries the result ( Appendix E ). Verification samples nothing.
Figure 2 : How the answer is read, not which model is asked, decides whether these models reach human agreement. (a) Quadratic-weighted κ with 95% image-bootstrap intervals under generative elicitation (open circles, dashed) and verification (filled squares). No generative interval lies entirely above the 0.468 human anchor. All eleven verification intervals do. (b) The between-model sd falls from 0.073 to 0.024 while the median interval narrows from 0.043 to 0.034 , so the spread moves inside the measurement precision rather than the measurement becoming noisier.
Figure 3 : Discard the graded logit and the gain goes with it, even though the question is unchanged. The same stored verification answers scored continuously from P(yes) (filled squares) and re-scored by thresholding them at 0.5 (open squares), and each gray line follows one model. No new inference is run and no prompt is altered, yet every model loses, and the binarized arm lands back inside the range spanned by generative elicitation (shaded, the eleven models with both elicitations).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : On real photographs the effect survives as six results, not a sweep. Change in six-way accuracy from generative to verification elicitation, with paired person-bootstrap 95% intervals ( 171 persons). Six models gain detectably (filled squares), three have intervals spanning zero (open circles), and Gemma-3-12b-it reverses (filled triangle). MiMo-VL-7B-RL ( † , grayed) is excluded by the validity gate of Sec. 3 ; we plot it because its delta is negative, so the exclusion favors our hypothesis. Alphabetical, not sorted by score.
In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more critical than ever. However, existing benchmarks fail to simultaneously achieve ecological validity, signal clarity, and reliable fine-grained labeling. We introduce EmoS, a high-fidelity bilingual benchmark designed to resolve the limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset. Supported by a rigorous dual-layer human annotation pipeline, EmoS provides trusted ground truth that captures continuous emotional evolution. Empirical results show that fine-tuning MLLMs (multimodal large language models) on EmoS yields significant gains over zero-shot baselines, laying the foundation for the training and evaluation of future emotion recognition models and empathy models. The dataset and code are publicly available at https://github.com/NLP2CT/EmoS.
Pengze Guo, Jingxi Liang, Zhiwen Xie +2
NLP2CT Lab, Department of Computer and Information Science, University of Macau · School of Computer Science, Central China Normal University · University of Macau
Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last few years for many visual tasks, potentially offering a promising solution for understanding emotions. However, it is surprising that even the most sophisticated contemporary VLMs struggle to recognize human emotions or to outperform even specialized vision-only classifiers. In this paper we ask the question "Why do VLMs struggle to recognize human emotions?", and observe that the inherently continuous and dynamic task of facial expression recognition (DFER) exposes two critical VLM vulnerabilities. First, emotion datasets are naturally long-tailed, and the web-scale data used to pre-train VLMs exacerbates this head-class bias, causing them to systematically collapse rare, under-represented emotions into common categories. We propose alternative sampling strategies that prevent favoring common concepts. Second, temporal information is critical for understanding emotions. However, VLMs are unable to represent temporal information over dense frame sequences, as they are limited by context size and the number of tokens that can fit in memory, which poses a clear challenge for emotion recognition. We demonstrate that the sparse temporal sampling strategy used in VLMs is inherently misaligned with the fleeting nature of micro-expressions (0.25-0.5 seconds), which are often the most critical affective signal. As a diagnostic probe, we propose a multi-stage context enrichment strategy that utilizes the information from "in-between" frames by first converting them into natural language summaries. This enriched textual context is provided as input to the VLM alongside sparse keyframes, preventing attentional dilution from excessive visual data while preserving the emotional trajectory.
Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara +1
Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. Building from 351K images collected from six public sources, we apply a rigorous multi-stage filtering pipeline to curate 138K high-confidence images. Each image is annotated at three hierarchical levels: perception QA for emotion and valence recognition, grounded understanding QA constructed from visual trigger extraction through constraint-guided generation, and cognition QA centered on response intent prediction and sequential insight reasoning. In total, InsightVQA contains 725K QA pairs. We further present \textbf{InsightVQA-Bench}, a high-quality evaluation benchmark comprising 30K samples for fine-grained evaluation. To support evaluation, we introduce \textbf{InsightNet}, an emotion-tuned baseline for MLLMs. Results demonstrate that InsightVQA poses significant challenges for grounded emotion understanding and reasoning.