cs.CVOct 6, 2026

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

Authors: Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth André

Organizations: Chair for Human-Centered Artificial Intelligence, University of Augsburg · Institute for Distributed Intelligent Systems, University of the Bundeswehr Munich

Abstract

Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a 4040-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at κw=0.468κ_w = 0.468 on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor (κw=0.268κ_w=0.268-0.4860.486). Under verification all eleven clear it, each of them significantly better at κw=0.507κ_w=0.507-0.5860.586. Three also significantly beat EIF sitting at κw=0.551κ_w = 0.551 (Small; 0.5340.534 Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to κw=0.254κ_w=0.254-0.4230.423. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

    May 9, 2026Pengze Guo, Jingxi Liang, Zhiwen Xie +2Emotion RecognitionEmpathy

  2. Why Do Vision Language Models Struggle To Recognize Human Emotions?

    Apr 16, 2026Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara +1Vision-Language Foundation ModelsEmotion

  3. InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

    Jun 1, 2026Shiyu Wang, Ziyu Liu, Chaoyi Yu +6Emotion RecognitionHuman-Annotated Benchmark