cs.CVMay 16, 2025

HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

Authors: Shaina RazaAravind NarayananVahid Reza KhazaieAshmal VayaniAhmed Y. RadwanMukund S. ChettiarAmandeep SinghMubarak Shah+1 more

Organizations: Vector Institute for Artificial Intelligence, Toronto, Canada · University of Central Florida, Orlando, USA

Abstract

Although recent large multimodal models (LMMs) show impressive progress on vision language tasks, their alignment with human centered (HC) principles such as fairness, ethics, inclusivity, empathy, and robustness is often overlooked. Existing LMM benchmarks are largely accuracy-agnostic. We present HumaniBench, a unified framework for characterizing HC alignment across realistic, socially grounded visual contexts. It contains 32,000 expert-verified image-question pairs from real-world news imagery, each mapped to one or more HC principles through explicit metrics. Comparing 15 state of the art LMMs reveals consistent trade -offs: proprietary systems lead on ethics, reasoning, and empathy, while open-source models show superior visual grounding and resilience. All models show persistent gaps in fairness and multilingual inclusivity. Chain-of-thought prompting and test-time scaling yield 8to 12 % gains on several HC dimensions. HumaniBench enables fine-grained analysis of alignment trade-offs not captured by conventional multimodal benchmarks. https://vectorinstitute.github.io/humanibench/

Explore similar work

CardsList
  1. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    Feb 13, 2025Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31Spatial Reasoning