cs.AIMay 28, 2026

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Authors: Yunjin QiZhaojun JiangXuan WuHanxi PanYixuan WangYanfang LiuXiang JiChuru Yu+5 more

Organizations: Department of Psychology and Behavioral Sciences, Zhejiang University · College of Artificial Intelligence, Zhejiang University · 3Human Machine Interaction Lab, Huawei Technologies Co., Ltd. · 4Zhejiang Key Laboratory of Neurocognitive Development and Mental Health

Abstract

As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social intelligence has become critical to the quality and safety of human-AI interaction. However, existing social intelligence benchmarks lack a unified framework that organizes social abilities into a unified structure, and therefore cannot enable fine-grained diagnosis. To build the first holistic diagnostic evaluation grounded in social theory, we first construct a social intelligence framework through a literature review and multi-stage expert validation guided by psychometric principles. The resulting framework includes 4 categories and 11 dimensions, each further specified by fine-grained capability facets. Building on this framework, we introduce NICE (Norm, Interaction, Cognition, Experience), a diagnostic benchmark of 137 items operationalized through representative Chinese contexts. Across 5 frontier LLMs and a human reference group, models score higher in aggregate accuracy yet show a consistent weakness in Communication, which the framework localizes to 3 specific capability facets: multi-turn communication, nonverbal communication, and synchrony. NICE thus reframes social intelligence evaluation toward theory-grounded diagnosis of socially consequential weaknesses in LLMs.

Explore similar work

CardsList
  1. Zing: Social Mind for LLMs

    Jul 26, 2026Zing Team, Ao Xiang, Bi Jingping +56Collective IntelligenceInternalization