cs.CLJul 8, 2026

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

Authors: Yiming GaiJunde LuXuefei Huang

Organizations: School of Computer Science and Engineering, Beihang University, Beijing, China · Data Science and Intelligent Computing Laboratory, Hangzhou International Innovation Institute, Beihang University, Hangzhou, Zhejiang 311115, P.R.China

Abstract

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

Explore similar work

CardsList
  1. Evaluation of Contextual Understanding in Large Language Models

    Sep 8, 2026Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe +4ContextualSemantic Similarity Metrics