LLM Evaluation

LLM: Large Language Model

Latest papers 1,643

All topics
CardsList
  1. LLM Benchmark Datasets Should Be Contamination-Resistant

    May 19, 2026Ali Al-Lawati, Jason Lucas, Dongwon Lee +1LLM EvaluationBenchmark Contamination

  2. LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation

    May 19, 2026Shanshan Xu, Johan Lindholm, Amogh Raina +2Legal NLPLLM Evaluation

  3. Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

    May 19, 2026Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad +1Mathematical Reasoning BenchmarksLLM Evaluation

  4. K-Quantization and its Impact on Output Performance

    May 19, 2026Robin Baki Davidsson, Pierre NuguesLLM QuantizationLLM Evaluation

  5. LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

    May 19, 2026Ming Zhang, Qiyuan Peng, Yinxi Wei +13Logical ReasoningLLM Evaluation

  6. Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

    May 19, 2026Atkia Mahila, Avinash Maurya, M. Mustafa Rafique +1LLM EvaluationInference-Time Search

  7. The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

    May 19, 2026David Pape, Jonathan Evertz, Lea SchönherrLLM EvaluationLLM Inference

  8. Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment

    May 19, 2026Grandee Lee, Yue Wang, Che Yee Lye +1LLM EvaluationEducational Assessment

  9. BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

    May 19, 2026Carla Castedo, Enrique Iglesias, Manuel Lama +3KG ConstructionLLM Evaluation

  10. SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models

    May 19, 2026Yiyang Gu, Junwei Yang, Junyu Luo +15AI for ScienceLLM Evaluation

  11. HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

    May 19, 2026Emmy Liu, Varun Gangal, Michael Yu +4LLM EvaluationHallucination in Language Models

  12. OpenCompass: A Universal Evaluation Platform for Large Language Models

    May 19, 2026Maosong Cao, Kai Chen, Haodong Duan +27LLM Evaluation

  13. Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

    May 19, 2026Yan Wang, Ziyi Guo, Christopher McCartyRetrieval-Augmented GenerationLLM Evaluation

  14. Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

    May 19, 2026Tiejin Chen, Longchao Da, Xiaoou Liu +1LLM EvaluationUncertainty Quantification

  15. Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

    May 18, 2026Matthew L. Smith, Jonathan P. Shock, Samuel T. Segun +2LLM EvaluationLanguage Model Scaling Laws

  16. GIM: Evaluating models via tasks that integrate multiple cognitive domains

    May 18, 2026Rohit Patel, Alexandre Rezende, Steven McClainLLM EvaluationTest-Time Scaling

  17. SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

    May 18, 2026Nithin Somasekharan, Youssef Hassan, Shiyao Lin +5AI for ScienceLLM Evaluation

  18. Forecasting Downstream Performance of LLMs With Proxy Metrics

    May 18, 2026Arkil Patel, Siva Reddy, Marius Mosbach +1Language Model PretrainingLLM Evaluation

  19. Estimating Item Difficulty with Large Language Models as Experts

    May 18, 2026Diana Kolesnikova, Kirill Fedyanin, Abe D. Hofman +2LLM EvaluationLLM-as-a-Judge

  20. QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi

    May 18, 2026Anthony G. Cohn, Robert E. BlackwellSpatial Reasoning BenchmarksLLM Evaluation

  21. Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

    May 18, 2026Tara Azin, Yongan Yu, Raj Singh +1LLM EvaluationPragmatic Reasoning in Language Models

  22. Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

    May 18, 2026Sterling Huang, Abigayle Brown, Jiyoo Noh +4LLM EvaluationLLM Compression

  23. Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

    May 17, 2026M. Mikail Demir, M. Abdullah CanbazLegal NLPLLM Evaluation

  24. Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models

    May 17, 2026Ethan TangLLM EvaluationLanguage Modeling

  25. CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models

    May 17, 2026Yuning Wu, Yingmin Liu, Yang ShuLLM Self-CorrectionLLM Evaluation

  26. A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

    May 17, 2026Qingchuan Ma, Yuexiao Ma, Yongkang Xie +3LLM EvaluationSynthetic Benchmark Generation

  27. CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

    May 16, 2026Zhaoyue Sun, Hainiu Xu, Andero Uusberg +3LLM EvaluationAffective Computing

  28. Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

    May 16, 2026Fangzhou Wu, Sandeep Silwal, Qiuyi ZhangLLM EvaluationClustering