LLM Evaluation

LLM: Large Language Model

Latest papers 1,643

All topics
CardsList
  1. Capability Self-Assessment: Teaching LLMs to Know Their Limits

    May 29, 2026Haoyan Yang, Reza Shirkavand, Yukai Jin +3LLM EvaluationLanguage Model Self-Assessment

  2. Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions

    May 29, 2026Wesley Scivetti, Ethan Wilcox, Nathan Schneider +2LLM EvaluationOpen-Weight Language Models

  3. Preference-Aware Rubric Learning for Personalized Evaluation

    May 29, 2026Yilun Qiu, Xiaoyan Zhao, Yang Zhang +7LLM EvaluationLLM Alignment

  4. RealityTest: How People Probe AI Identity and Whether Models Disclose It

    May 29, 2026Anna Gausen, Sarenne Wallbridge, Bessie O'Dell +2LLM EvaluationLanguage Model Safety Evaluation

  5. LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability

    May 29, 2026Tom Lucas, Alessio Buscemi, Alfredo Capozucca +2LLM EvaluationAI Accountability

  6. Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

    May 29, 2026Han Zhang, Zihao Tang, Xin Yu +8LLM EvaluationMemory-Augmented Language Models

  7. How Much Do LLMs Know About Chinese Zero Pronouns?

    May 29, 2026Yifei Li, Guanyi Chen, Tingting HeLLM Evaluation

  8. Office Comprehension Benchmark

    May 29, 2026Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi +17Document UnderstandingLLM Evaluation

  9. Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

    May 29, 2026Weijia Zhang, Ruiqi Chen, Yunze Xiao +1LLM EvaluationMoral Reasoning in Language Models

  10. Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

    May 28, 2026Munawar HasanLLM EvaluationLanguage Model Distillation

  11. SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

    May 28, 2026Sy-Tuyen Ho, Minghui Liu, Huy Nghiem +1AI for ScienceLLM Evaluation

  12. Resolution Diagnostics for Paired LLM Evaluation

    May 28, 2026Anany KotawalaLLM EvaluationPairwise Comparison

  13. ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

    May 28, 2026A. J. Lew, Y. Cao, M. J. BuehlerLLM EvaluationScientific Reasoning

  14. Protocol for evaluating ChatGPT in biomedical association generation and verification using a RAG-enabled, cross-model majority voting workflow

    May 28, 2026Ahmed Abdeen Hamed, Luis M. RochaLLM EvaluationLLM Answer Verification

  15. SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

    May 28, 2026Jiamin Chen, Yidi Wu, Qiexiang Wang +6LLM EvaluationLLM-as-a-Judge

  16. Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

    May 28, 2026Asaf Yehudai, Naama Rozen, Ariel GeraLLM EvaluationHuman Behavior Simulation

  17. Latent Performance Profiling of Large Language Models

    May 28, 2026Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya +7LLM EvaluationLLM Interpretability

  18. Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

    May 28, 2026Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong +2LLM EvaluationLanguage Model Bias Evaluation

  19. ExCAM: Explainable Cultural Awareness Metrics

    May 28, 2026Christoph Leiter, Haiyue Song, Hour Kaing +4LLM EvaluationLanguage Model Generation Evaluation

  20. NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

    May 28, 2026Anany KotawalaData LeakageLLM Evaluation

  21. PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

    May 28, 2026Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł +2LLM EvaluationScholarly Peer Review

  22. NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

    May 28, 2026Yunjin Qi, Zhaojun Jiang, Xuan Wu +10LLM Evaluation