LLM Evaluation

LLM: Large Language Model

Latest papers 1,643

All topics
CardsList
  1. MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media

    May 7, 2026Souvik Pramanik, S. M. Riaz Rahman Antu, Shak Mohammad Abyad +2LLM EvaluationLLM-Assisted Annotation

  2. Can LLMs Take Retrieved Information with a Grain of Salt?

    May 7, 2026Behzad Shayegh, Mohamed Osama Ahmed, Fred Tung +1Retrieval-Augmented GenerationLLM Evaluation

  3. How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem

    May 7, 2026Chun Zheng, Lianlong Wu, Bingqian Li +2LLM EvaluationLLM Reasoning

  4. IntentGrasp: A Comprehensive Benchmark for Intent Understanding

    May 7, 2026Yuwei Yin, Chuyuan Li, Giuseppe CareniniLLM EvaluationIntent Classification

  5. Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML

    May 7, 2026Jai Moondra, Ayela Chughtai, Bhargavi Lanka +1Multilingual Language Model EvaluationLLM Evaluation

  6. When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

    May 7, 2026Sushant Gautam, Finn Schwall, Annika Willoch Olstad +6LLM Safety BenchmarksLLM Evaluation

  7. Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

    May 7, 2026Hailey Onweller, Elias Lumer, Austin Huber +3LLM EvaluationCitation Verification

  8. How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation

    May 7, 2026Shai Feldman, Yaniv RomanoLLM EvaluationLanguage Model Safety Evaluation

  9. Ex Ante Evaluation of AI-Induced Idea Diversity Collapse

    May 7, 2026Nafis Saami Azad, Raiyan Abdul BatenLLM EvaluationComputational Creativity

  10. Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts

    May 7, 2026Sneha Oram, Ojaswita Bhushan, Pushpak BhattacharyyaLLM EvaluationLanguage Model Robustness

  11. SCRuB: Social Concept Reasoning under Rubric-Based Evaluation

    May 7, 2026Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11LLM EvaluationRubric-Based Evaluation

  12. Rethinking Vacuity for OOD Detection in Evidential Deep Learning

    May 7, 2026Claire McNamaraLLM EvaluationEvidential Deep Learning

  13. Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

    May 7, 2026Haoxiang Wang, Da Yu, Huishuai ZhangLLM EvaluationLanguage Model Generation Evaluation

  14. Teaching LLMs Program Semantics via Symbolic Execution Traces

    May 7, 2026Jonas Bayer, Stefan Zetzsche, Olivier Bouissou +3LLM EvaluationContinued Pretraining

  15. Visual Fingerprints for LLM Generation Comparison

    May 7, 2026Amal Alnouri, Andreas Hinterreiter, Christina Humer +2LLM EvaluationLanguage Model Fingerprinting

  16. More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs

    May 7, 2026Adrián Gude, Roi Santos-Ríos, Francis Bond +3LLM Evaluation

  17. Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals

    May 7, 2026Wenliang Huang, Zengyi YuLLM EvaluationLLM Alignment

  18. Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

    May 7, 2026Yang Xu, Jiefu Zhang, Haixiang Sun +3LLM EvaluationSelective Inference

  19. Logic-Regularized Verifier Elicits Reasoning from LLMs

    May 7, 2026Xinyu Wang, Changzhi Sun, Lian Cheng +4LLM EvaluationLLM Answer Verification

  20. CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency

    May 7, 2026Hirofumi Ota, Naoto Iwase, Yuki Ichihara +2LLM EvaluationAnytime-Valid Inference

  21. Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

    May 7, 2026Fahd Seddik, Fatemeh FardLLM EvaluationLLM Interpretability

  22. MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge

    May 6, 2026Perry E. RadauLLM EvaluationHealthcare

  23. The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences

    May 6, 2026Hubert Plisiecki, Sabina Siudaj, Kacper Dudzic +4LLM EvaluationPersonality Modeling in Language Models

  24. Curated AI beats frontier LLMs at pharma asset discovery

    May 6, 2026Łukasz Kidziński, Kevin ThomasLLM EvaluationHealthcare

  25. BenCSSmark: Making the Social Sciences Count in LLM Research

    May 6, 2026Arnault Chatelain, Étienne Ollion, Qianwen Guan +7LLM EvaluationBenchmark Design