Language Model Generation Evaluation

Latest papers 209

All topics
CardsList
  1. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

    Jul 31, 2026Mengting Chen, Yanshu Sun, Wanting Liang +5Open-Ended GenerationLLM Evaluation

  2. Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

    Jul 30, 2026Zheng Wu, Yibo Luo, Pu Zhang +2Human Preference EvaluationLLM-as-a-Judge

  3. Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

    Jul 27, 2026Zongyou Yang, Yinghan HouLLM EvaluationLLM Answer Verification

  4. QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

    Jul 23, 2026Emilio FerraraOpen-Ended GenerationLLM Quantization

  5. StabilityBench: Benchmarking Instability in LLMs

    Jul 17, 2026Emma Kondrup, Zachary Yang, Anne Imouza +1LLM EvaluationLanguage Model Generation Evaluation

  6. BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

    Jul 13, 2026Yuzhe Guo, Mengzhou Wu, Yuan Cao +4Tool-Augmented Language Model AgentsLanguage Model Generation Evaluation

  7. Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

    Jul 11, 2026Jinglan Gong, Jiefan Lu, Hewei Guo +5Language Model Generation EvaluationPersona Consistency

  8. Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

    Jul 6, 2026Sadia Kamal, Arefa Patwary, Anthony Marchiafava +2LLM EvaluationLanguage Model Generation Evaluation

  9. SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

    Jul 6, 2026Thomas Thebaud, Yuzhe Wang, Hao Zhang +5Audio-Language Model EvaluationLanguage Model Generation Evaluation

  10. Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

    Jul 6, 2026Robert Morabito, Tyler McDonald, Charitra Viswanath +4Human Preference EvaluationLLM Evaluation

  11. Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability

    Jul 5, 2026Christopher Nassif, Josh F. CooperAI-Generated Text DetectionLanguage Model Generation Evaluation

  12. Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

    Jul 2, 2026Jan DrchalLLM EvaluationLanguage Model Generation Evaluation

  13. Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

    Jul 2, 2026Dazhi Fu, Jiuding Yang, Yiwen Guo +1Reward ModelingLLM-as-a-Judge

  14. LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

    Jul 1, 2026Ruotong Zhao, Zhiyu Chen, Xurui Liu +7Human Preference EvaluationLLM-as-a-Judge

  15. Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

    Jul 1, 2026César Guerra-Solano, Xiang Lorraine LiLLM EvaluationLanguage Model Generation Evaluation

  16. "Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo

    Jul 1, 2026Sara Candussio, Francesca Padovani, Daniel Scalena +1LLM GroundingLanguage Model Generation Evaluation

  17. SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

    Jun 30, 2026Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou +5LLM EvaluationLanguage Model Generation Evaluation

  18. Little Brains, Big Feats: Exploring Compact Language Models

    Jun 29, 2026Dari Baturova, Elena Bruches, Ivan Chernov +3Retrieval-Augmented GenerationOn-Device Language Model Inference

  19. Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

    Jun 29, 2026Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie +2Language Model Generation EvaluationMultimodal Reasoning