LLM Evaluation

LLM: Large Language Model

Latest papers 1,631

All topics
CardsList
  1. ΦΦ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

    Sep 9, 2026Leilei Ding, Shumin Wang, Yuting Huang +10LLM EvaluationLarge Language Model-Guided Optimization

  2. RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

    Sep 9, 2026Yingqian Wu, Jingcong Liang, Siyuan Wang +4LLM Evaluation

  3. Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

    Sep 9, 2026Joana Rosa, Pedro Santos, Valdemar Oliveira +3LLM EvaluationLLM Planning

  4. Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format

    Sep 9, 2026Touchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit ModecruaLLM EvaluationLanguage Model Calibration

  5. Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

    Sep 8, 2026Peter Kochelka, Aleš Manuel Papáček, Vojtěch Dvořák +1LLM EvaluationKnowledge Conflicts in Language Models

  6. Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

    Sep 8, 2026Leyuan Tang, Kangda Wei, Tianyu Jiang +1LLM Safety BenchmarksLLM Evaluation

  7. Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    Sep 8, 2026Jaewoo Lim, Sungbok Shin, Sanghyun HongLLM EvaluationLLM Interpretability

  8. Evaluation of Contextual Understanding in Large Language Models

    Sep 8, 2026Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe +4LLM EvaluationLLM Grounding

  9. Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

    Sep 8, 2026Xingyu Deng, Mingzi Cao, Nikolaos Aletras +2LLM EvaluationLLM Grounding

  10. API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

    Sep 8, 2026Jennifer Wang, Joachim Baumann, Daniel E. Ho +1LLM EvaluationLanguage Model Generation Evaluation

  11. Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

    Sep 8, 2026Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah +1LLM EvaluationLLM Auditing

  12. A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

    Sep 8, 2026Maysam Khatib, Moysis Symeonides, Demetris Trihinas +2LLM Inference EfficiencyLLM Evaluation

  13. CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

    Sep 7, 2026Yonghong Zhang, Ricardo Correia, Isabel M. Parra +1LLM EvaluationCausal Effect Estimation

  14. Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

    Sep 7, 2026Ali Şenol, H. Russell Bernard, Huan LiuLLM EvaluationLLM Honesty

  15. How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

    Sep 7, 2026Riyadh Alnasser, Yusuf Mücahit Çetinkaya, Sumin Zhao +1LLM EvaluationLLM Sycophancy

  16. ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

    Sep 7, 2026Jun Xiang, Zhijie Bao, Rong Hu +3LLM EvaluationClinical Decision-Making

  17. FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect

    Sep 7, 2026Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10LLM EvaluationLLM Reliability

  18. Human-like moral judgments conceal divergent motive attributions in large language models

    Sep 7, 2026Xiaoyan Wu, Jean-Claude DreherLLM EvaluationMoral Reasoning in Language Models

  19. Quality Metrics for LLM-Generated Asset Administration Shells: A Perturbation-Based Evaluation Approach

    Sep 7, 2026Janek Groß, Elena Zentgraf, Jens HeidrichLLM Evaluation

  20. FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

    Sep 7, 2026Tyrone White, Yuki AraseLLM EvaluationLanguage Modeling

  21. PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

    Sep 5, 2026Gal Sapir, Alon Diament, Adva Wolf +7LLM EvaluationBenchmark Design

  22. Molecular Déjà Vu: Digit-Level Retrieval of Molecular Properties in Frontier Language Models

    Sep 4, 2026Matthias Busch, Marius Tacke, Sviatlana V. Lamaka +4LLM EvaluationLLM Reliability

  23. Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

    Sep 4, 2026Yumiao Li, Peixin Liu, Donglin Di +2LLM EvaluationPolicy Evaluation

  24. SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

    Sep 4, 2026Shenxi Wu, Yuhong Liu, Haosong Zhang +8Document UnderstandingLLM Evaluation

  25. PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

    Sep 3, 2026Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis +1LLM Evaluation

  26. Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

    Sep 3, 2026Haoyaun Zhu, Jie ZhangLLM EvaluationLLM-as-a-Judge

  27. Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

    Sep 3, 2026Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu +1LLM EvaluationHealthcare