LLM Evaluation

LLM: Large Language Model

Latest papers 1,631

All topics
CardsList
  1. ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation

    Sep 21, 2026Lijian Wu, Henry Hengyuan Zhao, Zijian Zhang +3VLM EvaluationLLM Evaluation

  2. From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification

    Sep 21, 2026Mai Mohamed Eida, Gunjan Anand, Ayush Singh +1Table QALLM Evaluation

  3. Financial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions

    Sep 20, 2026Kemal KirtacLLM EvaluationFinancial Sentiment Analysis

  4. Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

    Sep 20, 2026Jie Ying, Zhefan Wang, Zihong Chen +11LLM EvaluationOn-Policy Distillation

  5. Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

    Sep 18, 2026Nikhil Reddy Pottanigari, Ramin Fahimi, Noah Bolger +2LLM EvaluationText Summarization

  6. Chinese Competitive Debating Dataset and Benchmark

    Sep 18, 2026Zongrui Yang, Haoyuan Li, Zhongsheng Wang +5LLM Evaluation

  7. Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

    Sep 17, 2026Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak +3LLM EvaluationSelective Prediction

  8. Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

    Sep 17, 2026Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh DungLLM EvaluationRecurrent Transformers

  9. KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

    Sep 17, 2026Soha Lee, Soojin Lee, Heesung Yang +8LLM EvaluationLanguage Modeling

  10. PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

    Sep 17, 2026Pyrros Koussios, Benjamin Jäger, John Hua Yao +3LLM EvaluationTemporal Reasoning in Language Models

  11. Reproducibility is not construct validity: LLM measurement of institutionally situated communication

    Sep 17, 2026Veronika Batzdorfer, Carlo Romano Marcello Alessandro SantagiustinaLLM EvaluationLLM-Assisted Annotation

  12. F2^{2}DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

    Sep 17, 2026Bojian Xiong, Wentao Ding, Yujing Lu +11Reward ModelingLLM Evaluation

  13. A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    Sep 17, 2026Shaina Raza, Ahmed Y. Radwan, Imran Liaquat +1LLM EvaluationAI Agent Evaluation

  14. A Benchmark Framework for Screening Automation in Systematic Reviews

    Sep 16, 2026Gauransh Kumar, Luciano Marchezan, Guillaume Genois +2LLM EvaluationEvidence Selection

  15. Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

    Sep 16, 2026Leon Bergen, Usha Bhalla, Andrew Lee +15Reward HackingLLM Evaluation

  16. Which LLM is Best for Translating Natural Language Goals to PDDL

    Sep 16, 2026Tomas Balyo, Lukas Chrpa, G. Michael YoungbloodLLM EvaluationLLM Planning

  17. Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

    Sep 16, 2026Otto Segersven, Pentti HenttonenMultilingual Language Model EvaluationLLM Evaluation

  18. GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models

    Sep 16, 2026Alexander Didenko, Anna Shabanova, Vladislav Zapylikhin +2LLM EvaluationCross-Cultural Language Model Evaluation

  19. Made in Hungary: Comments on the performance of generative language models

    Sep 16, 2026Mátyás Osváth, Enikő Héja, Noémi Ligeti-NagyLLM EvaluationLanguage Model Generation Evaluation

  20. Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement

    Sep 16, 2026Xinglang Zhang, Yuanmeng Xiang, Yunyao Zhang +3LLM EvaluationPreference Alignment

  21. A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

    Sep 16, 2026Jerry KaplanLLM QuantizationLLM Evaluation

  22. Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

    Sep 15, 2026Adam Zachary Wasserman, David BeaucheminLLM EvaluationCross-Lingual Representation Alignment

  23. Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

    Sep 15, 2026Vanessa Lama, Sanjay Das, Emily Herron +5LLM EvaluationScientific QA

  24. Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising

    Sep 15, 2026Md. Samiul Alim, Mahir Shahriar Tamim, Tanvir Ahmed Khan +4LLM EvaluationSentiment Analysis