LLM Agent Evaluation

LLM: Large Language Model

Momentum

115 papers in the last four weeks, up 140% on the four weeks before. 1.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 793

All topics
CardsList
  1. Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

    May 28, 2026Francisco León Zúñiga BolívarMulti-Agent LLM SystemsLLM Agent Evaluation

  2. Realistic honeypot evaluations for scheming propensity

    May 28, 2026Victoria Krakovna, David Lindner, Lewis Ho +2LLM Agent EvaluationAI Agent Safety

  3. PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?

    May 28, 2026Dongdong Hua, Yifei Sun, Renhong Huang +3Game-Playing AgentsLLM Agent Evaluation

  4. BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

    May 28, 2026Jiahao Huang, Fei Cheng, Junfeng Jiang +2LLM Agent EvaluationAI Agent Benchmarks

  5. Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

    May 27, 2026Diego Gosmar, Deborah A. DahlAI Agent ReliabilityMulti-Agent LLM Systems

  6. Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

    May 27, 2026James P. Balhoff, Hilmar LappLLM-Assisted AnnotationLLM Agent Evaluation

  7. Do Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval

    May 27, 2026Shiyu Chen, Tarfah Alrashed, Alon Halevy +1LLM Agent EvaluationAgentic Retrieval

  8. LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

    May 27, 2026HuiMing Fan, Xiao Wang, Zheng Chu +5Computer-Use Agent BenchmarksLLM Agent Evaluation

  9. Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

    May 27, 2026Alejandro Buitrago López, Alberto Ortega Pastor, Javier Pastor-Galindo +1LLM Agent EvaluationLarge Language Model-Based Social Simulation

  10. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    May 27, 2026Tomer Keren, Nitay Calderon, Asaf Yehudai +3Computer-Use Agent BenchmarksTool-Use Evaluation

  11. Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

    May 27, 2026Jihyeong Park, Ingeol Baek, Jeonghyun Park +1LLM Agent EvaluationAI Agent Benchmarks

  12. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

    May 27, 2026Taojie Zhu, Wentao Zhao, Rui Sun +7Benchmark ContaminationLLM Agent Evaluation

  13. OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

    May 27, 2026Chenyu Zhou, Xinyun Lu, Jiangyue Zhao +3LLM Agent EvaluationAI Agent Benchmarks

  14. Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

    May 27, 2026Bin Wu, Guanyun Zou, Bingbing Wang +2LLM Agent MemoryLLM Agent Evaluation

  15. UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

    May 27, 2026Pengyu Zhu, Lijun Li, Yaxing Lyu +8LLM Agent EvaluationAI Agent Benchmarks

  16. Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

    May 26, 2026Aman Priyanshu, Supriti Vijay, Esha PahwaData LeakageMulti-Agent LLM Systems

  17. UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind

    May 26, 2026Cheng Qian, Jiayu Liu, Heng JiTheory of MindLLM Agent Evaluation

  18. Voluntary Collusion with Secret Tools in Competing LLM Agents

    May 26, 2026Xijie Zeng, Frank RudziczLLM Agent EvaluationLLM Agent Safety

  19. DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

    May 26, 2026Shijie Cao, Yuan Yuan, Jing LiuLLM Agent EvaluationCombinatorial Optimization

  20. Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

    May 26, 2026Yipeng Ouyang, Xin Huang, Bingjie Liu +3Software Engineering AgentsLLM Agent Evaluation

  21. QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

    May 26, 2026Ye Yuan, Rui Song, Weien Li +12Social Deduction GamesLLM Agent Evaluation

  22. Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

    May 26, 2026Yee Hin Chong, Jiaming Wu, Youhui Zhang +1GPU Kernel OptimizationLLM Agent Evaluation

  23. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

    May 26, 2026Anmol Agarwal, Natalie Neamtu, Pranjal Aggarwal +6LLM Agent EvaluationFormal Verification