Agent Evaluation

Momentum

23 papers in the last four weeks, up 229% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 97

All topics
CardsList
  1. Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

    Mar 4, 2026Ji-Lun Peng, Yun-Nung ChenPersonality Modeling in Language ModelsOOD Generalization

  2. MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration

    Mar 1, 2026Abdulhamid M. Mousa, Jinhui Pang, Rakhmonberdi Khajiev +5AI Agent EvaluationAgent Evaluation

  3. GameDevBench: Evaluating Agentic Capabilities Through Game Development

    Feb 11, 2026Wayne Chi, Yixiong Fang, Arnav Yayavaram +8Software Engineering AgentsAgent Evaluation

  4. DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems

    Jan 20, 2026Maojun Sun, Yifei Xie, Yue Wu +5LLM Agent EvaluationAgent Evaluation

  5. StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

    Jul 10, 2025Weihao Tan, Changjiu Jiang, Yu Duan +5Game-Playing AgentsAgent Evaluation

  6. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice Agent EvaluationSpoken Dialogue Systems