AI Agent Evaluation

Momentum

58 papers in the last four weeks, up 100% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 341

All topics
CardsList
  1. Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

    Nov 24, 2025Dayong Liu, Chao Xu, Weihong Chen +5AI Agent EvaluationMultimodal Large Language Models

  2. TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward

    Oct 10, 2025Yincen Qu, Huan Xiao, Feng Li +5Reward ModelingLLM Planning

  3. AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

    Jun 4, 2025Dhaval Patel, Shuxin Lin, James Rayfield +7LLM Agent OrchestrationAI Agent Evaluation

  4. When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

    Jul 15, 2024Chong Zhang, Xinyi Liu, Zhongmou Zhang +10Multi-Agent LLM SystemsHuman Behavior Simulation

  5. An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

    Date pendingHong Zhang, Barry Smith, Satish Balay +4High-Performance ComputingCode Generation

  6. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice Agent EvaluationSpoken Dialogue Systems

  7. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation

  8. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

    Date pendingYuqiao Tan, Shizhu He, Jun Zhao +1AI Agent EvaluationSparse Autoencoders