Agentic Benchmarks

Latest papers 306

All topics
CardsList
  1. ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

    Oct 5, 2026Fan Li, Xiangyu Gao, Zixuan Liu +4Agentic Benchmarks

  2. Sharpening Tax in Post-Training

    Oct 1, 2026Changdae Oh, Qi Zeng, Qi Qi +7Reinforcement Learning Post-TrainingPost-Training

  3. Finding the Right Fit: Model-Harness Interactions across Agent Tasks

    Oct 1, 2026Yixuan Li, Yiyun Zhou, Yao Long Teng +6Agent HarnessAgentic Benchmarks

  4. Cross-Benchmark Transfer from RL on Agentic Coding Tasks

    Oct 1, 2026Sushant Mehta, Logan Ritchie, Edwin ChenCoding AgentsAgentic Benchmarks

  5. EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

    Sep 30, 2026Zhixuan Tan, Pengjie Gu, Zhao Li +4Coding AgentsAgentic Benchmarks

  6. AgBench: Agentic AI Benchmarks for Personal AI Devices

    Sep 29, 2026Yizhou Han, Di Wu, Dhananjay Saikumar +1Agentic BenchmarksArtificial Intelligence Systems

  7. FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

    Sep 29, 2026Shantanu Dixit, Anson Bastos, Xuchao Zhang +2Large Language Model AgentsInteraction History

  8. Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

    Sep 29, 2026Rohith Reddy Bellibatlu, Zichong Wang, Wenbin ZhangAgentic BenchmarksModel Auditing

  9. MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

    Sep 29, 2026Lei Ma, Dennis Hofmann, Haowen Xu +4Multimodal Anomaly DetectionAgentic Benchmarks

  10. CheatBench: Measuring Reward Gaming in AI Agents

    Sep 28, 2026Long Phan, Stephen K. Yang, Jason J. Lim +10Agentic BenchmarksArtificial Intelligence Agents

  11. Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

    Sep 28, 2026Gunin Gupta, Nirmit Arora, Pavan Kalyan TankalaVideo EditingVideo Agent

  12. RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents

    Sep 28, 2026Hao Li, Hangfan Zhang, Zhiyao Cui +5Large Language Model RoutingInference Cost

  13. FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

    Sep 28, 2026Peiyu ZangLong-Horizon AgentsAgentic Benchmarks

  14. AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

    Sep 28, 2026Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6Agentic BenchmarksAgentic Inference

  15. AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

    Sep 28, 2026Chanhee Park, Jeongho Yoon, Sungbin Han +2Agentic BenchmarksDiagnostic Benchmark

  16. PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models

    Sep 28, 2026Shane K. A. Dalumura Hettige, Jonas OppenlaenderCreativityCanvas

  17. Agentic Multi-Turn Reasoning: A Fairness Approach

    Sep 27, 2026Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill +3Agentic BenchmarksCredit Assignment

  18. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

    Sep 27, 2026Dehai Min, Daoan Zhang, Yiming Zeng +13Agentic BenchmarksArtificial Intelligence Agents

  19. Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

    Sep 24, 2026Benjamin Gruenbaum, Doron Porat, Assaf Natanzon +3Agentic BenchmarksQuestion

  20. Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

    Sep 24, 2026Xincheng Yao, Haobo Fu, Weiming Liu +1Group-Based Reinforcement LearningAgentic Reinforcement Learning

  21. Reinforcement Learning with Decomposed Subtasks

    Sep 22, 2026Mattie Terzolo, Mikolaj Sacha, Ayan Sinha +1Policy GradientAgentic Benchmarks

  22. SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

    Sep 22, 2026Jennifer Williams, Dave Farris, Jeff Farris +1Agentic InferenceAgentic Benchmarks

  23. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

    Sep 22, 2026Wenbo Pan, Zhichao Liu, Shujie Liu +6Long-Horizon AgentsAgentic Benchmarks

  24. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

    Sep 21, 2026Yiran Wang, Xingyilang Yin, Junfu Pu +12Agentic BenchmarksVideo World Models

  25. MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

    Sep 21, 2026Demetris Paschalides, Moysis Symeonides, George Pallis +1Agentic BenchmarksLarge Language Model Agents

  26. BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

    Sep 20, 2026Peng Kuang, Yuchun Fan, Jiangnan Li +7Multilingual AgentsMultilingual Benchmark

  27. MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

    Sep 17, 2026Pritish Mishra, Ishaan Kumar, Akshat Mandoli +1Voice AgentsDialogue Benchmarks

  28. Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    Sep 16, 2026Xinshuai Guo, Junjie Wu, Dolly Deng +4Agentic BenchmarksLarge Language Model Benchmarks

  29. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    Sep 16, 2026Jeonghye Kim, Minseon Kim, Young Jin Kim +5Coding AgentsSoftware Engineering

  30. Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

    Sep 14, 2026Baoyang Jiang, Fengchun Zhang, Leyuan Wang +9Embodied AgentsAgentic Benchmarks

  31. MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

    Sep 14, 2026Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin +7Voice AgentsDialogue Benchmarks

  32. MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

    Sep 14, 2026Bosi Wen, Cunxiang Wang, Jiayi Gui +6Agentic Benchmarks

  33. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

    Sep 12, 2026Pingchen Lu, Xiangyi Wang, Xiang Li +6Self-Evolving Skill LibrariesAgent Skill Retrieval

  34. Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

    Sep 11, 2026Hanhua Hong, Yizhi Li, Luu Gia Huy +3Agentic BenchmarksReproducibility

  35. Benchmarking Hybrid Deep Research Across Database Querying and Web Search

    Sep 8, 2026Ruofan Wu, Peiran Xu, Xiaolong Li +9Deep ResearchAgentic Search

  36. NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management

    Sep 7, 2026Yulin Wei, Xiangchen Wang, Jianhui Pan +5DietAgentic Benchmarks

  37. ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

    Sep 7, 2026Weide Zhan, Qumu Shaqu, Yuanqing Liu +6Graphical User Interface AgentsAgentic Benchmarks

  38. ττ\tau^\tau-Bench: An Environment for End-To-End, Realistic Agent Construction

    Sep 7, 2026Quan Shi, Keshav Dhandhania, Karthik Narasimhan +1Agentic BenchmarksCoding Agents

  39. ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

    Sep 3, 2026Shidu Ren, Yunze Liu, Xing Liu +3Multimodal MemoryMultimodal Agents

  40. WorldBench: Culturally Grounded Benchmark for Multilingual Agents

    Sep 1, 2026Leonardo Ranaldi, Sherrie Shen, Jushi Kai +1Multilingual AgentsAgentic Benchmarks

  41. BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

    Aug 31, 2026Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3Agentic BenchmarksExploitation

  42. HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving

    Aug 31, 2026Boyang Mu, Zhiwei Wei, Mugen Peng +1Remote SensingMulti-Agent Communication

  43. SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

    Aug 31, 2026Jinshan Gao, Zhuoran Jin, Tianyi Men +2Multi-Agent OrchestrationAgentic Benchmarks

  44. AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

    Aug 30, 2026Xinke Jiang, Yue Fang, Zhibang Yang +12Agentic Reinforcement LearningAgentic Benchmarks

  45. Learning Generalizable Behaviors for Terminal Agents

    Aug 23, 2026Yihang Yao, Bo Pang, Xuan Phi Nguyen +3Agentic Reinforcement LearningAgentic Benchmarks

  46. DiG-bench: Discovery in Games

    Aug 12, 2026Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13Agentic BenchmarksArtificial Intelligence Benchmarks

  47. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

    Aug 12, 2026Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder +6Agentic BenchmarksReasoning Benchmark

  48. Benchmarking LLM Judges for Mobile Agent Evaluation

    Aug 11, 2026Ziqiang Wan, Li Gu, Zhixiang Chi +4Llm-As-A-JudgeAgentic Benchmarks