cs.AIOct 3, 2026

Agentic Cognitive Depth: Operational Criteria for Evaluating LLM Agents

Authors: Nijesh Upreti, Chris Sypherd, Vaishak Belle

Organizations: The University of Edinburgh 10 Crichton Street, Edinburgh EH8 9AB, UK

Abstract

Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view connects agentic LLM research with deployable systems and leaves open how such systems should be evaluated beyond end-to-end task success. Building on this view, we define agentic cognitive depth as a trajectory-level profile across five operational criteria. The profile contains context sensitivity (CC), temporal continuity (TT), multimodal coordination (MM), adaptive interaction (AA), and metacognitive monitoring (McMc). The first four criteria measure how well Control Flow, Memory, Tools, and Planning are used across a trajectory. The fifth measures whether the system monitors and regulates the full run. For each criterion, we give operational proxies and a perturbation procedure, then connect the profile to the agent's world model. We provide the structure needed to extend benchmarks such as GAIA, SWE-bench, WebArena, and TRIP-Bench with per-criterion diagnostics. Symbolic verifiers, structured memory, planner coupling, and tool constraints provide practical ways to build and test these capacities.

Explore similar work

CardsList
  1. UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

    May 27, 2026Pengyu Zhu, Lijun Li, Yaxing Lyu +8LLM Agent EvaluationAI Agent Benchmarks

  2. Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

    May 21, 2026Asaf Yehudai, Lilach Eden, Michal Shmueli-ScheuerLLM Agent EvaluationAI Agent Benchmarks

  3. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

    Aug 27, 2026Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian +1LLM-as-a-JudgeLLM Agent Evaluation