LLM Agent Verification

LLM: Large Language Model

Momentum

18 papers in the last four weeks, up 260% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 72

All topics
CardsList
  1. OpenComputer: Verifiable Software Worlds for Computer-Use Agents

    May 19, 2026Jinbiao Wei, Qianran Ma, Yilun Zhao +4Computer-Use Agent BenchmarksComputer-Use Agents

  2. No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

    May 13, 2026Ying Li, Hongbo Wen, Yanju Chen +3LLM GuardrailsLLM Agent Verification

  3. Behavioral Integrity Verification for AI Agent Skills

    May 12, 2026Yuhao Wu, Tung-Ling Li, Hongliang LiuAgent ReliabilityAI Agent Auditing

  4. MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

    May 7, 2026Ashwani Anand, Ivi Chatzi, Ritam Raha +1Computer-Use Agent BenchmarksAI Agent Reliability

  5. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

    Apr 20, 2026Wentao Shi, Yu Wang, Yuyang Zhao +8LLM-as-a-JudgeLLM Agent Evaluation

  6. Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

    Apr 1, 2026Alibek Kaliyev, Artem MaryanskyyAI Agent ReliabilityTool-Use Evaluation

  7. When Is Enough Not Enough? Illusory Completion in Search Agents

    Feb 7, 2026Dayoon Ko, Jihyuk Kim, Sohyeon Kim +5AI Agent ReliabilityLLM Answer Verification

  8. Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry

    Oct 29, 2025Run Peng, Ziqiao Ma, Amy Pang +5Multi-Agent LLM SystemsLLM Agent Verification