LLM Agent Verification

LLM: Large Language Model

Momentum

18 papers in the last four weeks, up 260% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 72

All topics
CardsList
  1. Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems

    Aug 2, 2026Juan Li, Wei Cai, Yan BaiMulti-Agent CollaborationLLM Agent Verification

  2. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

    Jul 31, 2026Bohan Chen, Shivam N. Patel, Richard Hoffmann +2LLM Agent VerificationTool-Using Agents

  3. VeriSkill: A Self-Evolution Framework for Program Verification Skills

    Jul 30, 2026Changguo Jia, Tianqi Zhao, Zhiyou Xiao +2LLM Agent Self-ImprovementFormal Verification

  4. Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

    Jul 28, 2026Genliang Zhu, Chu WangLLM Agent SecurityLLM Agent Verification

  5. Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

    Jul 27, 2026Diandian Guo, Cong Cao, Fangfang Yuan +3AI Agent EvaluationLLM Agent Self-Improvement

  6. Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

    Jul 20, 2026Yitao Wu, Si Shen, Rui Yang +2AI Agent ReliabilityLLM Self-Correction

  7. AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

    Jul 13, 2026Aritra Mazumder, Nusrat jahan LiaAI Agent ReliabilityLLM Agent Evaluation

  8. LLM-as-a-Verifier: A General-Purpose Verification Framework

    Jul 6, 2026Jacky Kwok, Shulu Li, Pranav Atreya +6LLM Answer VerificationLLM Agent Evaluation

  9. SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

    Jun 29, 2026Aojie Yuan, Yi Nian, Haiyue Zhang +2AI Agent ReliabilityProcess Reward Models

  10. The Verification Horizon: No Silver Bullet for Coding Agent Rewards

    Jun 24, 2026Binghai Wang, Chenlong Zhang, Dayiheng Liu +10Reward ModelingReward Hacking

  11. SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills

    Jun 14, 2026Ismail Hossain, Sai Puppala, Md Jahangir Alam +2LLM-as-a-JudgeLLM Agent Security

  12. MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

    Jun 10, 2026Abdelrahman Abdallah, AbdelRahim A. Elmadany, Sameh Al Natour +3Table QALLM Answer Verification

  13. OpenSkill: Open-World Self-Evolution for LLM Agents

    Jun 4, 2026Zhiling Yan, Dingjie Song, Hanrong Zhang +8LLM Agent Skill LearningLLM Agent Verification

  14. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

    Jun 2, 2026Ruida Wang, Jerry Huang, Pengcheng Wang +3Software Engineering AgentsFormal Verification

  15. Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

    Jun 2, 2026Thanh Luong Tuan, Abhijit SanyalLLM Agent VerificationAI Agent Governance

  16. LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

    May 27, 2026HuiMing Fan, Xiao Wang, Zheng Chu +5Computer-Use Agent BenchmarksLLM Agent Evaluation

  17. FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

    May 26, 2026Haoxuan Jia, Yang Liu, Bin Chong +10AI Agent ReliabilityLLM Guardrails

  18. QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

    May 26, 2026Ye Yuan, Rui Song, Weien Li +12Social Deduction GamesLLM Agent Evaluation