LLM Agent Verification

LLM: Large Language Model

Momentum

18 papers in the last four weeks, up 260% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 72

All topics
CardsList
  1. Skill-V: Verifiable Self-Evolving Skill Library for Interactive Agents

    Oct 8, 2026Jie Ma, Zhipeng Qian, Yufei Ma +6LLM Agent Skill LearningLLM Agent Skill Retrieval

  2. What Output-Only Review Cannot Verify: Study Contracts for Research Agents

    Oct 8, 2026Eitan Waks, Ben GlockerLLM Agent VerificationAI Agent Auditing

  3. Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents

    Oct 6, 2026Nikolaos Kekatos, Dimitrios Nikou, Anastasios Temperekidis +4LLM Agent VerificationRuntime Enforcement for AI Agents

  4. VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    Oct 1, 2026Caiqi Zhang, Rujun Han, Zifeng Wang +4LLM Agent VerificationLong-Horizon LLM Agents

  5. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

    Sep 30, 2026Minki Kang, Ryo Hachiuma, Shaokun Zhang +8Terminal AgentsTest-Time Scaling

  6. NarrativeSteward: Coordinating Delegation, Guidance, and Verification in Agent-Assisted Interactive Narrative Authoring

    Sep 30, 2026Wenjin Wang, Jiazhen Lei, Yuxin Sha +6Interactive StorytellingLLM Agent Orchestration

  7. Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification

    Sep 30, 2026Yingfeng Luo, Shaowei Wei, Daixin Wang +7LLM Agent VerificationLLM Agent Reliability

  8. Solver Agent: an Agentic AI Framework for Theoretical Physics Computations Applied to F-theory Uplifts of O3-planes and S-folds

    Sep 28, 2026Eliott Morgensztern, Cesar Fierro Cota, Alessandro MininnoAI Agents for Scientific DiscoveryLLM Agent Verification

  9. StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents

    Sep 28, 2026Wenle Liao, Zhao Wang, Jingchao Zhang +3LLM Agent VerificationGraph-Based Agent Memory

  10. Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

    Sep 28, 2026Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook +3Multi-Agent LLM SystemsLLM Grounding

  11. Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory

    Sep 27, 2026Jayant Parashar, Eugene F. Douglass, William C. Bastian +1LLM Agent MemoryAgent Harness Optimization

  12. Who Holds the Pen? Let Specifications, Not Agents, Sign Off

    Sep 24, 2026Haiqing Li, Xin Ma, Yinhao Wu +7LLM Agent VerificationRuntime Enforcement for AI Agents

  13. Calibration Is Not Verification: Falsifiability-Aware Conformal Routing for Mixture-of-Agents

    Sep 22, 2026Nada Rahali, Zijia Wang, Zhisong LiuFactual Consistency EvaluationConformal Prediction

  14. How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

    Sep 17, 2026Yukun Zhang, Kemu Xu, Yishen ChenAgent Harness OptimizationLLM Agent Verification

  15. Look Before You Leap: Pre-Action Verification for LLM Agents

    Sep 14, 2026Asaad AlthoubiLLM Agent VerificationAI Agent Benchmarks

  16. Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

    Sep 9, 2026Tianzhu Zhang, Chih-Kai Huang, Meikang QiuMulti-Agent OrchestrationLLM Agent Verification

  17. Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

    Sep 7, 2026Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2Tool Use in VLMsMultimodal QA

  18. VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

    Sep 2, 2026Wenzhuo Xu, Yuchen Zhu, Chongjian Ge +8Physical Consistency in Video GenerationLLM Agent Verification

  19. An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

    Aug 12, 2026Yuzhong Shen, Masha Sosonkina, Peng Xu +1Software Engineering AgentsAgentic Workflows

  20. MIRA: Medical Image Reflection for Agentic Diagnosis

    Aug 11, 2026Shengzhi Wang, Jun Yang, Kai Wu +11Medical Image AnalysisLLM Agent Verification

  21. SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

    Aug 6, 2026Yuru Feng, Yaoqi Chen, Beidi Zhao +7Test-Time AdaptationLLM Agent Self-Improvement

  22. VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space

    Aug 3, 2026Yu-Tung Liu, Cunxi YuLLM Agent VerificationRTL Code Generation