Large Language Model Agents

Latest papers 875

All topics
CardsList
  1. Learning to Accumulate Knowledge with Mutual Information

    Oct 7, 2026Yuyang Zhao, Lizi Liao, Leyang Shen +4Large Language Model Reinforcement LearningLarge Language Model Agents

  2. Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

    Oct 1, 2026Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed +4NativememLarge Language Model Agents

  3. Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    Oct 1, 2026Yu Luo, Jiamin Jiang, Yimin Zuo +9Long-Horizon AgentsLong-Horizon Task Planning

  4. DeFA: Dependency-Guided Failure Attribution for LLM Agents

    Oct 1, 2026Bo Deng, Xinlei Zheng, Yi Wei +6Large Language Model Agents

  5. Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

    Oct 1, 2026Haotian Chen, Bowen Ye, Yuning Zhang +1Large Language Model AgentsContracts

  6. ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

    Oct 1, 2026Sungho Park, Wonjoong Kim, Jue Zhang +8Agentic OptimizationAdaptive Curriculum

  7. PhantomEnvironments: Training LLM Agents in Fictional Worlds

    Sep 30, 2026Anmol Kabra, Swathi Saravana Selvam, Albert Gong +5Large Language Model AgentsSynthetic Environments

  8. OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation

    Sep 30, 2026Oana Madalina Fron, Ojas Shirekar, Chirag RamanLarge Language Model AgentsTactics

  9. Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

    Sep 30, 2026Haoyu Wang, Wei Zhao, Yedi Zhang +2Large Language Model Agents

  10. SkillFM: Generating Skills for LLM Agents via Latent Flow Matching

    Sep 30, 2026Zuming Zhang, Jie He, Yizhe Zhang +1SkillsLarge Language Model Agents

  11. Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

    Sep 30, 2026Wenxin Wu, Lingyong Yan, Lei Sha +2PoisoningMalicious Agents

  12. Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

    Sep 30, 2026Kaixing Zhang, Changming Li, Yingdong Shi +5Skill EvolutionSelf-Evolution

  13. Schema: Discovering Unknown Environments via Agentic Program Induction

    Sep 30, 2026Guanning Zeng, Jiani Wang, Wenjie Ma +8Large Language Model AgentsAgentic

  14. When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

    Sep 30, 2026Shuyao Xiao, Shengling Wang, Xuan Chen +7Large Language Model AgentsCausal Intervention

  15. Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

    Sep 30, 2026Yan Wang, Zhihao Zhang, Ke Chen +5AI TrustworthinessLarge Language Model Agents

  16. Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents

    Sep 29, 2026Jiacheng Qiu, Christopher E. Mower, Jan Peters +2Large Language Model AgentsAgentic

  17. From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs

    Sep 29, 2026Kunal Jha, Max Kleiman-Weiner, Natasha JaquesSelf-Improving AgentsLarge Language Model Agents

  18. Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

    Sep 29, 2026Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera +2Large Language Model PlanningLarge Language Model Agents

  19. When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

    Sep 29, 2026Yuqiao Meng, Luoxi Tang, Yingxue Zhang +2FuzzingPeak Memory

  20. EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

    Sep 29, 2026Min Yang, Yichen Pan, Jinghua Piao +3Large Language Model AgentsStrategic Reasoning

  21. FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

    Sep 29, 2026Shantanu Dixit, Anson Bastos, Xuchao Zhang +2Large Language Model AgentsInteraction History

  22. Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

    Sep 29, 2026Jio Oh, Seunghyun Do, Young-Jun Lee +2Large Language Model AgentsUser Simulation

  23. ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

    Sep 29, 2026Yanjie Li, Xiangyu He, Xuelong Dai +1AuthorizationLanguage-Model Agents

  24. SkillCome: Group Contrast Skill Optimization with Dual Memory

    Sep 29, 2026Haolin Li, Feng Hong, Ang Li +6SkillsLarge Language Model Agents

  25. SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

    Sep 29, 2026Yu Cheng, Yongkang Hu, Shuaijie Ma +12SaferLarge Language Model Agents

  26. LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

    Sep 29, 2026Hugo Lyons Keenan, Christopher Leckie, Sarah ErfaniModel ActivationsEvasion

  27. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    Sep 28, 2026Saswat Das, Parvati Viswanathan, Daniel Donnelly +3Self-Evolving AgentsSelf-Evolution

  28. Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

    Sep 28, 2026Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik +4Large Language Model AgentsAdversarial Robustness

  29. Continuous Context Management

    Sep 28, 2026William Hoy, Jingxuan Fan, Nurcin Celik +1Large Language Model AgentsProbe-Logit Distillation

  30. PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents

    Sep 28, 2026Lucas Biechy, Cédric Eichler, Héber H. Arcolezi +1PrivacyLarge Language Model Agents

  31. Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

    Sep 28, 2026Huachi Zhou, Yujing Zhang, Jiahe Du +7Data Science AgentsAgentic Workflow Design

  32. When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents

    Sep 28, 2026Geonwoo Kim, Brent ByungHoon KangAgentic DeploymentsCall

  33. Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

    Sep 28, 2026Yinhong Liu, Zhili Tan, Zilin Wang +1Large Language Model AgentsSchema

  34. PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents

    Sep 28, 2026Huayi Lai, Shichao Song, Qingchen Yu +4Large Language Model Agents

  35. CoSec: Benchmarking Agent Security in Communities

    Sep 28, 2026Hao Chen, Wenhui Dong, Ye Chen +12AuthorizationSecurity

  36. FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

    Sep 28, 2026Peiyu ZangLong-Horizon AgentsAgentic Benchmarks

  37. Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

    Sep 28, 2026Weiyuan Li, Jinghan Xu, Aili Chen +4Long-Horizon AgentsLarge Language Model Agents

  38. FlowState: Execution State as Memory for Long-Horizon LLM Agents

    Sep 28, 2026Minghao Li, Bangyan Li, Zifan Wang +5Long-Horizon AgentsLong-Horizon Task Planning

  39. Remember Before You're Asked: MemDream for Self-Probing Memory Evolution

    Sep 28, 2026Mingfei Lu, Mengjia Wu, Runsong Jia +2Large Language Model AgentsDream

  40. Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

    Sep 28, 2026Tong Zhao, Reed Li, Yuyang Hu +5Large Language Model AgentsNext

  41. PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents

    Sep 28, 2026Hanzhong Zhang, Ziwei Xiang, Weicheng Xie +2PersonalityRole-Playing Agents

  42. SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing

    Sep 28, 2026Zhiwei Chen, Tianchun Wang, Zhongtao Rao +3Strategic ReasoningLarge Language Model Agents

  43. ControlScope: Workflow Revision and Reliability in LLM Agents

    Sep 28, 2026Jingjie Ning, Xueqi Li, Yibo Kong +1Agentic Workflow DesignLarge Language Model Agents

  44. Certified Multi-Source Integrity for Structured Agent Actions

    Sep 28, 2026Anmol Pandey, Aditya Jain, Liang Chen +2IntegrityOn-Chain Attestations

  45. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5Task Success RateLarge Language Model Agents

  46. From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

    Sep 28, 2026Mingxi Zou, Langzhang Liang, Zhuo Wang +3Large Language Model AgentsHarms

  47. Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

    Sep 28, 2026Pengxin Wang, Yuanzhe LI, Yuxin Ren +2Frictive Policy OptimizationLarge Language Model Agents

  48. When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents

    Sep 27, 2026Zhihao Zhang, Chao Wang, Rujia Li +3AuthorizationAuthority

  49. When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

    Sep 27, 2026Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur +2Large Language Model AgentsRapid Adaptation

  50. LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

    Sep 27, 2026Haochen Luo, Yifan Li, Binh Minh An +4Algorithmic TradingLarge Language Model Agents

  51. HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents

    Sep 27, 2026Zhuowen Liu, Zhixuan WangTriageLarge Language Model Agents