LLM Guardrails

LLM: Large Language Model

Momentum

15 papers in the last four weeks, up 36% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 138

All topics
CardsList
  1. One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

    Oct 8, 2026Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud PedramLLM GuardrailsLLM Agent Security

  2. Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models

    Oct 8, 2026Youwei Feng, Yitong Zhang, Yuetong Liu +1LLM Agent SafetyLLM Guardrails

  3. Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

    Oct 7, 2026Georgios Koutidis, Nikolaos Kekatos, Tom Nianios +1LLMs for CybersecurityTool Access Control for LLM Agents

  4. POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

    Oct 6, 2026Yunju Kang, Seonghyeon Cho, Irene Li +2LLM GuardrailsRuntime Enforcement for AI Agents

  5. Benchmarking Jailbreak Guardrails for Embodied Agents

    Oct 5, 2026Xunguang Wang, Qingyue Wang, Yuguang Zhou +3LLM GuardrailsAI Agent Safety

  6. Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

    Sep 30, 2026Gou Tan, Pengfei Chen, Zhensu Sun +9Automated Program RepairLLM Guardrails

  7. PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

    Sep 28, 2026Ding Jia, Wei Liu, Xianglong Du +5LLM GuardrailsRuntime Enforcement for AI Agents

  8. AdaGuard: An Adaptive Guard Model with User-defined Policies

    Sep 28, 2026Yunhao Feng, Yifan Ding, Yuxiang Xie +4LLM GuardrailsAI Agent Safety

  9. Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

    Sep 27, 2026Gert Lek, Abele Malan, Chaoyi Zhu +3LLM Safety BenchmarksLLM Guardrails

  10. Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming

    Sep 24, 2026Madeleine Eastwood, Harshith Narne, Joseph Hilby +3LLM GuardrailsProgramming Education

  11. Visual Compliance via Executable Safety Rule Entailment

    Sep 16, 2026Jisoo Kim, TaeYoon Kwack, Jinwoo Jang +1Logical ReasoningLLM Guardrails

  12. SAGE: Governed Artifact Generation from Enterprise Guidelines

    Sep 15, 2026Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli +2LLM Hallucination MitigationLLM Guardrails

  13. HazardAuditor: From Executable Threats to Safer Computer-Use Agents

    Sep 14, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +5Language Model Safety EvaluationAI Agent Auditing

  14. Overflip: Repetition-Induced Label Flips in Guardrail Models

    Sep 14, 2026Xu He, Chih-Hsuan Lin, Hung-Mao Chen +3LLM GuardrailsAdversarial Attacks on LLMs

  15. DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

    Sep 11, 2026Abhinav Rajeev Kumar, Harshit Arora, Varun Singh +1LLM Safety BenchmarksLLM Guardrails

  16. CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

    Sep 9, 2026Jinyang Li, Mingyu Guo, Hung X. NguyenLLM GuardrailsAdversarial Attacks on LLMs

  17. Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

    Sep 8, 2026Yongxi Zhou, Wenbo Ye, Yuanzhe Liu +2LLM Safety BenchmarksLLM-as-a-Judge

  18. When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

    Sep 1, 2026Peiying Zhu, Sidi ChangLLM GuardrailsLLM Agent Evaluation

  19. The Safeguard Worked. Is the LLM System Safer?

    Sep 1, 2026Pingyu Wu, Weiming Zhang, Nenghai YuLLM GuardrailsLLM Safety

  20. SingProbe Technical Report

    Aug 31, 2026Sing TeamLanguage Model Safety EvaluationLLM Guardrails

  21. Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents

    Aug 30, 2026Tanzim Ahad, Ismail Hossain, Md Jahangir Alam +3LLM GuardrailsLLM Agent Security

  22. LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails

    Aug 27, 2026Ziyang Chen, Xing Wu, Songlin HuLanguage Model Safety EvaluationLong-Context Language Model Evaluation