Unsafe

Momentum

2 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 6Week of Sep 21

Latest papers 28

All topics
CardsList
  1. Hard-Gate Candidacy in a Deployed Validator Suite

    Sep 30, 2026Xin XuGatingValidation

  2. Inoculation Midtraining with Learned Neologisms

    Sep 14, 2026Kyle O'Brien, Edward James Young, Puria Radmard +4Large Language Model SafetyInoculation Prompting

  3. ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

    Sep 10, 2026Jianming Ma, Rongjun Jin, Xiaxi Si +3Safety ConstraintsSafer

  4. Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents

    Aug 30, 2026Tanzim Ahad, Ismail Hossain, Md Jahangir Alam +3Controlling Tool UseGuardrail

  5. DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

    Aug 6, 2026Wenhao Lin, Chenyu Yu, Xingwei Lin +6Runtime Safety FilteringLarge Language Model Agents

  6. Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems

    Aug 2, 2026Wanguang Li, Zhaoxin Wang, Handing WangAdversarial PromptsJailbreak Success Rates

  7. Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

    Aug 1, 2026Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4Large Language Model SafetyUnsafe

  8. Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

    Jul 28, 2026Elias Fernández Domingos, The Anh HanArtificial Intelligence SafetyArtificial Intelligence Systems

  9. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

    Jul 21, 2026Timothy TomashevskiySafety ConstraintsOffline Reinforcement Learning

  10. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    Jul 15, 2026Joshua A. Kroll, Andrew Smart, R. Stuart Geiger +1Socio-Technical SystemsArtificial Intelligence Risk

  11. Avoiding unsafe sets when training with Langevin Dynamics

    Jul 8, 2026Adam M. ObermanAnisotropic Loss LandscapesLangevin Dynamics

  12. Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

    Jul 1, 2026Jiaxian Lv, Shiyao Cui, Yingkang Wang +3Content ModerationToxicity

  13. Safe-RULE: Safe Reinforcement UnLEarning

    Jun 8, 2026Shixiong Jiang, Taozheng Zhu, Fanxin KongModel-Based Reinforcement LearningUnsafe

  14. AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

    Jun 3, 2026Yanjing Ren, Reza Ebrahimi, TengTeng MaArtificial Intelligence CompanionsLlm-As-A-Judge

  15. Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise

    May 15, 2026Zhen Huang, Zhihuang Liu, Mengxuan Luo +2Multi-Robot SystemsSafety Alignment

  16. 3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

    May 14, 2026Nicole Meng, Zheyuan Liu, Meng Jiang +13D Editing3D Generation

  17. Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

    May 8, 2026Zhengyang Tang, Yi Zhang, Chenxin Li +18Safety BenchmarksUnsafe

  18. Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy

    May 2, 2026Noushad Sojib, Ola Ghattas, Momotaz BegumImitation LearningBad

  19. Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

    Apr 29, 2026Diego F. Cuadros, Abdoul-Aziz MaigaAgentic DeploymentsEscalation

  20. The Safety-Aware Denoiser for Text Diffusion Models

    Apr 28, 2026Amman Yusuf, Zhejun Jiang, Mijung ParkDiffusion Language ModelsLarge Language Model Safety

  21. From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks

    Apr 25, 2026Nilanjana Das, Mathew Dawit, Aman Chadha +1Jailbreak Success RatesUnsafe

  22. RedVLA: Physical Red Teaming for Vision-Language-Action Models

    Apr 24, 2026Yuhao Zhang, Borong Zhang, Jiaming Fan +4Red-TeamingUnsafe

  23. Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

    Apr 17, 2026Wai Man Si, Mingjie Li, Michael Backes +1Large Language Model SafetyUnsafe

  24. Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation

    Apr 16, 2026Jacob Dang, Brian Y. Xie, Omar G. YounisUnsafeDataset Distillation

  25. Measuring Harmfulness of Computer-Using Agents

    Jul 31, 2025Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3Large Language Model SafetyUnsafe

  26. Learning-based Adaptive Safety-Critical Control With Evolving Unsafe Regions

    Mar 19, 2025Songqiao Hu, Zidong Wang, Zeyi Liu +2Control Barrier FunctionsSafety-Critical Scenarios

  27. SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

    Date pendingHaolong Hu, Hanyu Li, Tiancheng He +6Mllm-BasedLarge Language Model Alignment