AI Safety Evaluation

Momentum

3 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 91

All topics
CardsList
  1. AI Agents May Always Fall for Prompt Injections

    May 17, 2026Sahar Abdelnabi, Eugene BagdasarianContextual IntegrityAI Agent Security

  2. Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models

    May 17, 2026Nicanor Mayumu, Xiaoheng Deng, Patrick MukalaCoT FaithfulnessVLMs for Autonomous Driving

  3. Muse Spark Safety & Preparedness Report

    May 14, 2026Cristina Menghini, Peter Ney, Hamza Kwisaba +117CybersecurityAI Safety Evaluation

  4. Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

    May 14, 2026Eugene Koran, Yejun Yun, Samantha Tetef +2AI Agent MonitoringAI Control

  5. Training ML Models with Predictable Failures

    May 14, 2026Will Schwarzer, Scott NiekumLanguage Model Safety EvaluationLLM Fine-Tuning

  6. Fusion-fission forecasts when AI will shift to undesirable behavior

    May 14, 2026Neil F. Johnson, Frank Yingjie HuoAI Safety EvaluationLLM Safety

  7. VERA-MH: Validation of Ethical and Responsible AI in Mental Health

    May 13, 2026Luca Belli, Kate H. Bentley, Josh Gieringer +6HealthcareMental Health

  8. RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare

    May 13, 2026Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo +2Distribution Shift RobustnessHealthcare

  9. The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

    May 12, 2026Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1LLM AuditingAI Safety Evaluation

  10. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Opinion DynamicsAI Agent Safety

  11. Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance

    May 10, 2026Michael Alexander Riegler, Finn Schwall, Annika Willoch Olstad +4AI Agent SecurityAI Safety Evaluation

  12. Mental Health AI Safety Claims Must Preserve Temporal Evidence

    May 9, 2026Srimonti Dutta, Ratna KandalaHealthcareMental Health

  13. Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

    May 8, 2026Zhengyang Tang, Yi Zhang, Chenxin Li +18Computer-Use AgentsAgent Evaluation

  14. Automated alignment is harder than you think

    May 7, 2026Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau +1Scalable OversightAI Alignment

  15. Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone

    May 6, 2026Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1LLM EvaluationLLM Alignment

  16. NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

    May 5, 2026Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1AI AccountabilityAI Safety Evaluation

  17. Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

    May 1, 2026Tanav Singh Bajaj, Nikhil Singh, Karan Anand +1Opinion DynamicsAI Agent Safety

  18. Code World Model Preparedness Report

    May 1, 2026Daniel Song, Peter Ney, Cristina Menghini +21AI Safety EvaluationLLM Safety Evaluation

  19. Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty

    Apr 29, 2026Akshay Karjol, Shadi AlawnehCollision AvoidanceUncertainty Quantification

  20. Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

    Apr 29, 2026Mahiro Nakao, Kazuhiro TakemotoLanguage Model-Based ControlHealthcare

  21. A Comparative Evaluation of AI Agent Security Guardrails

    Apr 27, 2026Qi Li, Jiu Li, Pingtao Wei +8LLM GuardrailsAI Agent Security

  22. Ethics Testing: Proactive Identification of Generative AI System Harms

    Apr 23, 2026Shin Hwei Tan, Haibo Wang, Heng LiAI Safety EvaluationAutomated Test Generation

  23. How VLAs (Really) Work In Open-World Environments

    Apr 23, 2026Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4VLM RobustnessLong-Horizon Robotic Manipulation

  24. When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains

    Apr 21, 2026Ishita Kakkar, Enze Zhang, Rheeya Uppaal +1AI Safety EvaluationLLM Safety Evaluation

  25. HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

    Apr 14, 2026Zixing Chen, Yifeng Gao, Li Wang +8RoboticsAI Safety Evaluation

  26. SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

    Mar 18, 2026Rima Hazra, Bikram Ghuku, Ilona Marchenko +5Intelligent Tutoring SystemsAI Safety Evaluation