AI Safety

Momentum

10 papers in the last four weeks, up 150% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 88

All topics
CardsList
  1. AI Safety Considerations for Agents With Limited Time to Act

    Oct 7, 2026Leo Zeitler, Jack Richings, Victoria NocklesAI Agent SafetyAI Safety

  2. Comprehension Audits to Mitigate Risks from Automated AI Research

    Oct 7, 2026Ronald J. Bodkin, Bahrad A. Sokhansanj, Gillian K. HadfieldHuman Oversight of AIAI Safety

  3. Systemization of Knowledge (SoK): Human-Centered AI Safety for Youth

    Oct 6, 2026Pratyasha Saha, Yaman Yu, Yang WangAI Risk ManagementAI Safety Evaluation

  4. How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk

    Oct 6, 2026Mikołaj Sienicki, Krzysztof SienickiAI SafetyAI Risk Management

  5. Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI

    Oct 5, 2026Muhammad Rafiullah Memon, Viet Vo, Wanlun Ma +1Sign Language TranslationAI Safety Evaluation

  6. AI Safety via Debate is Compromised by Cognitive Biases

    Oct 4, 2026Gefei Liu, Sonya Rashkovan, Sophia Lloyd George +2Language Model Safety EvaluationAI Safety

  7. What if automating AI R&D triggers an intelligence explosion?

    Sep 28, 2026Alan Chan, Christoph Winter, Andrew Barto +19AI ControlAI Governance

  8. Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

    Sep 22, 2026Varshini Elangovan, James Wedgwood, Chhavi Yadav +3AI Agent SafetyHuman-AI Interaction

  9. AI Safety: Not Optional, Not Later

    Sep 14, 2026Qinghua Lu, Yoshua BengioAI AssuranceAI Accountability

  10. A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots

    Sep 13, 2026Mohammad Namvarpour, Tyler Chang, Afsaneh RaziAI CompanionsHuman-AI Interaction

  11. AI Persuasion as a Threat to Human Control

    Sep 13, 2026Joshua Levy, Mick Yang, Kellin PelrineAI ControlAI Governance

  12. LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

    Sep 9, 2026Hanjing Zhou, Mingze Yin, Ying Lian +3VLM EvaluationVisual Question Answering

  13. AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

    Aug 31, 2026Paulius Mui, Dean F. Sittig, Steve Labkoff +1AI Risk ManagementAI Safety

  14. Rules or Character? Scaling Laws for AI Safety Design

    Aug 13, 2026Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu +1AI AlignmentAI Safety

  15. SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

    Aug 10, 2026Yuanchi Zhu, Kang An, Tengyue Wang +11Multimodal ReasoningAI Safety

  16. A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

    Aug 8, 2026Jobst Heitzig, Ram PothamAI AlignmentAlgorithmic Fairness

  17. Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

    Aug 6, 2026Jessica Y. Bo, Paula Akemi Aoyagui, Shalaleh Rismani +3AI Safety EvaluationHuman-Centered AI

  18. Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

    Jul 28, 2026Elias Fernández Domingos, The Anh HanAI Safety

  19. On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

    Jul 25, 2026Muhammad Tukur, Hayatullahi B. Adeyemo, Tao Chen +5Software EngineeringAI Risk Management

  20. Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards

    Jul 22, 2026Gabriele Franchini, Giulio Mallardi, Michele De Carolis +1Cyber-Physical SystemsAI Agent Monitoring

  21. Hardware Mechanisms to Dynamically Throttle AI Performance

    Jul 20, 2026Haiyue Ma, Lauren Malek, Joseph Forzani +1AI ControlAI Safety