AI Alignment

Momentum

4 papers in the last four weeks, down 43% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 83

All topics
CardsList
  1. AI Safety Considerations for Agents With Limited Time to Act

    Oct 7, 2026Leo Zeitler, Jack Richings, Victoria NocklesAI Agent SafetyAI Safety

  2. How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk

    Oct 6, 2026Mikołaj Sienicki, Krzysztof SienickiAI SafetyAI Risk Management

  3. SPEAR: Five Principles for Interactive Human-Agent Alignment

    Oct 5, 2026Tao Long, Lydia B. ChiltonAI AlignmentHuman-AI Interaction

  4. Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

    Sep 12, 2026Marica Notte, Ludovica Marinucci, Vieri Giuliano SantucciAI AlignmentAI Agent Governance

  5. Mechanism Design for Alignment and Control

    Sep 1, 2026Dirk Bergemann, Andrew Koh, Stephen MorrisMechanism DesignAI Alignment

  6. The Constitutional Coverage Trilemma in AI Governance

    Sep 1, 2026Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly +1AI AlignmentAI Governance

  7. Automated Researchers Can Mitigate Well-characterized Alignment Failures

    Aug 28, 2026Chen Yueh-Han, Jiaxin Wen, Jan Hendrik KirchnerAI AlignmentLLM Safety Alignment

  8. AI Alignment through a Game-theoretic Lens: A Survey

    Aug 28, 2026Yanan Cai, Zhongrui Zhao, Zhigang Lu +6LLM AlignmentAI Alignment

  9. Rules or Character? Scaling Laws for AI Safety Design

    Aug 13, 2026Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu +1AI AlignmentAI Safety

  10. Toward a Theory of Value in AI Alignment

    Aug 10, 2026Andrew Smart, Shazeda Ahmed, Jackie Kay +3LLM AlignmentAI Alignment

  11. A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

    Aug 8, 2026Jobst Heitzig, Ram PothamAI AlignmentAlgorithmic Fairness

  12. Metanormative Theory for RL-Based Moral Agents

    Aug 8, 2026Aleks Knoks, Marija SlavkovikAI AlignmentMoral Reasoning

  13. AI Alignment and Fiduciary Obligation

    Aug 1, 2026Benjamin LangeAI AlignmentHuman-AI Interaction

  14. Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

    Jul 31, 2026Tyler Ashoff, Jordan RoduAI AlignmentCross-Modal Alignment

  15. Interactive Alignment

    Jul 27, 2026Sylvain ChassangAI AlignmentAI Agent Governance

  16. Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

    Jul 24, 2026Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao +1Creativity AssessmentLLM Evaluation

  17. Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

    Jul 20, 2026Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins +3Explainable Artificial IntelligenceAI Alignment

  18. Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development

    Jul 16, 2026Jean-Paul Van BelleAI AlignmentMoral Reasoning

  19. Align AI to Dynamic Human-AI Workflows

    Jul 15, 2026Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley +4AI AlignmentHuman-AI Interaction

  20. EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

    Jun 29, 2026Buğra Alperen Uluırmak, Rifat KurbanAI AlignmentLLM Safety Alignment

  21. Safety from Honesty in a Disinterested AI Predictor

    Jun 28, 2026Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13AI AlignmentAI Safety

  22. Agent Safety Is Action Alignment

    Jun 27, 2026Shawn Li, Yue ZhaoLanguage Model Safety EvaluationAI Alignment

  23. LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

    Jun 26, 2026Qinhong Zhou, Chuang Gan, Anoop CherianMulti-Agent CoordinationAI Alignment

  24. AI Alignment From Social Choice Perspectives

    Jun 19, 2026Daniel Halpern, Evi Micha, Ariel D. Procaccia +3LLM AlignmentAI Alignment

  25. Position: Align AI to Our Aspirations, Not Our Flaws

    Jun 11, 2026Nikita Kazeev, Bui Nhat Huyen PhanAI AlignmentPluralistic Alignment

  26. Exploring the Design Space of Reward Backpropagation for Flow Matching

    Jun 9, 2026Ruoyu Wang, Boye Niu, Xiangxin Zhou +3Flow MatchingAI Alignment

  27. Position: AI Must Become Planet-Centered, Not Just Human-Centered

    Jun 9, 2026Maria Perez-OrtizAI AlignmentHuman-Centered AI

  28. The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment

    Jun 9, 2026Filippo Tonini, Federico Torrielli, Anton Danholt Lautrup +3AI Agent AuditingLLM Auditing

  29. Emergent alignment and the projectability of ethical personas

    Jun 8, 2026Guillermo Del Pinal, Youngchan Lee, Calum McNamara +1Moral Reasoning in Language ModelsLanguage Model Safety Evaluation

  30. Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

    Jun 4, 2026Salvatore Greco, Hainiu Xu, Jacopo Domenicucci +2Moral Reasoning in Language ModelsAI Alignment

  31. Agentic Safety is an Epistemic Property, Not a Behavioral One

    Jun 2, 2026Charles L. Wang, Keir Dorchen, Peter JinAI AlignmentAI Agent Safety

  32. Structuring the Space of Sociotechnical Alignment

    Jun 2, 2026Esra Dönmez, Agnieszka FalenskaAI AlignmentHuman-Centered AI

  33. Solipsistic Superintelligence is Unlikely to be Cooperative

    Jun 2, 2026Rakshit S Trivedi, Natasha Jaques, Logan Cross +2AI AlignmentMulti-Agent Collaboration

  34. Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence

    Jun 2, 2026Karthik Barma, Anil Sanneboyina, V C Premchand YadavAI AlignmentAI Agent Safety

  35. Rationalize: Shared Semantic Reasoning for Human-AI Alignment

    May 28, 2026Aritra Dasgupta, Naga Datha Saikiran Battula, Avina Nakarmi +3AI AlignmentHuman-AI Interaction

  36. Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment

    May 28, 2026Toru TakahashiComputational Cognitive ModelingWorld Model Learning

  37. You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

    May 26, 2026Nicole Hsing, Asuka Yuxi Zheng, Yi Zhao +2Social DilemmasAI Alignment

  38. A Sober Look at Agentic Misalignment in Automated Workflows

    May 22, 2026Wenqian Ye, Bo Yuan, Zhichao Xu +4AI AlignmentAgentic Workflows

  39. Investigating Concept Alignment Using Implausible Category Members

    May 20, 2026Sunayana Rane, Brenden M. Lake, Thomas L. GriffithsAI Alignment

  40. Embedding-perturbed Exploration Preference Optimization for Flow Models

    May 15, 2026Sujie Hu, Chubin Chen, Jiashu Zhu +3AI AlignmentGroup Relative Policy Optimization

  41. Position: Assistive Agents Need Accessibility Alignment

    May 13, 2026Jie Hu, Changyuan Yan, Yu Zheng +2AI AlignmentHuman-Centered AI