AI Alignment

Momentum

4 papers in the last four weeks, down 43% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 83

All topics
CardsList
  1. AI Alignment From Social Choice Perspectives

    Jun 19, 2026Daniel Halpern, Evi Micha, Ariel D. Procaccia +3LLM AlignmentAI Alignment

  2. Position: Align AI to Our Aspirations, Not Our Flaws

    Jun 11, 2026Nikita Kazeev, Bui Nhat Huyen PhanAI AlignmentPluralistic Alignment

  3. Exploring the Design Space of Reward Backpropagation for Flow Matching

    Jun 9, 2026Ruoyu Wang, Boye Niu, Xiangxin Zhou +3Flow MatchingAI Alignment

  4. Position: AI Must Become Planet-Centered, Not Just Human-Centered

    Jun 9, 2026Maria Perez-OrtizAI AlignmentHuman-Centered AI

  5. The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment

    Jun 9, 2026Filippo Tonini, Federico Torrielli, Anton Danholt Lautrup +3AI Agent AuditingLLM Auditing

  6. Emergent alignment and the projectability of ethical personas

    Jun 8, 2026Guillermo Del Pinal, Youngchan Lee, Calum McNamara +1Moral Reasoning in Language ModelsLanguage Model Safety Evaluation

  7. Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

    Jun 4, 2026Salvatore Greco, Hainiu Xu, Jacopo Domenicucci +2Moral Reasoning in Language ModelsAI Alignment

  8. Agentic Safety is an Epistemic Property, Not a Behavioral One

    Jun 2, 2026Charles L. Wang, Keir Dorchen, Peter JinAI AlignmentAI Agent Safety

  9. Structuring the Space of Sociotechnical Alignment

    Jun 2, 2026Esra Dönmez, Agnieszka FalenskaAI AlignmentHuman-Centered AI

  10. Solipsistic Superintelligence is Unlikely to be Cooperative

    Jun 2, 2026Rakshit S Trivedi, Natasha Jaques, Logan Cross +2AI AlignmentMulti-Agent Collaboration

  11. Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence

    Jun 2, 2026Karthik Barma, Anil Sanneboyina, V C Premchand YadavAI AlignmentAI Agent Safety

  12. Rationalize: Shared Semantic Reasoning for Human-AI Alignment

    May 28, 2026Aritra Dasgupta, Naga Datha Saikiran Battula, Avina Nakarmi +3AI AlignmentHuman-AI Interaction

  13. Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment

    May 28, 2026Toru TakahashiComputational Cognitive ModelingWorld Model Learning

  14. You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

    May 26, 2026Nicole Hsing, Asuka Yuxi Zheng, Yi Zhao +2Social DilemmasAI Alignment

  15. A Sober Look at Agentic Misalignment in Automated Workflows

    May 22, 2026Wenqian Ye, Bo Yuan, Zhichao Xu +4AI AlignmentAgentic Workflows

  16. Investigating Concept Alignment Using Implausible Category Members

    May 20, 2026Sunayana Rane, Brenden M. Lake, Thomas L. GriffithsAI Alignment

  17. Embedding-perturbed Exploration Preference Optimization for Flow Models

    May 15, 2026Sujie Hu, Chubin Chen, Jiashu Zhu +3AI AlignmentGroup Relative Policy Optimization

  18. Position: Assistive Agents Need Accessibility Alignment

    May 13, 2026Jie Hu, Changyuan Yan, Yu Zheng +2AI AlignmentHuman-Centered AI