AI Alignment

Momentum

4 papers in the last four weeks, down 43% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 83

All topics
CardsList
  1. AI Safety Considerations for Agents With Limited Time to Act

    Oct 7, 2026Leo Zeitler, Jack Richings, Victoria NocklesAI Agent SafetyAI Safety

  2. How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk

    Oct 6, 2026Mikołaj Sienicki, Krzysztof SienickiAI SafetyAI Risk Management

  3. SPEAR: Five Principles for Interactive Human-Agent Alignment

    Oct 5, 2026Tao Long, Lydia B. ChiltonAI AlignmentHuman-AI Interaction

  4. Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

    Sep 12, 2026Marica Notte, Ludovica Marinucci, Vieri Giuliano SantucciAI AlignmentAI Agent Governance

  5. Mechanism Design for Alignment and Control

    Sep 1, 2026Dirk Bergemann, Andrew Koh, Stephen MorrisMechanism DesignAI Alignment

  6. The Constitutional Coverage Trilemma in AI Governance

    Sep 1, 2026Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly +1AI AlignmentAI Governance

  7. Automated Researchers Can Mitigate Well-characterized Alignment Failures

    Aug 28, 2026Chen Yueh-Han, Jiaxin Wen, Jan Hendrik KirchnerAI AlignmentLLM Safety Alignment

  8. AI Alignment through a Game-theoretic Lens: A Survey

    Aug 28, 2026Yanan Cai, Zhongrui Zhao, Zhigang Lu +6LLM AlignmentAI Alignment

  9. Rules or Character? Scaling Laws for AI Safety Design

    Aug 13, 2026Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu +1AI AlignmentAI Safety

  10. Toward a Theory of Value in AI Alignment

    Aug 10, 2026Andrew Smart, Shazeda Ahmed, Jackie Kay +3LLM AlignmentAI Alignment

  11. A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

    Aug 8, 2026Jobst Heitzig, Ram PothamAI AlignmentAlgorithmic Fairness

  12. Metanormative Theory for RL-Based Moral Agents

    Aug 8, 2026Aleks Knoks, Marija SlavkovikAI AlignmentMoral Reasoning

  13. AI Alignment and Fiduciary Obligation

    Aug 1, 2026Benjamin LangeAI AlignmentHuman-AI Interaction

  14. Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

    Jul 31, 2026Tyler Ashoff, Jordan RoduAI AlignmentCross-Modal Alignment

  15. Interactive Alignment

    Jul 27, 2026Sylvain ChassangAI AlignmentAI Agent Governance

  16. Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

    Jul 24, 2026Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao +1Creativity AssessmentLLM Evaluation

  17. Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

    Jul 20, 2026Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins +3Explainable Artificial IntelligenceAI Alignment

  18. Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development

    Jul 16, 2026Jean-Paul Van BelleAI AlignmentMoral Reasoning

  19. Align AI to Dynamic Human-AI Workflows

    Jul 15, 2026Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley +4AI AlignmentHuman-AI Interaction

  20. EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

    Jun 29, 2026Buğra Alperen Uluırmak, Rifat KurbanAI AlignmentLLM Safety Alignment

  21. Safety from Honesty in a Disinterested AI Predictor

    Jun 28, 2026Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13AI AlignmentAI Safety

  22. Agent Safety Is Action Alignment

    Jun 27, 2026Shawn Li, Yue ZhaoLanguage Model Safety EvaluationAI Alignment

  23. LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

    Jun 26, 2026Qinhong Zhou, Chuang Gan, Anoop CherianMulti-Agent CoordinationAI Alignment