AI Alignment

Momentum

4 papers in the last four weeks, down 43% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 83

All topics
CardsList
  1. No More, No Less: Task Alignment in Terminal Agents

    May 12, 2026Sina Mavali, David Pape, Jonathan Evertz +5Computer-Use Agent BenchmarksAI Alignment

  2. Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

    May 11, 2026Haonan Dong, Qiguan Feng, Kehan Jiang +3AI AlignmentAI Agent Evaluation

  3. Positive Alignment: Artificial Intelligence for Human Flourishing

    May 11, 2026Ruben Laukkonen, Seb Krier, Chloé Bakalar +13LLM AlignmentAI Alignment

  4. Automated alignment is harder than you think

    May 7, 2026Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau +1Scalable OversightAI Alignment

  5. Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

    May 5, 2026Jie Zhou, Qin Chen, Liang HeAI AlignmentAI Safety

  6. AI Alignment via Incentives and Correction

    May 2, 2026Rohit Agarwal, Joshua Lin, Mark Braverman +1AI Alignment

  7. MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

    May 2, 2026Jiyao Wang, Yunbiao Wang, Yubo Jiao +6AI AlignmentAI Agent Safety

  8. A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societies

    Apr 24, 2026Somyajit ChakrabortyAI AlignmentAI Governance

  9. Alignment has a Fantasia Problem

    Apr 23, 2026Nathanael Jo, Zoe De Simone, Mitchell Gordon +1AI AlignmentHuman-AI Interaction

  10. The Triadic Loop: A Framework for Negotiating Alignment in AI Co-hosted Livestreaming

    Apr 20, 2026Katherine Wang, Nadia Berthouze, Aneesha SinghAI AlignmentPluralistic Alignment

  11. The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem

    Apr 16, 2026Till Mossakowski, Helena Esther GrassAI AlignmentAI Agent Governance

  12. Evaluating Alignment of Behavioral Dispositions in LLMs

    Feb 11, 2026Amir Taubenfeld, Zorik Gekhman, Lior Nezry +8Human Preference EvaluationLLM Alignment

  13. Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs

    Jan 15, 2026Nan Li, Bo Kang, Tijl De BieMultilingual Language Model EvaluationLLM Alignment

  14. Labels have Human Values: Value Calibration of Subjective Tasks

    Jan 10, 2026Mohammed Fayiz Parappan, Ricardo HenaoAnnotator DisagreementAI Alignment

  15. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    Jun 4, 2025Akshat Naik, Emma Gouné, Patrick Quinn +4AI AlignmentLLM Agent Evaluation

  16. Societal Alignment Frameworks Can Improve LLM Alignment

    Feb 27, 2025Karolina Stańczak, Nicholas Meade, Mehar Bhatia +14LLM AlignmentAI Alignment

  17. Modelling Human Values for Value-Aware Multi-Agent Systems

    Feb 9, 2024Nardine Osman, Mark d'InvernoAI AlignmentMulti-Agent Systems

  18. Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

    Date pendingKevin Baum, Rūta Binkytė, Felix JahnAI AlignmentLLM Safety Alignment