Artificial Intelligence Alignment

Recent momentum

-33%

10 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

2 new papers

A weekly snapshot of new work published in Artificial Intelligence Alignment.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Artificial Intelligence Alignment.

Period ending 2026-09-07

3 new papers

A weekly snapshot of new work published in Artificial Intelligence Alignment.

97 papers

Latest in Artificial Intelligence Alignment

  1. How Well Do Models Follow Their Constitutions?

    May 22, 2026Arya Jakkli, Senthooran Rajamanoharan, Neel NandaArtificial Intelligence AlignmentAdversaries

  2. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Artificial Intelligence AlignmentConformity

  3. AI Alignment via Incentives and Correction

    May 2, 2026Rohit Agarwal, Joshua Lin, Mark Braverman +1Artificial Intelligence AlignmentIncentives

  4. AI Safety Training Can be Clinically Harmful

    Apr 25, 2026Suhas BN, Andrew M. Sherrill, Rosa I. Arriaga +2Mental Health SupportLarge Language Model Safety

  5. Alignment has a Fantasia Problem

    Apr 23, 2026Nathanael Jo, Zoe De Simone, Mitchell Gordon +1Human-Ai InteractionArtificial Intelligence Alignment

  6. Peer-Preservation in Frontier Models

    Mar 30, 2026Yujin Potter, Nicholas Crispino, Vincent Siu +2Frontier ModelsArtificial Intelligence Alignment