Value Alignment

Momentum

0 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 19

All topics
CardsList
  1. Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

    Sep 2, 2026Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6Value Alignment

  2. Contextual Value Alignment via Multilayer Combinatorial Fusion

    Aug 7, 2026Yuanhong Wu, Djallel Bouneffouf, D. Frank HsuLLM AlignmentMoral Reasoning in Language Models

  3. D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

    Jul 22, 2026Siyi Hao, Yidi Cao, Linhao Yu +2LLM EvaluationLLM Alignment

  4. Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila

    Jul 20, 2026Supryadi, Irfan, Julianti +4LLM EvaluationLLM Alignment

  5. Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

    Jul 12, 2026Asher Sprigler, Yang-Yang Feng, Iftach Amir +5LLM AlignmentMoral Reasoning in Language Models

  6. Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives

    Jun 26, 2026Young Yoon, Jimin Kim, Soyeon ParkGradient-Based AttributionValue Alignment

  7. Reinforcement Learning Towards Broadly and Persistently Beneficial Models

    Jun 22, 2026Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab +5LLM AlignmentLLM Safety Alignment

  8. Position: Align AI to Our Aspirations, Not Our Flaws

    Jun 11, 2026Nikita Kazeev, Bui Nhat Huyen PhanAI AlignmentPluralistic Alignment

  9. Sycophancy Towards Researchers Drives Performative Misalignment

    Jun 7, 2026David D. Baek, Xinnuo Li, Anay Gupta +4LLM SycophancyLLM Alignment

  10. Building Comparative Motivation Profiles with Instrumental Interventions

    Jun 6, 2026David Vella Zarb, Rustem Turtayev, Taywon Min +2LLM Safety EvaluationCausal Interventions in Language Models

  11. Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

    Jun 3, 2026Ajay Vishwanath, Christian OmlinReinforcement LearningCooperative MARL

  12. Behavioural Analysis of Alignment Faking

    May 26, 2026Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1LLM AlignmentValue Alignment

  13. DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

    May 14, 2026Pengyun Zhu, Yuqi Ren, Zhen Wang +2LLM AlignmentPluralistic Alignment

  14. Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance

    May 12, 2026Wenhao Chen, Sirui Sun, Shengyuan Bai +1LLM AlignmentLLM Safety Alignment

  15. Positive Alignment: Artificial Intelligence for Human Flourishing

    May 11, 2026Ruben Laukkonen, Seb Krier, Chloé Bakalar +13LLM AlignmentAI Alignment

  16. Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

    Apr 29, 2026Matteo Leonesi, Francesco Belardinelli, Flavio Corradini +1Deception in Language ModelsLLM Safety Alignment

  17. Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

    Date pendingKevin Baum, Rūta Binkytė, Felix JahnAI AlignmentLLM Safety Alignment