cs.AIAug 28, 2026

AI Alignment through a Game-theoretic Lens: A Survey

Authors: Yanan CaiZhongrui ZhaoZhigang LuIckjai LeeWei Emma ZhangMinhui XueYihong ZhangShuchao Pang+1 more

Organizations: James Cook University · Western Sydney University · Adelaide University · CSIRO · The University of Osaka · Nanjing University of Science and Technology · Macquarie University · La Trobe University

Abstract

As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.

Explore similar work

CardsList
  1. Conformity Generates Collective Misalignment in AI Agents Societies

    May 11, 2026Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2Artificial Intelligence AlignmentConformity