cs.AIJun 29, 2026

Open Problems in Constitutional Preference Reconstruction

Authors: Eleanor CliffordMichael AmirArduin FindeisAaron ZhaoRobert Mullins

Organizations: Imperial College London · University of Cambridge

Abstract

Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods such as Inverse Constitutional AI (ICAI) attempt to improve interpretability by compressing datasets into short ``constitutions'' of natural-language principles. We argue this framing is under-specified: a flat list of principles is not yet an executable decision rule because it leaves principle composition implicit. We use the pairwise setting as a testbed to empirically characterize three open problems in constitutional methods. First, principle quality is hard to measure: coverage and accuracy are useful but incomplete proxies for end-to-end reconstruction. Second, \emph{composition is ambiguous}: holding principles fixed, different executors (LLM judge versus majority vote) agree only 73%73\% of the time. Third, \emph{constitutions differ between LLMs}: cross-model vote agreement is 73%73\%, whereas intra-model agreement is 81%81\%. Across PRISM, AlpacaEval, and Chatbot Arena, we show that principle refinement (ICAI+) may be a first step towards ameliorating these problems: inter-executor agreement rises to 78%78\%, and transparent executors match LLM judge accuracy (66%66\% vs.\ 67%67\%). Our results highlight that constitutions should be evaluated as \emph{constitution--executor systems}, with implications for LLMs-as-a-judge broadly.

Explore similar work

CardsList
  1. Democratic ICAI: Debating Our Way to Steering Principles from Preferences

    Jun 26, 2026Kevin Kingslin, Anish Natekar, Ashutosh Ranjan +3DemocracyDeliberation

  2. The Constitutional Coverage Trilemma in AI Governance

    Sep 1, 2026Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly +1Artificial Intelligence GovernanceAutonomy