cs.LGJan 20, 2026

Self-Improvement as Coherence Optimization: A Theoretical Account

Authors: Tianyi QiuAhmed Hani IsmailZhonghao HeShi Feng

Abstract

Can language models improve their accuracy without external supervision? Methods such as debate, bootstrap, and internal coherence maximization achieve this surprising feat, even matching golden finetuning performance. Yet why they work remains theoretically unclear. We show that they can all be understood as coherence optimization, the search for a context-to-behavior mapping that is most compressible and jointly predictable, with debate an exact instance and bootstrap and internal coherence maximization closely related to it. We prove that coherence optimization is equivalent to description-length regularization, and that among all such regularization schemes, coherence regularization with a prior derived from a pretrained model optimizes a lower bound of worst-case accuracy for semi-supervised learning. Our theory, supported by preliminary experiments, explains why feedback-free self-improvement works and predicts when it should succeed or fail.

Explore similar work

CardsList
  1. Position: It's Time to Optimize LLMs for Self-Consistency

    Jul 31, 2026Itamar Pres, Belinda Z. Li, Laura Ruis +6Self-ConsistencyPosition

  2. On the Generalization Gap in Self-Evolving Language Model Reasoning

    May 31, 2026Zhenting Qi, Susanna Maria Baby, Stefanie Anna Baby +5Self-EvolutionReasoning Benchmark