cs.CLSep 27, 2026

Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

Authors: Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen

Organizations: Universit´e de Neuchˆatel · Universität Bern · IBM Research · Delft University of Technology

Abstract

Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Selecting The Most Informative Tokens in Natural Language Autoencoders

    Sep 29, 2026Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez +4Explainable AI MethodsModel Activations

  2. Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

    May 25, 2026Haiyan Zhao, Zirui He, Guanchu Wang +3VerbalizationModel Activations