cs.CLSep 27, 2026

Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

Authors: Haiyan Zhao, Zirui Hei, Wei Shi, Huiqi Deng, Na Zou, Mengnan Du

Organizations: New Jersey Institute of Technology · Shanghai AI Laboratory · Xi’an JiaoTong University · Chinese Univerity of Hongkong, Shenzhen

Abstract

Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

    May 25, 2026Haiyan Zhao, Zirui He, Guanchu Wang +3VerbalizationModel Activations

  2. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

    Sep 27, 2026Gert Lek, Zixuan Xia, Pin-Yu Chen +1Model ActivationsExplainable AI Methods