cs.CLSep 29, 2026

Selecting The Most Informative Tokens in Natural Language Autoencoders

Authors: Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, Lukas Galke Poech

Organizations: Department of Computer Science, University of Turin, Italy · Department of Mathematics and Computer Science, University of Southern Denmark, Denmark

Abstract

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.74.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5%5\% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

    Sep 27, 2026Gert Lek, Zixuan Xia, Pin-Yu Chen +1Model ActivationsExplainable AI Methods

  2. Measurement Under Selection: Decoy-Calibrated Failure Audits for Language Models

    Jun 8, 2026Vyzantinos Repantis, Ameya Gawde, Harshvardhan SinghLarge Language Models FailModel Auditing