Selecting The Most Informative Tokens in Natural Language Autoencoders
Organizations: Department of Computer Science, University of Turin, Italy · Department of Mathematics and Computer Science, University of Southern Denmark, Denmark
Abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Figures & tables
| Signal | Definition | Elevated at positions where |
| predictive distribution | ||
| surprisal | the observed token was improbable | |
| entropy | the continuation was uncertain | |
| varentropy | the uncertainty is concentrated on few alternatives | |
| temporal_kl | the observed token changed the prediction | |
| attention pattern | ||
| OpenPromptInjection | Tensor Trust | Liars’ Bench | Taboo organisms | |||||||||||
| signal | q7 | g12 | g27 | l70 | q7 | g12 | g27 | l70 | g27 | l70 | q7 | g12 | g27 | l70 |
| predictive distribution | ||||||||||||||
| surprisal | 0.511 | 0.506 | 0.495 | 0.513 | 0.522 | 0.629 | 0.599 | 0.612 | 0.599 | 0.548 | 0.522 | 0.504 ∘ | 0.571 | 0.491 ∘ |
| entropy | 0.372 | 0.397 | 0.452 | 0.451 | 0.510 ∘ | 0.598 | 0.589 | 0.611 | 0.700 | 0.609 | 0.566 | 0.557 | 0.604 | 0.522 |
| varentropy | 0.366 | 0.391 | 0.442 | 0.424 | 0.498 ∘ | 0.583 | 0.582 | 0.583 | 0.712 | 0.596 | 0.551 | 0.601 | 0.622 | 0.535 |
| temporal_kl | 0.606 | 0.508 | 0.509 | 0.482 | 0.627 | 0.648 | 0.601 | 0.650 | 0.396 | 0.375 | 0.499 ∘ | 0.500 ∘ | 0.402 | 0.515 ∘ |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Model | Positions | Base rate | Lift, all positions | Lift, within role |
| Open PromptInjection | q7 | ||||
| g12 | |||||
| g27 | |||||
| l70 | |||||
| Taboo organisms | q7 | n/a | n/a | ||
| g12 | n/a | n/a |
| Strategy | Qwen2.5-7B | Gemma-3-12B | Gemma-3-27B | Llama-3.3-70B |
| naive | ||||
| escape | ||||
| fake_comp | ||||
| ignore | ||||
| combine |
| All positions | Shared prompt | Own word, by concealed word | |||||||
| Model | Positions | On-task | Own word | Other | Own | Other | moon | ship | snow |
| q7 | |||||||||
| g12 | |||||||||
| g27 | |||||||||
| l70 | |||||||||
| OpenPromptInjection | Tensor Trust | Liars’ Bench | Taboo organisms | |||||||||||||
| signal | q7 | g12 | g27 | l70 | q7 | g12 | g27 | l70 | g27 | l70 | q7 | g12 | g27 | l70 | Agreeing | Minority |
| predictive distribution | ||||||||||||||||
| surprisal | H | H | L | H | H | H | H | H | H | H | H | H | H | L | 2/4 | 2 |
| entropy | L | L | L | L | H | H | H | H | H | H | H | H | H | H | 4/4 | 4 |
| varentropy | L | L | L | L | L | H | H | H | H | H | H | H | H | H | 3/4 | 5 |
| temporal_kl | H | H | H | L | H | H | H | H | L | L | L | H | L | H | 2/4 | 5 |
| signal | Correlation | Pooled | Within structure |
| predictive distribution | |||
| surprisal | to | ||
| entropy | to | ||
| varentropy | to | ||
| temporal_kl | to | ||
| attention pattern | |||
| Dataset | Model | Signal | Direction | Pooled [95% CI] | Case-macro | Within structure | Free structure |
| OpenPromptInjection | q7 | peak_ratio | lower | ||||
| g12 | resid_jump_nla | higher | |||||
| g27 | resid_jump_nla | higher | |||||
| l70 | sink_drain | higher | |||||
| Tensor Trust | q7 | resid_jump | higher | ||||
| g12 | norm_ratio | higher |
| Signal | Ensemble | Structure | ||||||
| Dataset | Model | Base | ||||||
| OpenPromptInjection | q7 | |||||||
| g12 | ||||||||
| g27 | ||||||||
| l70 | ||||||||
| Taboo organisms | q7 | |||||||
| Dataset | Model | Signal 1 | Signal 2 | Held-out | CI |
| OpenPromptInjection | q7 | lookback_ratio | sink_drain | ||
| g12 | lookback_ratio | sink_drain | |||
| g27 | lookback_ratio | sink_drain | |||
| l70 | peak_ratio | sink_drain | |||
| Tensor Trust | q7 | resid_jump | w | ||
| g12 | norm_ratio | resid_jump |
| Complete-data pooled | Held-out pooled | Held-out case-macro | ||||||
| Dataset | Model | Ensemble | Ens. | Single | Ens. | Single | Ens. | Single |
| Open Prompt- Injection | q7 | lookback_ratio / sink_drain | ||||||
| g12 | lookback_ratio / sink_drain | |||||||
| g27 | lookback_ratio / sink_drain | |||||||
| l70 | peak_ratio / sink_drain | |||||||
| Tensor Trust | q7 | resid_jump / w | ||||||
| Dataset | Shared ensemble | Full-data pooled | Held-out pooled | Held-out case-macro |
| OpenPromptInjection | lookback_ratio + sink_drain | |||
| Tensor Trust | resid_jump_nla + w | |||
| Liars’ Bench | dominant_mass + head_disagreement | |||
| Taboo organisms | dominant_mass + norm_ratio |