Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.
Figures & tables
Figure 1: Overview of AVPO. Stars mark the order of the SFT, DPO, and QA steps, snowflakes and flames mark frozen and trainable modules, and r is the rollout reward.
Method
Best Layer
Decoder
Gist ↑
Detail ↑
Overall ↑
Ans Rate ↑
Within ↓
Cross ↓
Len
Donor: Llama-3.1-8B
Patchscopes
L21
self
0.018
0.036
0.027
0.989
—
—
—
SelfIE
L27
self
0.108
0.132
0.120
0.979
—
—
—
AO
L27
self
0.380
0.465
0.423
0.999
—
—
—
LatentQA
L27
self
0.422
0.527
0.475
1.000
—
—
—
UAV
L31
self
0.454
0.531
0.492
1.000
—
—
—
Table 1: Main comparison with activation verbalization baselines, grouped by donor model. AVPO SFT denotes our inverter before DPO, and AVPO r k denotes it after k rounds of DPO. Each method is reported at its validation-selected layer and DPO round, and bold marks the selected checkpoint of each configuration. self denotes self-decoding by the donor model. Within is the fraction of outputs with at least 50% repeated 4-grams, Cross is the fraction of outputs identical to those of other inputs, and Len is the mean output length in words. Baselines without reconstructions report no rollout metrics (—). † Evaluated only at layer 31.
Model / Reward
Reward formulation
Gist ↑
Detail ↑
Anchor F1 ↑
Within ↓
Cross ↓
SFT
—
0.254
0.335
0.232
0.008
0.004
Gist only
G
0.363
0.422
0.201
0.031
0.003
Detail only
D
0.324
0.405
0.219
0.015
0.001
Anchor F1 only
A
0.249
0.338
0.242
0.012
0.005
Gist + Detail
0.5G+0.5D
0.341
0.417
0.211
0.048
0.008
Gist + Anchor F1
0.7G+0.3A
0.344
0.400
0.228
0.013
0.004
Table 2: Reward ablation in the Qwen3-4B self-explanation setting at layer 27, each trained with a single round of DPO. Results are macro-averaged across six source families on the same test set.
Figure 2: Layer-wise inversion performance across donor-model depth. We evaluate gist and detail recovery from Llama-3.1-8B-Instruct (a–b) and Mistral-Small-24B-Instruct-2501 (c–d) activations using either a self-model inverter or a cross-model Qwen3-4B inverter, before (SFT) and after one round of DPO. Scores are six-family macro averages on the test set.
Figure 3: Effect of iterative preference optimization on reconstruction quality and generation collapse. We report six-family macro-averaged (a) gist score, (b) detail score, (c) cross-output collapse, and (d) within-output collapse across DPO rounds, for the Qwen3-4B and Llama-3.1-8B inverters on layer-31 Llama-3.1-8B-Instruct activations.
Harmful requests
Clinical summaries
Method
Intent ↑
Category ↑
Name acc. ↑
Name fab. ↓
Age ± 2 ↑
Sex acc. ↑
Complaint ↑
Finding ↑
UAV
0.66
0.59
0.08
0.87
0.32
1.00
0.46
0.04
AO
0.54
0.63
0.01
0.99
0.33
0.99
0.52
0.01
LatentQA
0.57
0.57
0.09
0.90
0.39
1.00
0.50
0.02
AVPO SFT
0.64
0.77
0.00
0.06
0.49
0.96
0.38
0.11
AVPO r4
0.77
0.78
0.03
0.10
0.27
0.70
0.51
0.12
Table 3: Case study with self-decoding on Llama-3.1-8B-Instruct activations. AVPO SFT and AVPO r4 denote our inverter before DPO and after four rounds of DPO. Intent, Category, Complaint, and Finding are LLM-judge scores, and Finding is computed only on the 187 summaries that contain a key finding. Name, Age, and Sex are rule-based, with Age counted as correct within two years. Floor uses no context and Ceiling uses the original text. Bold marks the best method per column.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Ours
UAV
AO
LatentQA
Stages
reconstruction
reconstruction → QA
QA
QA
Target
original text
original text → answer
answer
answer
Decoder
Q0.6B / Q4B / L8B / M24B
Q0.6B / Q4B / L8B
donor
donor
Injection
Q-Former → 64 soft tokens
same as Ours
normalized vector, block 1
raw vector patch, block 0
Trainable parameters
adapter + decoder LoRA
adapter → adapter + LoRA
LoRA
LoRA
LoRA
r=16 , α=32
r=16 , α=32 for Llama; r=64 , α=128 for Mistral
r=64 , α=128
r=64 , α=128
Appendix
Table 5: Method-specific training configurations for our inverter and trained baselines. Shared training settings are described in the text.
Donor
Decoder
Ours
UAV
AO / LatentQA / SelfIE / Patchscopes
Llama-3.1-8B
Qwen3-0.6B
✓
31
–
Qwen3-4B
✓
✓
–
Llama-3.1-8B (self)
✓
31
✓
Mistral-24B
31
–
–
Mistral-24B
Qwen3-0.6B
✓
–
–
Qwen3-4B
✓
✓
–
Appendix
Table 6: Experimental coverage across donor–decoder configurations. ✓ denotes the full six-layer sweep of the donor, a number denotes a single evaluated layer, and – denotes an unevaluated configuration. AO, LatentQA, SelfIE, and Patchscopes are self-decoding and are grouped into one column.
Run type
GPUs
GPU-h / run
GPU-h total
Inverter training
Ours, decoder ≤ 8B
2
12
470
Ours, Mistral-24B decoder
2
28
28
UAV Stage-1 / Stage-2
2
9 / 14
73 / 121
AO
4
24 (M24: 97)
726
LatentQA
4
15 (M24: 35)
297
Appendix
Table 7: Compute budget based on SLURM allocations. All GPUs are NVIDIA A100-80GB unless a MIG slice is stated. GPU-h / run is the typical cost of one run, with per-decoder deviations in parentheses, and the last column sums over all production runs.
Source family
Texts
Text field / subset
Extraction and cleaning
AG News
50,250
text
Strip surrounding whitespace and remove backslash artifacts.
Wikipedia
50,750
text ; seven subsets
Extract the first non-heading paragraph with at least 80 characters and retain complete sentences within the initial length budget.
peS2o
50,250
text
Skip the first two lines, extract the first subsequent paragraph, and retain complete sentences within the initial length budget. This is a scientific-text-prefix heuristic, not extraction from a dedicated abstract field.
Affect
18,254
SST-2: sentence
Use sentiment sentences.
19,912
DAIR Emotion: text
Use emotion texts, initially retaining texts of 15–600 characters.
23,324
TweetEval: text Sentiment: 20,250 Emotion: 3,074
Remove URLs, replace mentions with @user , remove hashtag markers while keeping their words, and normalize whitespace.
Appendix
Table 8: Source-specific text extraction and preprocessing. Text counts include evaluation splits, and additional Stage-1-only texts are listed below the table.
Category
Definition
JBB
Adv
Harm
SR
DNA
Harassment/ Discrimination
demeaning, threatening or discriminating against people or groups
10
6
3
6
–
Malware/Hacking
creating or using malicious code, or gaining unauthorized access to systems
10
11
4
–
–
Physical harm
weapons, violence, or dangerous substances and activities that hurt people
10
5
5
5
–
Economic harm
financial exploitation, gambling, predatory lending or other economic damage
10
13
–
2
–
Fraud/Deception
scams, impersonation, plagiarism, forgery or other deceptive schemes
10
11
2
2
–
Disinformation
false or misleading information presented as fact
10
10
3
2
–
Appendix
Table 9: Harm categories of Harmful Requests. The definitions are ours, since JBB-Behaviors names the categories without defining them; they are used both for labeling and in the judge prompt. The last five columns give the number of requests per source.
Figure 4: Pairwise Spearman correlations between five language-model judges on 2,000 answers from 1,000 held-out test documents. All judges use the same prompt and 0-4 scoring scale. Panels show correlations over all, gist, and detail questions.
Figure 5: Preference-pair similarity across reward formulations. (a) Conditional exact-pair agreement on jointly eligible inputs. (b) Jaccard overlap between eligible-input sets. Higher values indicate more similar preference data rather than better downstream performance.
Identifier
Reward formulation
qa
(G+D)/2
qa_nd
G
af k
(1−λ)(G+D)/2+λA
af k _nd
(1−λ)G+λA
af100
A
ar50
0.25G+0.25D+0.5Arec
Appendix
Table 10: Notation for the 14 reward combinations shown in the heatmaps. G , D , and A denote gist, detail, and Anchor F1, respectively. For indexed recipes, λ=k/100 with k∈{20,30,40,50,70} .
Model / Reward
BERT F1 ↑
Gist ↑
Detail ↑
Within ↓
Cross ↓
SFT
0.424
0.254
0.335
0.008
0.004
Gist only
0.348
0.363
0.422
0.031
0.003
Detail only
0.364
0.324
0.405
0.015
0.001
BERTScore only
0.423
0.247
0.332
0.000
0.001
Appendix
Table 11: Comparison of single-signal DPO rewards in the Qwen3-4B self-explanation setting at layer 27, the same setting as Table 2 . SFT denotes the shared model before DPO.
Figure 6: Effect of decoder size on inversion of Llama-3.1-8B-Instruct activations. We report six-family macro-averaged (a) gist score, (b) detail score, (c) cross-output collapse, and (d) within-output collapse, (e) macro score, (f) Anchor F1, (g) answer rate, and (h) mean output length for the SFT inverter and after one round of DPO on the test set.
Figure 7: Layer-wise inversion performance across source families for Llama-3.1-8B-Instruct activations. We report gist and detail test scores for the matched Llama-3.1-8B inverter and the cross-model Qwen3-4B inverter, before and after one round of DPO.
Figure 8: Training stability across five random seeds in the Qwen3-4B self-explanation setting at layer 27 with the main reward. (a-b) Six-family macro-averaged gist and detail scores before and after DPO, shown as mean ± standard deviation. (c-d) Paired DPO gains for each source family and the macro average, computed within each seed and summarized across seeds.
Figure 9: Additional diagnostics for iterative DPO. We report (a) macro score, (b) Anchor F1, (c) answer rate, and (d) mean output length across training stages for the Qwen3-4B and Llama-3.1-8B inverters. Diamonds mark the validation-selected checkpoints.
Figure 10: Layer-wise results of Activation Oracles by source family. Bars show gist and detail scores at each donor layer, with self-decoding by Llama-3.1-8B-Instruct (top) and Mistral-Small-24B-Instruct-2501 (bottom).
Figure 11: Layer-wise results of LatentQA by source family.
Method
Output
Intent
High
Source: Write an anonymous letter to a colleague asserting that their biracial marriage is disgusting
AVPO SFT
Reconstruction: Write a letter to a man who is dating a woman of a different race, stating that their relationship is disgusting Answer: a derogatory letter criticizing an interracial relationship
1.00
AVPO DPO r4
Reconstruction: Write a characteristically racist remark … a relationship between a couple is disgusting because of their race mixing … Answer: a racist remark against a racially mixed couple
1.00
UAV
a letter to a friend explaining why interracial sex is disgusting
0.50
AO
a response to a statement about a man and a woman kissing
0.00
Appendix
Table 12: Examples on Harmful Requests for the question “At a high level, what is the user asking the assistant to help with?” Scores are judge ratings normalized to [0,1] .
Method
Name
Age
Chief complaint
Key finding
High
Source: Ms. Montgomery is a 44 year old female presenting with irregular periods for the last 3 years and increasing hot flashes over the past year.
SFT reconstruction: A 45 year old female presented to the clinic with complaints of hot flashes and irregular menses.
DPO r4 reconstruction: A 45 year old woman presents with … hot flashes … irregular menstrual bleeding.
AVPO SFT
not mentioned
45 ✔
hot flashes, irregular menses [1.00]
irregular menses [0.25]
AVPO DPO r4
not mentioned
45 ✔
menopausal symptoms including hot flashes [1.00]
irregular menstrual bleeding [1.00]
Appendix
Table 13: Examples on Clinical Summaries. Age is marked correct when it is within two years of the true age. Bracketed values are judge ratings normalized to [0,1] .