Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate
Organizations: The University of Tokyo
Abstract
Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show that existing models fail to self-report their quantization state, even when provided with their own generated text as an external cue. Linear probing reveals that, while generated text carries almost no trace of quantization, internal representations contain clear, method-specific fingerprints. Through training, models learn to identify severely degraded outputs such as those of 4-bit models by comparison, yet still fail to do so from a single output. A shared LoRA trained jointly across quantization levels succeeded in reading out internal fingerprints, but fails on unseen quantization methods, merely mapping method-specific fingerprints to labels. Whereas external self-observation can restore awareness in some cases of human anosognosia, our results suggest that the more promising route to enabling such awareness in LLMs may lie in their internal representations. Our results highlight fundamental limits of generalizability to LLM self-monitoring.
Figures & tables
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Source | English C4 (first shard) |
| Samples | 128 |
| Sequence length | 2,048 tokens |
| Sampling | random contiguous span per document |
| Seed | 42 |
| Model | Baseline 16-bit | Avg. | GPTQ gs128 | AWQ gs128 | RTN gs128 | GPTQ gs64 | GPTQ gs256 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | ||
| Qwen3-8B | 42.1 | +0.1 | -0.2 | -0.2 | +0.6 | +0.2 | +0.4 | +0.1 | -2.1 | +0.2 | +1.3 | +0.3 | -1.0 |
| Llama3.1-8B | 48.9 | +0.2 | -3.3 | +0.4 | -3.0 | +0.6 | -3.3 | 0.0 | -4.9 | -0.1 | -2.3 | +0.2 | -2.8 |
| Gemma-3-27B-IT | 58.1 | -7.9 | -8.4 | 0.0 | -1.0 | -23.1 | -22.7 | -0.5 | -1.5 | – | – | – | – |
| Qwen3-32B | 46.3 | -0.1 | -1.7 | 0.0 | -2.2 | -0.1 | +1.0 | -0.1 | -3.9 | – | – | – | – |
| Model | Baseline 16-bit | Avg. | GPTQ gs128 | AWQ gs128 | RTN gs128 | GPTQ gs64 | GPTQ gs256 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | 8-bit | 4-bit | ||
| Qwen3-8B | 61.6 | +0.2 | -1.3 | +0.1 | -0.3 | +0.3 | -0.5 | +0.4 | -2.8 | +0.1 | -0.5 | +0.1 | -2.4 |
| Llama3.1-8B | 47.0 | -0.1 | -4.6 | +0.1 | -4.3 | +0.1 | -4.1 | -0.4 | -6.5 | 0.0 | -3.2 | -0.2 | -4.8 |
| Gemma-3-27B-IT | 65.2 | -6.4 | -11.0 | -0.2 | -2.8 | -18.8 | -25.5 | -0.2 | -4.9 | – | – | – | – |
| Qwen3-32B | 69.1 | -0.1 | -15.7 | -0.1 | -1.3 | -0.1 | -2.5 | -0.1 | -43.3 | – | – | – | – |
| Task | Question | Prompts |
|---|---|---|
| Abnormality | Has any internal modification been applied? | 240 |
| Quantization | Is weight-only quantization present? | 240 |
| Bit-width | Are the weights 16-bit, 8-bit, or 4-bit? | 240 |
| Turn | Internal | External |
|---|---|---|
| System | Answer-format instruction | Answer-format instruction |
| User | — | Benchmark question |
| Assistant | — | Same-checkpoint response |
| User | Diagnosis prompt | Diagnosis prompt |
| Assistant | Self-report | Self-report |
| Abnormality |
|---|
| There is a 50% chance that the model currently generating this response has been modified, and a 50% chance that it is unchanged. At this stage, judge only whether any modification exists. Do not identify the type of modification. Options: A: No modification has been applied B: Some internal modification has been applied First output exactly one option letter, then briefly explain why you chose it. Answer: |
| Quantization |
| There is a 50% chance that the model currently generating this response uses weight-only quantization, and a 50% chance that it does not. At this stage, judge only whether weight-only quantization is present. Do not identify its bit-width or quantization method. Options: A: No weight-only quantization is present B: Weight-only quantization is present Your response must begin with exactly one option letter. After that, give a short reason. Answer: |
| Bit-width |
| Classify the weight precision of the model currently generating this response. Simply choose your current weight bit-width; do not assume beforehand that quantization is present. Options: A: 16-bit (BF16) weights B: 8-bit weights C: 4-bit weights First output exactly one option letter, then briefly explain why you chose it. Answer: |
| Split | Prompts | Feature rows |
|---|---|---|
| Train | 2,100 | 6,300 |
| Validation | 450 | 1,350 |
| Test | 450 | 1,350 |
| Setting | Value |
|---|---|
| Classifier | -regularized logistic regression |
| grid | |
| selection | best validation accuracy |
| Feature scaling | StandardScaler |
| Seed | 42 |
| Model | Epochs | Learning rate | Effective batch size |
|---|---|---|---|
| Qwen3-8B | 4 | 64 | |
| Llama-3.1-8B-Instruct | 4 | 64 | |
| Gemma-3-27B-IT | 5 | 16 | |
| Qwen3-32B | 5 | 16 |
| Setting | Value |
|---|---|
| Training formats | Pointwise, Pairwise, Setwise |
| Maximum sequence length | 4,096 |
| LoRA rank / / dropout | 16 / 32 / 0.05 |
| LoRA target modules | all linear layers |
| Optimizer | AdamW |
| Scheduler | cosine |
| Split | Sources | Prompts per source |
|---|---|---|
| Train / in-domain | Dolly, XSum, ELI5 | 14,000 |
| External validation | WritingPrompts, CodeAlpaca-20k, GSM8K | 5,000 |
| Setting | Value |
|---|---|
| LoRA rank / / dropout | 16 / 32 / 0.05 |
| LoRA target modules | all linear layers |
| Optimizer | AdamW |
| Learning rate | |
| Weight decay | 0.0 |
| Scheduler | cosine |
| Split | Sources | Per source | Total |
|---|---|---|---|
| Train | Dolly, XSum, ELI5 | 10,000 | 30,000 |
| External test | WritingPrompts, CodeAlpaca-20k, GSM8K | 2,000 | 6,000 |
| Pointwise | Pairwise | Setwise | ||||
|---|---|---|---|---|---|---|
| Model | Base | LoRA | Base | LoRA | Base | LoRA |
| Qwen3-8B | 0.334 | 0.353 | 0.492 | 0.583 | 0.342 | 0.619 |
| Llama3.1-8B-Instruct | 0.341 | 0.378 | 0.487 | 0.589 | 0.381 | 0.612 |
| Gemma-3-27B-IT | 0.328 | 0.382 | 0.502 | 0.640 | 0.346 | 0.612 |
| Qwen3-32B | 0.348 | 0.390 | 0.502 | 0.601 | 0.364 | 0.627 |