Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Organizations: Lexsi Labs
Abstract
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
Figures & tables
| Model | Params | Elicitation | Scope |
|---|---|---|---|
| Qwen3-8B | 8B | Native | Full battery |
| Qwen3-14B | 14B | Native | Full battery |
| Llama-3.1-8B | 8B | Prompted | Full battery |
| Llama-3-8B | 8B | Prompted | Full battery |
| Gemma-3-12B | 12B | Prompted | Full battery |
| Llama-3.3-70B | 70B | Prompted | ECHR/SCOTUS Pair A |
| Metric | Qwen3-8B | Qwen3-14B | Llama-3.1-8B | Llama-3-8B | Gemma-3-12B | Legal- -14B |
|---|---|---|---|---|---|---|
| Convergence (target/control) | 90.0 / 100 | 96.7 / 100 | 96.7 / 100 | 90.0 / 93.3 | 100 / 100 | 100 / 100 |
| Accuracy (of answered cases) | 88.9 | 82.8 | 69.0 | 48.1 | 66.7 | 83.3 |
| Verdict-swap sensitivity | 33.3 | 31.0 | 58.6 | 40.7 | 36.7 | 33.3 |
| 95% bootstrap CI | [14.8, 51.9] | [13.8, 48.3] | [41.4, 75.9] | [22.2, 59.3] | [20.0, 56.7] | [16.7, 50.0] |
| Surface engagement (target) | 100 | 100 | 100 | 80.0 | 100 | 96.7 |
| Post-commitment re-engagement | 53.3 | 70.0 | 33.3 | 46.7 | 26.7 | 20.0 |
| Metric | Qwen3-8B | Qwen3-14B | Llama-3.1-8B | Llama-3-8B | Gemma-3-12B |
|---|---|---|---|---|---|
| Paraphrase instability | 13.8 | 13.8 | 23.3 | 14.8 | 16.7 |
| Injection compliance | 86.2 | 73.3 | 90.0 | 96.4 | 86.7 |
| Injection verdict change | 41.4 | 27.6 | 30.0 | 18.5 | 20.0 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Domain | Authority type | Target / Control | Signal |
|---|---|---|---|---|
| ECHR | Judicial (ECtHR) | Convention article | Art. 8 / Art. 3 | Keyword classifier |
| CaseHOLD | Judicial (US) | Cited precedent | real / Doe v. Roe | Logit-lens (0–4) |
| SCOTUS | Judicial (US) | Issue area | Crim. Proc. / Econ. Activity | Logit-lens (Y/N) |
| ContractNLI | Commercial (NDAs) | Contract clause | Disclosure / No solicitation | Logit-lens (Y/N) |
| Model | ECtHR-A | SCOTUS | CaseHOLD |
| BERT (fine-tuned) | 63.6 | 58.3 | 70.8 |
| RoBERTa (fine-tuned) | 59.0 | 62.0 | 71.4 |
| Legal-BERT (fine-tuned) | 64.0 | 66.5 | 75.3 |
| CaseLaw-BERT (fine-tuned) | 62.9 | 65.9 | 75.4 |
| Qwen3-8B (zero-shot) | 88.9 | 83.3 | 83.3 |
| Qwen3-14B (zero-shot) | 82.8 | 73.3 | 76.9 |
| Metric | Qwen3-8B | Qwen3-14B | Llama-3.1-8B | Llama-3-8B | Gemma-3-12B | Legal- -14B |
|---|---|---|---|---|---|---|
| Convergence | 80.0 | 86.7 | 100 | 100 | 100 | 100 |
| Accuracy (of answered cases) | 83.3 | 76.9 | 73.3 | 76.7 | 73.3 | 76.7 |
| Answer-swap sensitivity | 0.0 | 5.6 | 16.7 | 21.7 | 8.3 | 8.3 |
| 95% bootstrap CI | [0.0, 0.0] † | [0.0, 16.7] | [4.2, 33.3] | [8.7, 39.1] | [0.0, 20.8] | [0.0, 20.8] |
| Metric | Qwen3-8B | Qwen3-14B | Llama-3.1-8B | Llama-3-8B | Gemma-3-12B | Legal- -14B |
|---|---|---|---|---|---|---|
| Convergence (target/control) | 100/100 | 100/100 | 93.3/100 | 100/100 | 100/100 | 100/100 |
| Accuracy (of answered cases) | 83.3 | 73.3 | 92.9 | 73.3 | 83.3 | 90.0 |
| Verdict-swap sensitivity | 43.3 | 30.0 | 53.6 | 33.3 | 60.0 | 70.0 |
| 95% bootstrap CI | [23.3, 60.0] | [13.3, 46.7] | [35.7, 71.4] | [16.7, 50.0] | [40.0, 76.7] | [53.3, 86.7] |
| Surface mention rate | 100 | 100 | 100 | 66.7 | 96.7 | 93.3 |
| Post-commitment re-engagement | 63.3 | 100 | 23.3 | 16.7 | 50.0 | 0.0 |
| Metric | Qwen3-8B | Qwen3-14B | Llama-3.1-8B | Llama-3-8B | Gemma-3-12B | Legal- -14B |
|---|---|---|---|---|---|---|
| Convergence (target/control) | 100/100 | 100/96.7 | 86.7/96.7 | 100/93.3 | 100/100 | 100/100 |
| Accuracy (of answered cases) | 76.7 | 73.3 | 57.7 | 73.3 | 76.7 | 76.7 |
| Verdict-swap sensitivity | 43.3 | 48.3 | 48.0 | 50.0 | 43.3 | 43.3 |
| 95% bootstrap CI | [26.7, 60.0] | [31.0, 65.5] | [28.0, 68.0] | [32.1, 67.9] | [26.7, 63.3] | [26.7, 60.0] |
| Surface engagement (target) | 100 | 100 | 100 | 93.3 | 100 | 70.0 |
| Post-commitment re-engagement | 60.0 | 83.3 | 66.7 | 60.0 | 80.0 | 3.3 |
| ECHR | SCOTUS | |
|---|---|---|
| Model | Pair A / Pair B | Pair A / Pair B |
| Qwen3-8B | 33.3 / 55.6 | 43.3 / 56.7 |
| Qwen3-14B | 31.0 / 46.4 | 30.0 / 53.3 |
| Llama-3.1-8B | 58.6 / 65.5 | 53.6 / 43.3 |
| Llama-3-8B | 40.7 / 56.7 | 33.3 / 33.3 |
| Gemma-3-12B | 36.7 / 63.3 | 60.0 / 50.0 |
| ECHR | SCOTUS | |||
|---|---|---|---|---|
| Model | Orig. | Differ | Orig. | Differ |
| Qwen3-8B | 33.3 | 48.1 | 43.3 | 33.3 |
| Qwen3-14B | 31.0 | 51.9 | 30.0 | 36.7 |
| Llama-3.1-8B | 58.6 | 53.3 | 53.6 | 51.7 |
| Llama-3-8B | 40.7 | 21.4 | 33.3 | 46.7 |
| Gemma-3-12B | 36.7 | 43.3 | 60.0 | 66.7 |
| Original ( ) | Confound-free ( ) | |
|---|---|---|
| Qwen3-8B | 43.3 | 66.7 (12) |
| Qwen3-14B | 48.3 | 66.7 (12) |
| Llama-3.1-8B | 48.0 | 50.0 (10) |
| Llama-3-8B | 50.0 | 63.6 (11) |
| Gemma-3-12B | 43.3 | 66.7 (12) |
| Legal- -14B | 43.3 | 66.7 (12) |
| Model | ECHR (null-authority) | SCOTUS (null-authority) |
|---|---|---|
| Qwen3-8B | 43.3 | 33.3 |
| Qwen3-14B | 51.7 | 26.7 |
| Llama-3.1-8B | 46.4 | 36.7 |
| Llama-3-8B | 17.9 | 36.7 |
| Gemma-3-12B | 10.0 | 43.3 |
| Model A | Model B | |||
|---|---|---|---|---|
| Qwen3-8B | Qwen3-14B | 7 | 6 | 1.000 |
| Qwen3-8B | Llama-3.1-8B | 4 | 11 | 0.118 |
| Qwen3-8B | Llama-3-8B | 5 | 7 | 0.774 |
| Qwen3-8B | Gemma-3-12B | 9 | 9 | 1.000 |
| Qwen3-8B | Legal- -14B | 7 | 6 | 1.000 |
| Qwen3-14B | Llama-3.1-8B | 2 | 10 | 0.039 |
| Model A | Model B | |||
|---|---|---|---|---|
| Qwen3-8B | Qwen3-14B | 6 | 2 | 0.289 |
| Qwen3-8B | Llama-3.1-8B | 3 | 5 | 0.727 |
| Qwen3-8B | Llama-3-8B | 7 | 4 | 0.549 |
| Qwen3-8B | Gemma-3-12B | 1 | 6 | 0.125 |
| Qwen3-8B | Legal- -14B | 1 | 9 | 0.021 |
| Qwen3-14B | Llama-3.1-8B | 4 | 10 | 0.180 |
| Metric | Qwen2.5-14B-Instruct | Legal- -14B |
|---|---|---|
| (base, unmerged) | (LoRA-merged) | |
| Convergence (target/control) | 100 / 100 | 100 / 100 |
| Accuracy (of answered cases) | 73.3 | 83.3 |
| Verdict-swap sensitivity | 26.7 | 33.3 |
| Surface engagement (target) | 100 | 96.7 |
| Post-commitment re-engagement | 56.7 | 20.0 |