J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
Authors: Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang
Organizations: Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University · Shanghai Key Laboratory of Data Science, College of Computer Science and Artificial Intelligence, Fudan University · Meituan · College of Design and Innovation, Tongji University
Task-fine-tuned large language model (LLM) classifiers acquire task-specific decision knowledge, but this knowledge remains implicit in distributed internal computations, making their decision logic difficult to interpret. We introduce the Executable Decision Compression (EDC) framework and propose J-Miner, which mines vocabulary-named variables from internal readouts and learns rules shared across inputs to produce executable explanations. Analysis reveals that a small set of these variables captures much of the classifier's decision behavior, holding for both varying parameter scales within a family and distinct families. Across six binary tasks, a rule using just one variable reproduces 76.7% of source-classifier decisions on average, rising to 88.8% with 16 variables. Most of the decision information retained by these variables comes from internal activations beyond literal surface matching. A lightweight text reader predicts the variable states, allowing the same fixed rules to execute independently of the source classifier.
Figures & tables
Figure 1: A fine-tuned LLM classifier outputs only a label, leaving decision logic implicit. J-Miner fits a shared executable rule over vocabulary-named internal readouts to reproduce its predictions.
Figure 2: Overview of Executable Decision Compression (EDC) and J-Miner.
Figure 3: Compact rule recovery and layer-wise label readout. (a) Held-out fidelity to the source classifier, averaged over five tasks. (b) Held-out J-Lens verdict agreement with final predictions.
Binary
Multiclass
Method
Metric (%)
SMS
Sentiment
Formality
IMDB
Toxicity
Sarcasm
HateXplain
SNIPS-3
SNIPS-7
Source
Acc.
98.0
90.5
94.8
94.5
85.5
88.8
62.2
100.0
98.2
Zero-shot
Acc.
51.3
66.8
54.3
76.2
60.7
50.0
35.2
69.4
74.6
Majority
Fid.
50.0
56.2
51.5
52.2
57.2
54.8
47.7
33.3
14.0
Surface
Fid.
92.3
62.3
78.7
69.8
61.0
66.3
65.3
98.0
81.7
AST
Acc.
96.7
80.7
91.2
82.5
84.3
68.8
51.2
98.7
85.1
Table 1: Recovering classifier decisions with compact rules.
Figure 4: Recovering and executing decision rules. (a) Source fidelity (pale: Surface). (b) Boolean and linear replay. (c) Reader replacement (diamonds: net change). Mean weights tasks equally.
Source-classifier fidelity (%)
Concept F1
Cover (%)
Keep (%)
Task
g(cθ)
J-only
J+Aux
Plain KD
J-only
J+Aux
J-only
J+Aux
J-only
J+Aux
SMS
98.3
97.3
98.7
98.3
.919
.914
98.0
98.3
98.3
98.3
Sentiment
84.3
90.2
92.0
91.3
.571
.564
79.8
81.0
82.8
83.3
Formality
95.3
92.3
95.5
95.5
.738
.696
89.3
93.5
92.7
93.8
IMDB
89.5
92.7
94.2
93.2
.738
.721
86.2
88.2
90.5
89.3
Toxicity
90.5
89.0
89.8
88.5
.682
.672
88.7
88.3
90.2
89.7
Table 2: Standalone rule execution on text-predicted concept states.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Exhaustive K=1
Exhaustive K=2
Prefix test
Task
Evidence
Constr.
Test
Constr.
Test
K=1
K=2
K0.10⋆
SMS
J-Lens
96.7
95.0
97.7
96.7
95.0
96.7
1
Literal
69.3
73.3
79.1
81.7
73.3
80.7
≥3
Sentiment
J-Lens
72.2
70.7
78.8
77.5
69.8
75.8
≥3
Literal
55.5
55.3
57.8
58.3
50.0
52.2
≥3
Formality
J-Lens
85.3
82.8
89.8
87.3
83.3
85.3
3
Appendix
Table 3: Small-budget linear-rule fidelity (%) and certified minimum budgets ( K0.10⋆ ) across the six binary tasks.
Task
Train / Val. / Test
∣Dc∣
Epochs
Batch × Accum.
Max Tokens
SMS
896 / 298 / 300
1,194
4
8×2
192
Sentiment
1,800 / 600 / 600
2,400
4
8×2
256
Formality
1,800 / 600 / 600
2,400
4
8×2
256
IMDB
1,800 / 600 / 600
2,400
4
4×4
512
Toxicity
1,800 / 600 / 600
2,400
4
8×2
256
Sarcasm
1,800 / 600 / 600
2,400
4
8×2
256
Appendix
Table 4: Dataset partitions, construction sample sizes ( ∣Dc∣ ), and fine-tuning schedules for Qwen3.5-0.8B source classifiers.
Classifier
Tasks
Transported layers
Identity
Qwen3.5-0.8B
Binary
14–22
23
Qwen3.5-2B
Binary
14–22
23
Qwen3.5-4B
Binary
19–30
31
Gemma-3-1B
Binary
16–24
25
Llama-3.2-1B
Binary
10–14
15
Phi-4-mini
Binary
20–30
31
Appendix
Table 5: Fixed zero-based layer windows for message-level variable extraction.
Task
Task phrase
Labels (1 / 0)
Field
Additional excluded names
SMS
SMS spam
spam / ham
Message
spam, ham, sms, message, classification
Sentiment
sentiment
positive / negative
Text
sentiment, text
Formality
formality
formal / informal
Sentence
formal, informal, formality, sentence
IMDB
sentiment
positive / negative
Review
sentiment, review
Toxicity
toxicity
toxic / non-toxic
Comment
toxic, toxicity, non, comment, classification
Sarcasm
sarcasm
sarcastic / not sarcastic
Headline
sarcastic, sarcasm, not, headline
Appendix
Table 6: Prompt template fields and task-specific excluded vocabulary for the six binary classification tasks.
Task
FJ (%)
FS (%)
ΔF (pp)
Paired 95% CI (pp)
McNemar p
SMS
98.3
92.3
+6.00
[3.00, 9.33]
1.21×10−4
Sentiment
84.3
62.3
+22.00
[17.50, 26.50]
4.37×10−20
Formality
95.3
78.7
+16.67
[13.33, 20.17]
4.52×10−21
IMDB
89.5
69.8
+19.67
[15.50, 24.00]
2.24×10−18
Toxicity
90.5
61.0
+29.50
[25.33, 33.67]
3.14×10−39
Sarcasm
75.0
66.3
+8.67
[3.50, 13.67]
1.27×10−3
Appendix
Table 7: Binary recovery fidelity (%) and paired statistical tests on the held-out test sets.
Task
FgoldJ (%)
Error count n
FerrJ (%)
FerrJgold (%)
FerrS (%)
SMS
98.0
6
50.0
50.0
66.7
Sentiment
84.7
57
68.4
63.2
49.1
Formality
95.5
31
77.4
77.4
58.1
IMDB
88.8
33
63.6
54.5
54.5
Toxicity
90.0
87
71.3
69.0
58.6
Sarcasm
75.2
67
55.2
56.7
35.8
Appendix
Table 8: Gold-fitted control and error-following fidelity (%) on the held-out binary test sets.
Task
Classifier
FJ (%)
FS (%)
AJ (%)
AS (%)
ΔF (pp)
95% CI (pp)
Sentiment
Qwen3.5-0.8B
84.3
62.3
80.8
62.5
+22.00
[17.50, 26.50]
Qwen3.5-2B
82.2
63.3
79.2
64.0
+18.83
[14.50, 23.17]
Qwen3.5-4B
78.0
64.2
78.8
63.7
+13.83
[9.17, 18.50]
Gemma-3-1B
87.5
63.2
83.0
64.0
+24.33
[20.00, 28.83]
Llama-3.2-1B
77.3
62.8
76.2
63.0
+14.50
[9.67, 19.33]
Phi-4-mini
82.5
66.0
80.7
64.8
+16.50
[11.83, 21.17]
Appendix
Table 9: Recovery fidelity and gold accuracy (%) across model scales and families at K=16 .
Figure 5: Source-classifier fidelity across matched feature budgets. Each panel compares J-Miner and Surface scorecards from K=1 to K=32 under the same source-classifier target and prediction head. The dotted line marks the primary K=16 setting.
Task
K=1
K=2
K=4
K=8
K=16
K=32
K=64
SMS
95.0 (55.3)
96.7 (59.3)
98.0 (66.8)
98.3 (78.3)
98.3 (88.3)
98.7 (93.5)
98.7 (96.3)
Sentiment
69.8 (55.5)
75.8 (56.3)
79.5 (55.9)
83.7 (58.7)
84.3 (63.9)
84.7 (69.4)
87.3 (75.0)
Formality
83.3 (52.7)
85.3 (54.8)
92.3 (59.9)
94.0 (66.7)
95.3 (76.2)
96.5 (83.8)
96.5 (88.3)
IMDB
76.0 (52.2)
79.0 (52.7)
81.7 (53.7)
86.3 (55.4)
89.5 (59.0)
90.2 (63.3)
91.8 (69.4)
Toxicity
77.8 (55.0)
86.0 (54.3)
88.8 (55.3)
88.5 (56.4)
90.5 (61.1)
90.7 (68.8)
90.8 (78.0)
Sarcasm
58.3 (48.7)
63.7 (52.7)
69.5 (53.1)
71.8 (56.8)
75.0 (61.3)
78.7 (66.5)
82.8 (71.7)
Appendix
Table 10: Held-out fidelity (%) of Δ(v) -ranked prefixes and random orderings across budgets K .
Figure 6: Held-out fidelity grids FP(ℓ,K) across the six binary tasks. Upper heatmaps report held-out source-classifier fidelity over cumulative cutoff layers ℓ∈{14,…,23} and variable budgets K∈{1,…,32} , with cells at or above 90% fidelity outlined; lower curves plot budget slices at cutoffs L14, L18, and L23.
Task
Test n
L14 (%)
L23 (%)
Range across cutoffs (%)
SMS
300
96.7
98.3
96.7–99.0
Sentiment
600
83.0
84.3
81.5–86.3
Formality
600
93.3
95.3
93.3–95.7
IMDB
600
87.7
89.5
86.5–89.5
Toxicity
600
89.0
90.5
89.0–90.8
Sarcasm
600
72.3
75.0
72.3–76.7
Appendix
Table 11: Held-out source-classifier fidelity (%) at K=16 across observation depth cutoffs.
Share (%)
Held-out fidelity (%)
Task
Non-literal gap
Literal activations
All
Literal
Non-literal
Surface
SMS
94.8
5.8
98.3
76.7
98.7
92.3
Sentiment
95.6
4.1
84.3
49.3
84.3
62.3
Formality
72.6
27.0
95.3
83.5
94.0
78.7
IMDB
91.3
6.6
89.5
65.7
88.5
69.8
Toxicity
99.1
1.5
90.5
45.8
90.7
61.0
Appendix
Table 12: Literal and non-literal activation shares (%) and refitted LR-16 held-out fidelity (%) across the six binary tasks.
Excluded positions
Sentiment (%)
Toxicity (%)
None, full content sequence
84.3
90.5
Sentence-ending punctuation only
82.5
85.3
Final content position
82.3
85.5
All tokens without letters or digits
80.5
84.7
Surface rule, full input
62.3
61.0
Appendix
Table 13: Held-out fidelity (%) of frozen LR-16 scorecards after excluding structural or boundary readout positions prior to pooling.
Figure 7: Recovered 16-variable linear scorecards for SMS, Sentiment, and Toxicity on Qwen3.5-0.8B, showing eight representative coefficients per task.
Figure 8: Direction-agreement gap between rule-guided and matched-random word deletions across the six binary classification tasks. Points report percentage-point gains in the rate at which edits shift the source classifier’s margin in the rule-predicted direction, with horizontal bars showing paired 95% confidence intervals and the vertical dashed line marking the six-task mean of +12.45 pp. Absolute rates and eligible sample counts appear in Table 14 .
Agreement (%)
Task
Eligible n/N
Targeted
Random
Gap [95% CI]
SMS
164/300
73.78
59.15
+14.63 [6.71, 22.56]
Sentiment
401/600
55.61
46.80
+8.81 [3.99, 13.63]
Formality
263/600
57.03
46.39
+10.65 [4.18, 17.11]
IMDB
482/600
51.45
39.00
+12.45 [8.44, 16.60]
Toxicity
219/600
59.36
37.90
+21.46 [12.79, 30.14]
Appendix
Table 14: Direction-agreement rates (%) and paired differences (pp) for rule-guided versus matched-random word deletions.
Observation protocol within EDC (%)
Input reference (%)
English words
Task
J-Lens
Vanilla
SAE
Surface
J-Lens / Vanilla
Toxicity
90.50
91.83
93.33
61.00
16 / 3
Formality
95.33
96.33
97.83
78.67
15 / 10
Sentiment
84.33
86.17
81.67
62.33
16 / 6
IMDB
89.50
89.00
94.00
69.83
16 / 5
Mean / Total
89.92
90.83
91.71
67.96
63 / 24
Appendix
Table 15: Held-out source-classifier fidelity (%) and complete English word counts for 16-variable linear scorecards across internal observation protocols.
Task
Reader
pcorr (%)
pharm (%)
Fϕ−FΠ (pp)
Toward (%)
SMS
J-only
0.33
1.33
−1.00
91.3
J+Aux
1.00
0.67
+0.33
94.9
Sentiment
J-only
11.50
5.67
+5.83
89.4
J+Aux
12.17
4.50
+7.67
90.9
Formality
J-only
2.17
5.17
−3.00
88.1
J+Aux
3.17
3.00
+0.17
94.9
Appendix
Table 16: Decision-change decomposition and directional alignment under standalone reader replacement.