J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
Authors: Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang
Organizations: Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University · Shanghai Key Laboratory of Data Science, College of Computer Science and Artificial Intelligence, Fudan University · Meituan · College of Design and Innovation, Tongji University
Task-fine-tuned large language model (LLM) classifiers acquire task-specific decision knowledge, but this knowledge remains implicit in distributed internal computations, making their decision logic difficult to interpret. We introduce the Executable Decision Compression (EDC) framework and propose J-Miner, which mines vocabulary-named variables from internal readouts and learns rules shared across inputs to produce executable explanations. Analysis reveals that a small set of these variables captures much of the classifier's decision behavior, holding for both varying parameter scales within a family and distinct families. Across six binary tasks, a rule using just one variable reproduces 76.7% of source-classifier decisions on average, rising to 88.8% with 16 variables. Most of the decision information retained by these variables comes from internal activations beyond literal surface matching. A lightweight text reader predicts the variable states, allowing the same fixed rules to execute independently of the source classifier.
Figures & tables
Figure 1: A fine-tuned LLM classifier outputs only a label, leaving decision logic implicit. J-Miner fits a shared executable rule over vocabulary-named internal readouts to reproduce its predictions.
Figure 2: Overview of Executable Decision Compression (EDC) and J-Miner.
Figure 3: Compact rule recovery and layer-wise label readout. (a) Held-out fidelity to the source classifier, averaged over five tasks. (b) Held-out J-Lens verdict agreement with final predictions.
Binary
Multiclass
Method
Metric (%)
SMS
Sentiment
Formality
IMDB
Toxicity
Sarcasm
HateXplain
SNIPS-3
SNIPS-7
Source
Acc.
98.0
90.5
94.8
94.5
85.5
88.8
62.2
100.0
98.2
Zero-shot
Acc.
51.3
66.8
54.3
76.2
60.7
50.0
35.2
69.4
74.6
Majority
Fid.
50.0
56.2
51.5
52.2
57.2
54.8
47.7
33.3
14.0
Surface
Fid.
92.3
62.3
78.7
69.8
61.0
66.3
65.3
98.0
81.7
AST
Acc.
96.7
80.7
91.2
82.5
84.3
68.8
51.2
98.7
85.1
Table 1: Recovering classifier decisions with compact rules.
Figure 4: Recovering and executing decision rules. (a) Source fidelity (pale: Surface). (b) Boolean and linear replay. (c) Reader replacement (diamonds: net change). Mean weights tasks equally.
Source-classifier fidelity (%)
Concept F1
Cover (%)
Keep (%)
Task
g(cθ)
J-only
J+Aux
Plain KD
J-only
J+Aux
J-only
J+Aux
J-only
J+Aux
SMS
98.3
97.3
98.7
98.3
.919
.914
98.0
98.3
98.3
98.3
Sentiment
84.3
90.2
92.0
91.3
.571
.564
79.8
81.0
82.8
83.3
Formality
95.3
92.3
95.5
95.5
.738
.696
89.3
93.5
92.7
93.8
IMDB
89.5
92.7
94.2
93.2
.738
.721
86.2
88.2
90.5
89.3
Toxicity
90.5
89.0
89.8
88.5
.682
.672
88.7
88.3
90.2
89.7
Table 2: Standalone rule execution on text-predicted concept states.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Exhaustive K=1
Exhaustive K=2
Prefix test
Task
Evidence
Constr.
Test
Constr.
Test
K=1
K=2
K0.10⋆
SMS
J-Lens
96.7
95.0
97.7
96.7
95.0
96.7
1
Literal
69.3
73.3
79.1
81.7
73.3
80.7
≥3
Sentiment
J-Lens
72.2
70.7
78.8
77.5
69.8
75.8
≥3
Literal
55.5
55.3
57.8
58.3
50.0
52.2
≥3
Formality
J-Lens
85.3
82.8
89.8
87.3
83.3
85.3
3
Appendix
Table 3: Small-budget linear-rule fidelity (%) and certified minimum budgets ( K0.10⋆ ) across the six binary tasks.
Task
Train / Val. / Test
∣Dc∣
Epochs
Batch × Accum.
Max Tokens
SMS
896 / 298 / 300
1,194
4
8×2
192
Sentiment
1,800 / 600 / 600
2,400
4
8×2
256
Formality
1,800 / 600 / 600
2,400
4
8×2
256
IMDB
1,800 / 600 / 600
2,400
4
4×4
512
Toxicity
1,800 / 600 / 600
2,400
4
8×2
256
Sarcasm
1,800 / 600 / 600
2,400
4
8×2
256
Appendix
Table 4: Dataset partitions, construction sample sizes ( ∣Dc∣ ), and fine-tuning schedules for Qwen3.5-0.8B source classifiers.
Classifier
Tasks
Transported layers
Identity
Qwen3.5-0.8B
Binary
14–22
23
Qwen3.5-2B
Binary
14–22
23
Qwen3.5-4B
Binary
19–30
31
Gemma-3-1B
Binary
16–24
25
Llama-3.2-1B
Binary
10–14
15
Phi-4-mini
Binary
20–30
31
Appendix
Table 5: Fixed zero-based layer windows for message-level variable extraction.
Task
Task phrase
Labels (1 / 0)
Field
Additional excluded names
SMS
SMS spam
spam / ham
Message
spam, ham, sms, message, classification
Sentiment
sentiment
positive / negative
Text
sentiment, text
Formality
formality
formal / informal
Sentence
formal, informal, formality, sentence
IMDB
sentiment
positive / negative
Review
sentiment, review
Toxicity
toxicity
toxic / non-toxic
Comment
toxic, toxicity, non, comment, classification
Sarcasm
sarcasm
sarcastic / not sarcastic
Headline
sarcastic, sarcasm, not, headline
Appendix
Table 6: Prompt template fields and task-specific excluded vocabulary for the six binary classification tasks.
Task
FJ (%)
FS (%)
ΔF (pp)
Paired 95% CI (pp)
McNemar p
SMS
98.3
92.3
+6.00
[3.00, 9.33]
1.21×10−4
Sentiment
84.3
62.3
+22.00
[17.50, 26.50]
4.37×10−20
Formality
95.3
78.7
+16.67
[13.33, 20.17]
4.52×10−21
IMDB
89.5
69.8
+19.67
[15.50, 24.00]
2.24×10−18
Toxicity
90.5
61.0
+29.50
[25.33, 33.67]
3.14×10−39
Sarcasm
75.0
66.3
+8.67
[3.50, 13.67]
1.27×10−3
Appendix
Table 7: Binary recovery fidelity (%) and paired statistical tests on the held-out test sets.
Task
FgoldJ (%)
Error count n
FerrJ (%)
FerrJgold (%)
FerrS (%)
SMS
98.0
6
50.0
50.0
66.7
Sentiment
84.7
57
68.4
63.2
49.1
Formality
95.5
31
77.4
77.4
58.1
IMDB
88.8
33
63.6
54.5
54.5
Toxicity
90.0
87
71.3
69.0
58.6
Sarcasm
75.2
67
55.2
56.7
35.8
Appendix
Table 8: Gold-fitted control and error-following fidelity (%) on the held-out binary test sets.
Task
Classifier
FJ (%)
FS (%)
AJ (%)
AS (%)
ΔF (pp)
95% CI (pp)
Sentiment
Qwen3.5-0.8B
84.3
62.3
80.8
62.5
+22.00
[17.50, 26.50]
Qwen3.5-2B
82.2
63.3
79.2
64.0
+18.83
[14.50, 23.17]
Qwen3.5-4B
78.0
64.2
78.8
63.7
+13.83
[9.17, 18.50]
Gemma-3-1B
87.5
63.2
83.0
64.0
+24.33
[20.00, 28.83]
Llama-3.2-1B
77.3
62.8
76.2
63.0
+14.50
[9.67, 19.33]
Phi-4-mini
82.5
66.0
80.7
64.8
+16.50
[11.83, 21.17]
Appendix
Table 9: Recovery fidelity and gold accuracy (%) across model scales and families at K=16 .
Figure 5: Source-classifier fidelity across matched feature budgets. Each panel compares J-Miner and Surface scorecards from K=1 to K=32 under the same source-classifier target and prediction head. The dotted line marks the primary K=16 setting.
Task
K=1
K=2
K=4
K=8
K=16
K=32
K=64
SMS
95.0 (55.3)
96.7 (59.3)
98.0 (66.8)
98.3 (78.3)
98.3 (88.3)
98.7 (93.5)
98.7 (96.3)
Sentiment
69.8 (55.5)
75.8 (56.3)
79.5 (55.9)
83.7 (58.7)
84.3 (63.9)
84.7 (69.4)
87.3 (75.0)
Formality
83.3 (52.7)
85.3 (54.8)
92.3 (59.9)
94.0 (66.7)
95.3 (76.2)
96.5 (83.8)
96.5 (88.3)
IMDB
76.0 (52.2)
79.0 (52.7)
81.7 (53.7)
86.3 (55.4)
89.5 (59.0)
90.2 (63.3)
91.8 (69.4)
Toxicity
77.8 (55.0)
86.0 (54.3)
88.8 (55.3)
88.5 (56.4)
90.5 (61.1)
90.7 (68.8)
90.8 (78.0)
Sarcasm
58.3 (48.7)
63.7 (52.7)
69.5 (53.1)
71.8 (56.8)
75.0 (61.3)
78.7 (66.5)
82.8 (71.7)
Appendix
Table 10: Held-out fidelity (%) of Δ(v) -ranked prefixes and random orderings across budgets K .
Figure 6: Held-out fidelity grids FP(ℓ,K) across the six binary tasks. Upper heatmaps report held-out source-classifier fidelity over cumulative cutoff layers ℓ∈{14,…,23} and variable budgets K∈{1,…,32} , with cells at or above 90% fidelity outlined; lower curves plot budget slices at cutoffs L14, L18, and L23.
Task
Test n
L14 (%)
L23 (%)
Range across cutoffs (%)
SMS
300
96.7
98.3
96.7–99.0
Sentiment
600
83.0
84.3
81.5–86.3
Formality
600
93.3
95.3
93.3–95.7
IMDB
600
87.7
89.5
86.5–89.5
Toxicity
600
89.0
90.5
89.0–90.8
Sarcasm
600
72.3
75.0
72.3–76.7
Appendix
Table 11: Held-out source-classifier fidelity (%) at K=16 across observation depth cutoffs.
Share (%)
Held-out fidelity (%)
Task
Non-literal gap
Literal activations
All
Literal
Non-literal
Surface
SMS
94.8
5.8
98.3
76.7
98.7
92.3
Sentiment
95.6
4.1
84.3
49.3
84.3
62.3
Formality
72.6
27.0
95.3
83.5
94.0
78.7
IMDB
91.3
6.6
89.5
65.7
88.5
69.8
Toxicity
99.1
1.5
90.5
45.8
90.7
61.0
Appendix
Table 12: Literal and non-literal activation shares (%) and refitted LR-16 held-out fidelity (%) across the six binary tasks.
Excluded positions
Sentiment (%)
Toxicity (%)
None, full content sequence
84.3
90.5
Sentence-ending punctuation only
82.5
85.3
Final content position
82.3
85.5
All tokens without letters or digits
80.5
84.7
Surface rule, full input
62.3
61.0
Appendix
Table 13: Held-out fidelity (%) of frozen LR-16 scorecards after excluding structural or boundary readout positions prior to pooling.
Figure 7: Recovered 16-variable linear scorecards for SMS, Sentiment, and Toxicity on Qwen3.5-0.8B, showing eight representative coefficients per task.
Figure 8: Direction-agreement gap between rule-guided and matched-random word deletions across the six binary classification tasks. Points report percentage-point gains in the rate at which edits shift the source classifier’s margin in the rule-predicted direction, with horizontal bars showing paired 95% confidence intervals and the vertical dashed line marking the six-task mean of +12.45 pp. Absolute rates and eligible sample counts appear in Table 14 .
Agreement (%)
Task
Eligible n/N
Targeted
Random
Gap [95% CI]
SMS
164/300
73.78
59.15
+14.63 [6.71, 22.56]
Sentiment
401/600
55.61
46.80
+8.81 [3.99, 13.63]
Formality
263/600
57.03
46.39
+10.65 [4.18, 17.11]
IMDB
482/600
51.45
39.00
+12.45 [8.44, 16.60]
Toxicity
219/600
59.36
37.90
+21.46 [12.79, 30.14]
Appendix
Table 14: Direction-agreement rates (%) and paired differences (pp) for rule-guided versus matched-random word deletions.
Observation protocol within EDC (%)
Input reference (%)
English words
Task
J-Lens
Vanilla
SAE
Surface
J-Lens / Vanilla
Toxicity
90.50
91.83
93.33
61.00
16 / 3
Formality
95.33
96.33
97.83
78.67
15 / 10
Sentiment
84.33
86.17
81.67
62.33
16 / 6
IMDB
89.50
89.00
94.00
69.83
16 / 5
Mean / Total
89.92
90.83
91.71
67.96
63 / 24
Appendix
Table 15: Held-out source-classifier fidelity (%) and complete English word counts for 16-variable linear scorecards across internal observation protocols.
Task
Reader
pcorr (%)
pharm (%)
Fϕ−FΠ (pp)
Toward (%)
SMS
J-only
0.33
1.33
−1.00
91.3
J+Aux
1.00
0.67
+0.33
94.9
Sentiment
J-only
11.50
5.67
+5.83
89.4
J+Aux
12.17
4.50
+7.67
90.9
Formality
J-only
2.17
5.17
−3.00
88.1
J+Aux
3.17
3.00
+0.17
94.9
Appendix
Table 16: Decision-change decomposition and directional alignment under standalone reader replacement.
A central goal of explainable AI is to express large language model (LLM) decision logic symbolically and ground it in internal mechanisms. Existing rule-extraction methods usually learn ungrounded symbolic surrogates, while mechanistic interpretability links behavior to neurons but often requires hand-crafted hypotheses and costly interventions. We introduce MechaRule, a pipeline that grounds rule extraction in LLM circuits by localizing sparse agonist activations whose ablation disrupts rule-related behavior. MechaRule rests on two findings. First, in a fixed baseline/flip regime, sparse agonist effects can exhibit overtopping: a few high-effect activations remain detectable within larger groups, dominate weaker ones, and flip many of the same examples. In such regimes, adaptive group testing with confidence-guided conservative pruning requires O(k log(N/k) + k) interventions over N candidates when k << N are agonists. Second, agonists are localized more reliably on data splits aligned with close-to-faithful rule behavior; spectral splits provide a rule-free fallback, whereas unfaithful splits degrade localization. Empirically, on arithmetic and jailbreaking, MechaRule recalls 97.0% of highest-effect agonists in matched brute-force validations at only 2.14% of exhaustive-ablation cost on average. Ablating the localized agonists eliminates 97.6--100.0% of eligible correct arithmetic answers and jailbreaks, and can correct arithmetic errors or induce jailbreaks by up to 72.8% and 32.5%.
Francesco Sovrano, Gabriele Dominici, Marc Langheinrich
Università della Svizzera italiana Lugano, Switzerland
LLMs have advanced text classification, yet existing paradigms face a trade-off: supervised (label only) fine-tuning is scalable but offers limited reasoning on complex text and lacks broader model transparency, while discrete prompt optimization offers human-readable instructions but struggles with performance and scalability. We introduce eXTC (eXplainable Text Classifier) with three progressive stages: (1) learning a Standard Operating Procedure (SOP, or rulebook) in natural language via a new Structured Prompt Optimization algorithm; (2) SOP-grounded reasoning distillation from a large teacher LLM into a compact LM; and (3) expanding reasoning capabilities beyond the initial SOP via reinforcement learning. This design enables eXTC to provide (i) fast inference via a compact LM, with (ii) inference-time local reasoning traces, alongside a global, modular explanation of its learned domain rules, while (iii) significantly outperforming existing paradigms across diverse benchmarks in both classification performance and explanation quality, with stage-by-stage gains.
Tianyang Zhou, Wenbo Chen, Pierre Jinghong Liang +1
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.