Organizations: Guangdong Institute of Smart Education, Jinan University, Guangzhou, China · School of Physical Education, Jinan University, Guangzhou, China · Guangdong Provincial Key Laboratory of Speed Capability Research, Guangzhou, China · Sapient Intelligence Pte Ltd, Singapore
Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top-K predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: https://github.com/Nicozwy/SRJudge.
Figures & tables
Figure 1: Illustration of the knowledge concept tagging task.
Model
S_Math
S_Bio
S_Phy
ACC
ACC@5
ACC
ACC@5
ACC
ACC@5
Chinese-BERT
0.7200
0.9107
0.7216
0.9377
0.6344
0.9154
Chinese-RoBERTa
0.7191
0.9148
0.7176
0.9247
0.6352
0.9100
Conan-embedding
0.7240
0.9156
0.7226
0.9296
0.6367
0.9051
Table 1: Performance of BERT models with supervised fine-tuning , where ACC@5 denotes the accuracy of including the correct answer among the top 5 predictions.
Figure 2: The proposed SRJudge framework for fine-grained knowledge concept tagging. RL denotes reinforcement learning, “Reference” denotes a frozen counterpart of the policy model (i.e., LLM Reasoner), Op is a binary pair defined as Op=(y^i,j,ψi,j),1≤p≤G .
Dataset
Sum
Train
Valid
Test
Category
Level
S_Math
12,205
9,763
1,221
1,221
155
G7–G9
S_Bio
9,941
7,952
994
995
72
G10–G12
S_Phy
18,440
14,752
1,844
1,844
135
G10–G12
Table 2: Dataset statistics.
Model
Strategy
S_Math
S_Bio
S_Phy
ACC
P
R
F1
ACC
P
R
F1
ACC
P
R
F1
Chinese-RoBERTa
SFT
0.7191
0.7213
0.7062
0.6992
0.7176
0.6889
0.6624
0.6613
0.6352
0.6378
0.6232
0.6227
Chinese-BERT
SFT
0.7200
0.7213
0.6938
0.6911
0.7216
0.6900
0.6743
0.6682
0.6344
0.6359
0.6217
0.6203
Conan-embedding
SFT
0.7240
0.7164
0.7118
0.7017
0.7226
0.7003
0.6725
0.6786
0.6367
0.6267
0.6177
0.6165
Qwen2.5-7B-Instruct
SFT
0.7428
0.7273
0.7165
0.7041
0.7286
0.6989
0.6808
0.6790
0.6524
0.6555
0.6340
0.6336
GRPO
0.7117
0.6989
0.6833
0.6600
0.6613
0.5874
0.5923
0.5674
0.6410
0.5721
0.5598
0.5388
Table 3: Performance comparison on the S_Math, S_Bio, and S_Phy datasets. The bold numbers denote the best results, and the underlined numbers denote the best performance of baselines using different strategies (statistically significant at p<0.05 ).
Model
S_Math
S_Bio
S_Phy
ACC
P
R
F1
ACC
P
R
F1
ACC
P
R
F1
Chinese-BERT (Selector, S )
0.7412
0.7503
0.7386
0.7289
0.7337
0.7045
0.6865
0.6829
0.6752
0.6748
0.6519
0.6437
Qwen2.5-1.5B-Instruct (Reasoner, R )
0.6454
0.6045
0.6057
0.5785
0.5879
0.4704
0.5468
0.4952
0.5548
0.4930
0.4705
0.4539
Qwen3-14B (Judger, J 1 )
0.5913
0.6161
0.5694
0.5488
0.5508
0.5182
0.5099
0.4883
0.3791
0.3831
0.3565
0.3373
Qwen3-32B (Judger, J 2 )
0.6093
0.6179
0.5913
0.5660
0.5688
0.5210
0.5251
0.4940
0.4018
0.4181
0.3952
0.3705
S + R
0.7649
0.7738
0.7573
0.7491
0.7417
0.7128
0.6986
0.6925
0.6871
0.6779
0.6667
0.6629
Table 4: Ablation Study of SRJudge. We provide additional ablation results of SRJudge variants using different Judgers.
Position reward
S_Math
S_Bio
S_Phy
ACC
F1
ACC
F1
ACC
F1
No
0.7559
0.7372
0.7397
0.6862
0.6806
0.6528
Static (Ours)
0.7608
0.7451
0.7387
0.6888
0.6853
0.6619
Dynamic (Ours)
0.7649
0.7491
0.7417
0.6925
0.6871
0.6629
Table 5: Ablation study on position reward strategies in RL. Static position reward sets Rposition(op)=3 , whereas Dynamic position reward follows the standard computation in Equation ( 5 ).
Pruning rate
S_Math
S_Bio
S_Phy
F1
Time (h)
F1
Time (h)
F1
Time (h)
0.00
0.7404
14.1
0.6821
10.4
0.6603
24.3
0.25
0.7481
10.4
0.6912
6.7
0.6612
18.7
0.50
0.7491
7.2
0.6925
5.9
0.6629
13.4
0.75
0.7408
5.5
0.6902
4.2
0.6607
9.1
Table 6: Performance for RL using different pruning rates
Figure 3: Performance of SRJudge on sample difficulty intervals.