Intelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well-structured printed documents or isolated mathematical expressions, lacking datasets that capture the complex characteristics inherent to student answer sheets, including multi-line derivation processes, heterogeneous mixtures of text and mathematical formulae, and noise artifacts such as strikethroughs. To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research. Building upon HANS, we propose NA-GOT, an end-to-end framework that achieves two-stage noise suppression through a lightweight noise suppression module operating at the feature level, complemented by a noiseaware attention mechanism incorporated into the decoding stage. Experimental results demonstrate that HANS poses substantial challenges to existing methods, while NA-GOT achieves significant improvements in both accuracy and stability for answer process recognition. The dataset will be made publicly available upon publication.
Figures & tables
Figure 1: Illustration of representative characteristics of the HANS dataset. HANS is a real-world student handwritten mathematics answer sheet dataset, encompassing the dual challenges of complex formula-text interleaving (blue bounding boxes) and correction noise artifacts (red bounding boxes).
Dataset
HS
Level
HME
HT
NA
OmniDocBench
116
Pages
✓
✓
✗
CC-OCR [ 26 ]
200
Pages
✓
✗
✗
FOX
0
Pages
✗
✗
✗
IAM
1539
Lines
✗
✓
✗
CROHME
8836
Lines
✓
✗
✗
HANS(ours)
5213
Pages
✓
✓
✓
Table 1: Comparison of HANS against existing datasets.HS denotes the number of handwritten samples,and Level indicates the sample granularity(page or line). HME, HT, and NA indicate the presence of handwritten mathematical expressions,handwritten text, and noise annotations,respectively.
Figure 2: Characteristics of the HANS annotated dataset. Each sample in the dataset is annotated with both the complete problem-solving process transcription and the corresponding correction noise regions. Blue bounding boxes denote the annotation of interleaved mathematical expressions and natural language text, while red bounding boxes indicate the spatial annotations of correction noise artifacts.
Figure 3: Overall architecture of the NA-GOT framework.
Method
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
PaddleOCR PP-V3
0.177
0.659
0.336
0.624
0.546
0.566
MinerU2.7.6-pipeline
0.064
0.734
0.218
0.482
0.431
0.439
Qwen3-VL
0.433
0.445
0.586
0.734
0.731
0.724
InternVL3.5
0.196
0.650
0.364
0.660
0.554
0.650
Paddle-VL1.5
0.349
0.526
0.497
0.722
0.665
0.678
MinerU2.7.6-VLM
0.274
0.584
0.485
0.611
0.677
0.629
Table 2: Experimental results of NA-GOT and baseline models on the HANS dataset.
Method
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
NA-GOT
0.594
0.306
0.701
0.821
0.813
0.810
w/o WAC
0.589
0.314
0.698
0.820
0.809
0.808
w/o Token Gating
0.580
0.321
0.691
0.815
0.808
0.803
Table 3: Ablation study of token gating and Weighted Attention Correction (WAC) in NA-GOT.
α
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
0.5
0.592
0.307
0.698
0.818
0.806
0.807
0.7
0.594
0.306
0.701
0.821
0.813
0.810
0.9
0.593
0.306
0.699
0.820
0.810
0.809
Table 4: Analysis of the influence of hyperparameter α on model performance.
β
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
1.0
0.593
0.307
0.700
0.819
0.811
0.810
2.0
0.594
0.306
0.701
0.821
0.813
0.810
3.0
0.591
0.308
0.697
0.817
0.807
0.808
Table 5: Analysis of the influence of hyperparameter β on model performance.
Figure 4: Visualization Results.Figure 4 presents a sample of answer sheet image recognition. The upper part shows the original answer sheet image, with the red box indicating the noise areas due to student corrections. The lower part displays the recognition results from four models—NA-GOT, Qwen3-VL, DeepSeek-OCR2, and MinerU2.7.6-pipeline—with red boxes highlighting the recognition errors caused by noise regions in the image