Intelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well-structured printed documents or isolated mathematical expressions, lacking datasets that capture the complex characteristics inherent to student answer sheets, including multi-line derivation processes, heterogeneous mixtures of text and mathematical formulae, and noise artifacts such as strikethroughs. To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research. Building upon HANS, we propose NA-GOT, an end-to-end framework that achieves two-stage noise suppression through a lightweight noise suppression module operating at the feature level, complemented by a noiseaware attention mechanism incorporated into the decoding stage. Experimental results demonstrate that HANS poses substantial challenges to existing methods, while NA-GOT achieves significant improvements in both accuracy and stability for answer process recognition. The dataset will be made publicly available upon publication.
Figures & tables
Figure 1: Illustration of representative characteristics of the HANS dataset. HANS is a real-world student handwritten mathematics answer sheet dataset, encompassing the dual challenges of complex formula-text interleaving (blue bounding boxes) and correction noise artifacts (red bounding boxes).
Dataset
HS
Level
HME
HT
NA
OmniDocBench
116
Pages
✓
✓
✗
CC-OCR [ 26 ]
200
Pages
✓
✗
✗
FOX
0
Pages
✗
✗
✗
IAM
1539
Lines
✗
✓
✗
CROHME
8836
Lines
✓
✗
✗
HANS(ours)
5213
Pages
✓
✓
✓
Table 1: Comparison of HANS against existing datasets.HS denotes the number of handwritten samples,and Level indicates the sample granularity(page or line). HME, HT, and NA indicate the presence of handwritten mathematical expressions,handwritten text, and noise annotations,respectively.
Figure 2: Characteristics of the HANS annotated dataset. Each sample in the dataset is annotated with both the complete problem-solving process transcription and the corresponding correction noise regions. Blue bounding boxes denote the annotation of interleaved mathematical expressions and natural language text, while red bounding boxes indicate the spatial annotations of correction noise artifacts.
Figure 3: Overall architecture of the NA-GOT framework.
Method
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
PaddleOCR PP-V3
0.177
0.659
0.336
0.624
0.546
0.566
MinerU2.7.6-pipeline
0.064
0.734
0.218
0.482
0.431
0.439
Qwen3-VL
0.433
0.445
0.586
0.734
0.731
0.724
InternVL3.5
0.196
0.650
0.364
0.660
0.554
0.650
Paddle-VL1.5
0.349
0.526
0.497
0.722
0.665
0.678
MinerU2.7.6-VLM
0.274
0.584
0.485
0.611
0.677
0.629
Table 2: Experimental results of NA-GOT and baseline models on the HANS dataset.
Method
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
NA-GOT
0.594
0.306
0.701
0.821
0.813
0.810
w/o WAC
0.589
0.314
0.698
0.820
0.809
0.808
w/o Token Gating
0.580
0.321
0.691
0.815
0.808
0.803
Table 3: Ablation study of token gating and Weighted Attention Correction (WAC) in NA-GOT.
α
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
0.5
0.592
0.307
0.698
0.818
0.806
0.807
0.7
0.594
0.306
0.701
0.821
0.813
0.810
0.9
0.593
0.306
0.699
0.820
0.810
0.809
Table 4: Analysis of the influence of hyperparameter α on model performance.
β
BLEU ↑
Edit Distance ↓
METEOR ↑
Precision ↑
Recall ↑
F1-score ↑
1.0
0.593
0.307
0.700
0.819
0.811
0.810
2.0
0.594
0.306
0.701
0.821
0.813
0.810
3.0
0.591
0.308
0.697
0.817
0.807
0.808
Table 5: Analysis of the influence of hyperparameter β on model performance.
Figure 4: Visualization Results.Figure 4 presents a sample of answer sheet image recognition. The upper part shows the original answer sheet image, with the red box indicating the noise areas due to student corrections. The lower part displays the recognition results from four models—NA-GOT, Qwen3-VL, DeepSeek-OCR2, and MinerU2.7.6-pipeline—with red boxes highlighting the recognition errors caused by noise regions in the image
Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noisy, multi-page handwritten work. This paper addresses the missing evaluation setting of page-aware, two-level answer-region grounding: given a sequence of homework page images, a model must localize complete answer regions and their ordered step-level subregions. We introduce HG-Bench, a benchmark of 500 human-annotated K-12 homework samples curated from a 1,489,278-image source pool, with question-level and step-level boxes linked by a hierarchical containment constraint. HG-Bench is paired with a page-aware evaluation protocol that separately measures complete-answer localization (FA) and step-level decomposition (FSm), revealing whether models truly ground the spatial structure of student reasoning rather than merely parse visible text. Across frontier closed-source APIs and competitive open-weight VLMs, no zero-shot system exceeds 55.22% on FA or 48.22% on FSm, while a GLM-4.6V 9B reference model fine-tuned on ~10k in-domain examples reaches 74.97/72.26. These results identify step-level handwritten grounding as a concrete capability gap and provide a reproducible benchmark, evaluation protocol, and trained reference point for future work on automated homework assessment.
Correcting handwritten exams by hand is time-consuming and error-prone, particularly for large cohorts, while fully digital exams tend to force a didactic narrowing towards closed question formats. A practical middle ground keeps paper-based, problem-oriented tasks but records the assessment-relevant answers as single capital letters in a table that a machine can read. The open question is whether this reading can be made accurate and, above all, fair enough for unsupervised grading. Earlier automated approaches reached only about 88%--91% recognition -- too low -- and failed on the cases that matter most: answers placed outside the cell, crossed out, or written in cursive. We show that general-purpose vision-language foundation models (VLMs), which interpret the page rather than match pixel templates, close this gap. On a benchmark of 61 anonymised exams (3141 answer positions) the best model reaches 98.4% accuracy, well above the previous baseline. Crucially, we centre the evaluation on fairness: we distinguish false negatives (a correct answer marked wrong, which disadvantages the student) from false positives, and a lightweight prompt that supplies the reference solution as context lowers the false-negative rate to 0.58%. Under an exemplary grading scheme only three of the 61 exams would be graded worse, all caught by a student self-review step. Fully automated, fairness-aware exam grading at scale is therefore defensible; we release the anonymised benchmark to support reproducibility.
Hartwig Grabowski
Institute for Machine Learning and Analytics (IMLA), Offenburg University, Offenburg, Germany.
Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle with exposure bias and syntax inconsistency. We present GryphOne, a discrete diffusion framework which reformulates HMER as iterative symbolic refinement instead of sequential generation. GryphOne progressively refines symbols and relations, removing autoregression and improving consistency. Symbol-aware tokenization and random-masking mutual learning further enhance robustness to handwriting diversity. On the MathWriting benchmark, GryphOne achieves 5.51% CER and 59.9% EM (ExpRate), outperforming all reimplemented models in the matched setting as well as the commercial HMER system. Held-out evaluation on CROHME 2014-2023 further shows strong cross-dataset generalization.
Takaya Kawakatsu, Ryo Ishiyama
Preferred Networks, Inc., Otemachi, Tokyo, Japan · Kyushu University, Fukuoka, Japan