Evaluating Hierarchy-Aware Deep Learning for the Recognition of Tironian Notes
Authors: Yule Kang, Thomas Gorges, Janne van der Loop, Franziska Marske, Nikolaus Weichselbaumer, Tino Licht, Vincent Christlein
Organizations: Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany · Buchwissenschaft, Johannes Gutenberg-Universität Mainz, Germany · Abteilung Handschriften und Historische Drucke, Staatsbibliothek zu Berlin, Germany · Lateinische Philologie des Mittelalters und der Neuzeit am Historischen Seminar, Universität Heidelberg, Germany
Tironian notes are generally regarded as the first Latin shorthand system and are notable for their large, fine-grained symbol inventory. Their high visual similarity and large class set make manual reading time-consuming, leaving manuscripts that contain Tironian notes inaccessible to many researchers. Automatic recognition is also challenging because models must distinguish subtle differences in stroke shape and sign structure while realistic training data remain scarce. However, standard flat classifiers do not explicitly use visual or structural relations between related signs. This paper investigates whether structural relationships between Tironian notes can support automatic recognition. We use the Supertextus Notarum Tironianarum (SNT) by Martin Hellmann, which provides idealized sign forms and a hierarchical organization of Tironian notes. We compare flat ResNet18, ConvNeXt, Shifted Window Transformer (Swin), and Vision Transformer (ViT) classifiers with Hierarchical Deep Convolutional Neural Network (HD-CNN)-style coarse-to-fine models and hierarchy-aware routing models based on visual class cleaning and similarity-based re-clustering. The models are evaluated on handwritten samples and manuscript-domain samples from Vergilius Turonensis, both with and without limited few-shot adaptation to the manuscript domain. The results show that the relative performance of flat and hierarchical models depends on adaptation. On Vergilius Turonensis, HD-CNN achieves the best non-adapted Top-1 result with 45.43%, while flat classification reaches the best Top-1 result after few-shot adaptation with 82.09%. Overall, the results indicate that hierarchical structure can support Tironian note recognition, especially under non-adapted conditions.
Figures & tables
Figure 1: Examples of visually corresponding Tironian symbols from the three main visual domains used in this paper: historical manuscript crops from Vergilius Turonensis [ 5 ] , modern handwritten samples from the internal handwritten notes dataset created by jgu Mainz and Heidelberg University, and clean standard forms from the snt [ 7 ] .
Source
Description
Classes
Samples
snt
Standard symbol images, metadata, and reference hierarchy
14,445
15,898
Augmented snt
Synthetic variants of snt symbols
14,445
254,368
Handwritten notes
Modern handwritten samples from jgu Mainz and Heidelberg University
310
57,179
Vergilius Turonensis
Cropped manuscript symbols from the e-codices facsimile
142
691
Table 1: Overview of the data sources used in this paper.
Table 2: Flat classification performance without few-shot fine-tuning.
Vergilius Turonensis
Model
Top-1
Top-5
Top-10
Top-30
resnet 18
68.73
82.36
85.68
89.14
ConvNeXt
79.14
86.14
87.32
89.14
swin
82.09
87.59
89.36
90.59
vit
81.77
88.09
88.96
90.55
Table 3: Flat classification performance after few-shot fine-tuning.
Without few-shot
With few-shot
Model
Level
Top-1
Top-30
Top-1
Top-30
resnet 18
Leaf
37.59
56.45
48.27
70.82
Cluster group
44.76
65.41
55.68
76.27
snt group
44.14
77.56
53.55
82.65
ConvNeXt
Leaf
36.99
54.34
71.82
85.05
Cluster group
40.59
57.91
74.18
85.95
Table 4: Hierarchy-aware routing performance on Vergilius Turonensis using the shared tree generated from the trained resnet 18 flat model.
Without few-shot
With few-shot
Model
Top-1
Top-5
Top-10
Top-30
Top-1
Top-5
Top-10
Top-30
resnet 18
37.68
46.78
50.13
57.52
67.21
76.91
80.91
86.43
ConvNeXt
36.53
45.53
47.88
52.27
66.43
78.55
81.45
86.12
swin
45.43
56.82
60.47
67.37
75.76
82.24
85.15
87.94
vit
43.68
56.27
59.62
63.57
74.24
80.49
82.49
86.43
Table 5: hd-cnn coarse-to-fine classification performance on Vergilius Turonensis.
Label
Image
Flat Predict
Hierarchical Predict
True
True-label Rank (Flat)
True-label Rank (Hier.)
sit
4
1
deus
13
3
templum
>30
>30
Table 6: Representative misclassified examples, including manuscript crops from Vergilius Turonensis [ 5 ] and corresponding standard symbols from snt [ 7 ] .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Tree type
Level 1
Level 2
Level 3
Level 4
SNT tree
26
392
4,968
10,678
ResNet18 clustering tree
108
355
576
182
ConvNeXt clustering tree
40
612
537
58
Swin clustering tree
40
556
596
7
ViT clustering tree
40
655
521
8
Appendix
Table S1: Structural comparison of the SNT reference tree and generated visual clustering trees.
Vergilius Turonensis
Model
Top-1
Top-5
Top-10
Top-30
ResNet18
36.28 ± 0.74
46.67 ± 1.02
50.45 ± 0.77
56.30 ± 1.13
ConvNeXt
40.37 ± 0.97
50.75 ± 1.53
54.87 ± 1.01
62.93 ± 1.21
Swin
43.82 ± 0.65
55.17 ± 0.51
60.65 ± 0.40
69.83 ± 1.64
ViT
42.62 ± 0.93
52.81 ± 0.80
59.15 ± 0.34
67.70 ± 0.62
Handwritten
Appendix
Table S2: Flat classification performance without few-shot fine-tuning.
Vergilius Turonensis
Model
Top-1
Top-5
Top-10
Top-30
ResNet18
68.73 ± 1.26
82.36 ± 1.13
85.68 ± 0.91
89.14 ± 0.35
ConvNeXt
79.14 ± 0.52
86.14 ± 0.32
87.32 ± 0.49
89.14 ± 0.43
Swin
82.09 ± 0.69
87.59 ± 0.08
89.36 ± 0.42
90.59 ± 0.41
ViT
81.77 ± 0.39
88.09 ± 0.46
88.96 ± 0.20
90.55 ± 0.34
Handwritten
Appendix
Table S3: Flat classification performance after few-shot fine-tuning.
Model
Level
Top-1
Top-5
Top-10
Top-30
ResNet18
Leaf
37.59 ± 0.56
46.74 ± 0.40
51.05 ± 0.94
56.45 ± 0.77
Cluster group
44.76 ± 0.86
54.34 ± 0.40
58.39 ± 0.98
65.41 ± 0.25
SNT group
44.14 ± 0.32
60.96 ± 0.99
67.68 ± 0.67
77.56 ± 0.92
ConvNeXt
Leaf
36.99 ± 0.32
47.23 ± 0.45
50.00 ± 0.44
54.34 ± 0.57
Cluster group
40.59 ± 0.29
50.48 ± 0.39
52.88 ± 0.36
57.91 ± 0.43
SNT group
46.89 ± 0.43
60.81 ± 0.79
65.92 ± 0.84
72.74 ± 0.38
Appendix
Table S4: Hierarchy-aware routing performance on Vergilius Turonensis using the shared tree generated from the trained ResNet18 flat model, without few-shot fine-tuning.
Model
Level
Top-1
Top-5
Top-10
Top-30
ResNet18
Leaf
48.27 ± 1.37
60.50 ± 1.08
64.50 ± 0.98
70.82 ± 0.30
Cluster group
55.68 ± 1.17
66.59 ± 0.52
70.68 ± 0.33
76.27 ± 0.71
SNT group
53.55 ± 1.06
68.62 ± 0.41
73.91 ± 1.40
82.65 ± 0.50
ConvNeXt
Leaf
71.82 ± 0.59
79.68 ± 0.15
81.73 ± 0.42
85.05 ± 0.43
Cluster group
74.18 ± 0.47
81.09 ± 0.36
83.00 ± 0.65
85.95 ± 0.51
SNT group
75.91 ± 0.52
83.83 ± 0.20
85.93 ± 0.41
89.07 ± 0.34
Appendix
Table S5: Hierarchy-aware routing performance on Vergilius Turonensis using the shared tree generated from the trained ResNet18 flat model, after few-shot fine-tuning.
Without few-shot fine-tuning
With few-shot fine-tuning
Model
Top-1
Top-5
Top-10
Top-30
Top-1
Top-5
Top-10
Top-30
ResNet18
37.68 ± 0.39
46.78 ± 0.53
50.13 ± 0.99
57.52 ± 0.57
67.21 ± 0.48
76.91 ± 0.51
80.91 ± 0.53
86.43 ± 0.43
ConvNeXt
36.53 ± 0.25
45.53 ± 0.51
47.88 ± 0.31
52.27 ± 0.67
66.43 ± 0.45
78.55 ± 0.65
81.45 ± 0.39
86.12 ± 0.82
Swin
45.43 ± 1.17
56.82 ± 0.21
60.47 ± 0.19
67.37 ± 0.31
75.76 ± 0.70
82.24 ± 0.45
85.15 ± 0.73
87.94 ± 0.17
ViT
43.68 ± 0.92
56.27 ± 0.70
59.62 ± 0.39
63.57 ± 0.21
74.24 ± 0.48
80.49 ± 0.60
82.49 ± 0.34
86.43 ± 0.56
Appendix
Table S6: HD-CNN coarse-to-fine classification performance on Vergilius Turonensis.
Without few-shot fine-tuning
With few-shot fine-tuning
Model
Vergilius
Handwritten
Vergilius
Handwritten
ResNet18
18.85 ± 1.49
98.98 ± 0.11
35.15 ± 1.75
97.99 ± 0.09
ConvNeXt
22.83 ± 1.01
99.36 ± 0.07
45.56 ± 1.54
99.35 ± 0.07
Swin
29.37 ± 1.20
99.02 ± 0.06
51.13 ± 0.98
99.03 ± 0.08
ViT
25.57 ± 0.80
99.29 ± 0.04
50.27 ± 1.23
99.34 ± 0.07
Appendix
Table S7: Flat classification balanced accuracy.
Without few-shot fine-tuning
With few-shot fine-tuning
Model
Vergilius
Handwritten
Vergilius
Handwritten
ResNet18
18.71 ± 0.22
98.98 ± 0.04
22.48 ± 0.90
98.37 ± 0.05
ConvNeXt
21.13 ± 0.77
99.12 ± 0.03
33.68 ± 0.95
98.57 ± 0.07
Swin
25.62 ± 0.57
98.07 ± 0.06
31.04 ± 0.97
97.16 ± 0.02
ViT
19.87 ± 0.37
98.07 ± 0.06
31.67 ± 0.56
97.16 ± 0.02
Appendix
Table S8: Hierarchy-aware routing balanced accuracy with self-generated visual trees.
Without few-shot fine-tuning
With few-shot fine-tuning
Model
Vergilius
Handwritten
Vergilius
Handwritten
ResNet18
18.71 ± 0.22
98.98 ± 0.04
22.48 ± 0.90
98.37 ± 0.05
ConvNeXt
18.59 ± 0.73
99.26 ± 0.05
36.85 ± 0.72
99.17 ± 0.04
Swin
26.34 ± 0.35
98.81 ± 0.07
38.95 ± 0.96
98.73 ± 0.04
ViT
22.13 ± 1.14
99.23 ± 0.06
37.42 ± 0.82
99.10 ± 0.04
Appendix
Table S9: Hierarchy-aware routing balanced accuracy with the shared tree from the trained ResNet18 flat model.
Without few-shot fine-tuning
With few-shot fine-tuning
Model
Vergilius
Handwritten
Vergilius
Handwritten
ResNet18
17.14 ± 0.71
98.75 ± 0.08
21.91 ± 0.74
98.29 ± 0.09
ConvNeXt
22.82 ± 1.25
99.17 ± 0.02
37.43 ± 0.44
98.97 ± 0.02
Swin
25.17 ± 1.37
98.69 ± 0.03
33.82 ± 0.23
97.99 ± 0.08
ViT
20.01 ± 1.21
99.03 ± 0.03
32.62 ± 0.59
98.79 ± 0.06
Appendix
Table S10: Hierarchy-aware routing balanced accuracy with the shared tree from ImageNet-pretrained ResNet18.
Optical recognition of Gregorian notation has recently been attempted with end-to-end methods, with four datasets introduced. However, each of these datasets is in a different encoding. We design a common encoding based on the S-GABC proposal, convert all four datasets to this common encoding, and train a shared end-to-end foundational model for diastematic Gregorian notation that establishes a new state of the art across all four datasets.
Daniel Kurek, Jan Hajič
Institute of Formal and Applied Linguistics, Charles University, Prague, Czechia
Fine-tuning transformer-based handwritten text recognition (HTR) models on medieval manuscripts is challenging because these models are pre-trained on modern text and must adapt to a very different visual domain. This paper studies how three controllable fine-tuning choices (contrast normalization, data augmentation, and layer freezing) affect recognition accuracy when adapting TrOCR to small historical datasets. We run controlled experiments on a 13th-century Italian manuscript (I-CT 91 "Cortonese") and replicate the same experimental grid on the public READ-16 benchmark as robustness evidence. On Cortonese, our best configuration achieves 8.03% character error rate (CER). Statistical comparisons across 13 configurations show that freezing up to three encoder layers or six decoder layers does not significantly harm accuracy, while deeper freezing becomes progressively detrimental. Removing contrast normalization (CLAHE) yields 7.84% CER, comparable to a domain-specialized baseline, suggesting strong optimization can reduce reliance on image preprocessing. Cross-dataset validation on READ-16 shows that decoder freezing thresholds transfer more robustly than encoder thresholds, and combined freezing strategies require dataset-specific re-validation. Finally, we use Grad-CAM gradient attributions and decoder cross-attention maps to diagnose error patterns and failure modes revealed by the ablations. Source code is available at https://github.com/LaudareProject/TrOCR-analysis
We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, Legato 2 consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. Legato 2 establishes new state-of-the-art performance in both OMR and downstream sheet music understanding.
Guang Yang, Brian Siyuan Zheng, Victoria Ebert +1
Paul G. Allen School of Computer Science & Engineering, University of Washington · Allen Institute for AI