We propose an effective Hybrid Deep Learning (HDL) architecture for the task of determining the probability that a questioned handwritten word has been written by a known writer. HDL is an amalgamation of Auto-Learned Features (ALF) and Human-Engineered Features (HEF). To extract auto-learned features we use two methods: First, Two Channel Convolutional Neural Network (TC-CNN); Second, Two Channel Autoencoder (TC-AE). Furthermore, human-engineered features are extracted by using two methods: First, Gradient Structural Concavity (GSC); Second, Scale Invariant Feature Transform (SIFT). Experiments are performed by complementing one of the HEF methods with one ALF method on 150000 pairs of samples of the word "AND" cropped from handwritten notes written by 1500 writers. Our results indicate that HDL architecture with AE-GSC achieves 99.7% accuracy on seen writer dataset and 92.16% accuracy on shuffled writer dataset which out performs CEDAR-FOX, as for unseen writer dataset, AE-SIFT performs comparable to this sophisticated handwriting comparison tool.
Figures & tables
Fig. 1 : Dataset Description
Fig. 2 : CNN Architecture
Fig. 3 : Auto Encoder Architecture
Fig. 4 : Hybrid Deep Learning Architecture
Fig. 5 : SIFT Feature Extractor with FLANN Feature Matching
Method
Deep Learning
Handcrafted
Hybrid Deep Learning
Baseline
Seen dataset Partitioning
Setups
CNN
AE
SIFT
GSC
CNN_SIFT
AE_SIFT
CNN_GSC
AE_GSC
CEDAR-FOX
f1 - f2
93.47
93.21
71.49
87.94
84.9
86.47
95.35
96.11
f1 ⋃ f2
98.78
99.56
76
97.38
78.5
82.39
98.95
99.78
81.36
Shuffled dataset Partitioning
f1 - f2
84.21
87.14
70.15
89.65
72.41
76.17
89.05
91.43
TABLE I : Experimental results of comparison between Deep Learning, Handcrafted and HDL methods. Here notation ”f1 - f2” means the results belong to setup when network was trained on the vector difference of the features obtained, from the two input samples, using the methods in the third row of column header. And ‘‘f1⋃f2" means the results belong to setup when network was trained on the union of the features obtained, from the two input samples, using the methods in the third row of column header. The first rowc code of column header tells the category of feature extraction.
Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluency, and legibility, often hindering academic achievement and early learning development. This motor coordination disorder is typically diagnosed through subjective assessments based on clinician observation, which can be timeconsuming and prone to variability. In this paper, we introduce a deep learning-based framework for objective dysgraphia detection using online handwriting data captured via digitizing tablets. The proposed framework relies on two complementary branches: the first pipeline extracts both handcrafted and embedding-based kinematic features directly from raw temporal signals, while the second leverages image-based representations of the temporal signals generated using continuous wavelet transforms (CWT) and Gramian Angular Fields (GAF). The resulting features are then fused to leverage the complementary strengths of both representations. The four representations were evaluated separately and jointly using the publicly available DiaGraMo dataset, showing that the fusion of GAF, MOMENT, and hand-crafted kinematic features outperforms each individual representation, as well as other fusion schemes. These findings highlight the potential of the complementarity of image and signal based representations for more objective dysgraphia detection.
Lydia Ouhib, Yassine Ouzar, Zoé Pinseel +2
LIASD · LIASD Laboratory, University of Paris 8, Saint-Denis, France · Centre Jacques Calv´e, Fondation Hopale, Berck, France
Non-Latin handwritten character recognition (HCR) remains understudied. Dominant methods consider it as generic image classification, which uses model scale to implicitly learn stroke structure. Structural-prior efficiency---the principle that explicitly encoding script-geometric regularities as architectural inductive biases can be both more accurate and require fewer parameters. We introduce GraphemeNet, a unified multi-script architecture, governed by two orthogonal binary axes. Axis 1 operationalises stroke-level geometric regularity via Persistent Scaffold Injection (PSI): a script-specific asymmetric convolution injects a stroke scaffold as a weighted residual at every encoder stage, continuously anchoring learned features to script geometry---distinct from skip connections, auxiliary losses, or attention reweighting. Axis 2 selects between global average pooling with gated fusion and cross-scale attention with a Stroke Topology Module (STM), depending on whether glyph discrimination requires spatial relational reasoning. A Linear Capsule Routing (LCR) with O(n) routing is shared universally. On fourteen benchmarks across eight writing systems, the architecture generalises with only scaffold and decoder topology varying per script, consistently challenging, outperforming published baselines, and establishing structural-prior efficiency as a broadly applicable principle for multi-script HCR.
Ranjit Raut, Aarav Subedi, Ashim Shrestha
Department of Artificial Intelligence Kathmandu University Dhulikhel, Nepal
Online handwriting recognition using inertial measurement units opens up handwriting on paper as input for digital devices. Doing it on edge hardware improves privacy and lowers latency, but entails memory constraints. To address this, we propose Error-enhanced Contrastive Handwriting Recognition (ECHWR), a training framework designed to improve feature representation and recognition accuracy without increasing inference costs. ECHWR utilizes a temporary auxiliary branch that aligns sensor signals with semantic text embeddings during the training phase. This alignment is maintained through a dual contrastive objective: an in-batch contrastive loss for general modality alignment and a novel error-based contrastive loss that distinguishes between correct signals and synthetic hard negatives. The auxiliary branch is discarded after training, which allows the deployed model to keep its original, efficient architecture. Evaluations on the OnHW-Words500 dataset show that ECHWR significantly outperforms state-of-the-art baselines, reducing character error rates by up to 7.4% on the writer-independent split and 10.4% on the writer-dependent split. Finally, although our ablation studies indicate that solving specific challenges require specific architectural and objective configurations, error-based contrastive loss shows its effectiveness for handling unseen writing styles.
Jindong Li, Dario Zanca, Vincent Christlein +4
Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany · Munich Center for Machine Learning, Munich, Germany · STABILO International GmbH, Heroldsberg, Germany +2