cs.CVSep 28, 2026

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Authors: Carlos Garrido-Munoz, Jorge Calvo-Zaragoza

Organizations: University of Alicante, Spain

Abstract

In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition

    Aug 3, 2026Xiubo Liang, Jinxing Han, Yuke Li +3Handwritten Text RecognitionSpiking Neural Networks

  2. TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation

    Jun 23, 2026Sachin Sharma, Michele Flammini, Federico SimonettaHandwritten Text RecognitionHistorical Manuscripts

  3. On the Design Fundamentals of Pixel Text Representation Learning

    Sep 1, 2026Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4Weak Visual GroundingEncoders