cs.CLJun 30, 2026

LV-ROVER-MLT: Low-Resource Maltese OCR by Multi-Stream Voting

Authors: Adam Darmanin

Organizations: Independent Researcher

Abstract

Maltese, although a low-resource language, has its own text corpora and pretrained language models, but we are aware of only one real labelled PDF corpus for OCR training, 57 pages, far below what paragraph-level training needs. With no real corpus to train on at scale, we built a synthetic training pipeline and a 5-stream Tesseract ensemble voted under a lexicon-anchored, ROVER-style scheme adapted for a low-resource setting. We call the Maltese submission LV-ROVER-MLT: an engineered adaptation of LV-ROVER's voting algorithm, not a new one, submitted to the DocEng 2026 competition. All results below are dev-set figures from the competition's own benchmark; the held-out real test CER is unknown at the time of writing and this paper does not claim one. We report results on a 422-paragraph benchmark against a fine-tuned Tesseract baseline with a character error rate of 0.0234. Ensemble recognition alone, scored under the same label convention as the baseline, improves character error rate by 44 percent to 0.01317. A post-processing chain that aligns Tesseract's straight-quote and dash output to the benchmark's curly-quote convention, plus one stage that recovers misread diacritics, brings the full pipeline to a character error rate of 0.00700, a 70 percent reduction. We also tested the same method, unchanged, on Hungarian and Luxembourgish: a bootstrap and permutation audit confirms a 33.7 percent character error rate improvement on Luxembourgish, while the Hungarian margin, 0.8 percent, is not statistically significant.

Explore similar work

CardsList
  1. DocAtlas: Multilingual Document Understanding Across 80+ Languages

    May 12, 2026Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6MultilingualDeepseek

  2. PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction

    Jun 17, 2026João Cardeira, Diogo Glória-Silva, Manuel Letras da Luz +4Optical Character RecognitionPortuguese