cs.CVSep 27, 2026

A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations

Authors: Yoav Evron, Michal Bar-Asher Siegal, Michael Fire

Organizations: Faculty of Computer and Information Science, Ben-Gurion University of the Negev, Be’er Sheva, Israel · The Goldstein-Goren Department of Jewish Thought, Ben-Gurion University of the Negev, Be’er Sheva, Israel

Abstract

Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.

Figures & tables

Explore similar work

CardsList
  1. Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

    May 25, 2026Kateryna Lutsai, Dana Křivánková, Pavel Straňák +1Image ClassificationOptical Character Recognition

  2. HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

    Jul 5, 2026Yinsheng Yao, Yan Liu, Chen YeHistorical ManuscriptsHandwriting

  3. Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

    May 12, 2026Gengluo Li, Shangpin Peng, Xingyu Wan +16Lingdt-Vl-OcrMultimodal Benchmarks