Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts
Organizations: Independent Researcher, Berlin, Germany. · University of Vienna, Vienna, Austria.
Abstract
Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in papyri.info, we imitate a letters-only "perfect HTR" output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.
Figures & tables
| Task | Labels from | Train / validation / test | Target | Headline metric | Baseline |
|---|---|---|---|---|---|
| T1 Document form | Grammateus type | 1,776 (5-fold CV) | 4 types | macro-F1 | 0.121 |
| T1 subtype | Grammateus subtype | 1,665 (5-fold CV) | 16 subtypes ( 30 papyri) | macro-F1 | 0.028 |
| T2 Dating | HGV date interval | 43,803 / 4,629 / 4,681 | year (interval midpoint) | MAE in years | 184 |
| T3 Documentary/literary | DDbDP vs DCLP | 47,328 / 5,906 / 5,950 | 2 classes | macro-F1 | 0.491 |
| T4 Search | dictionary, names, formulae, substrings | 2,534 queries; 5,961 docs, 96,410 lines | relevant units | recall@20, filter F1 | – |
| Edition (L0) as printed: restorations, accents, case, word division | 1 | ’Απίων ’Επιμάχῳ τω̃ι πατρὶ ϰαὶ |
| 2 | ϰυρίῳ πλει̃ςτα χαίρειν. πρ`ο μὲν πάν- | |
| 19 | Καπίτων[α] π̣ο̣λλὰ ϰαὶ το`̣υς̣ ἀδελφο´υς | |
| 20 | [μ]ου ϰαὶ Σε[ρηνί]λλαν ϰαὶ το[`υς] φίλους μο[υ]. | |
| Ink only editorial layer removed; marks lost text | 1 | ’Απίων ’Επιμάχῳ τω̃ι πατρὶ ϰαὶ |
| 2 | ϰυρίῳ πλει̃ςτα χαίρειν. πρ`ο μὲν πάν |
| Model | Tasks | Description |
|---|---|---|
| TF-IDF + logistic / ridge | T1–T3 | letter 1–5-gram TF-IDF (sublinear), multinomial logistic regression (classes) or ridge regression (dates); , or and class weights tuned Pedregosa et al (2011) |
| fastText | T1–T3 | supervised fastText Joulin et al (2017) with one token per letter, so its word -grams are letter -grams (3–5); dating as 25-year classes, prediction = probability-weighted bin midpoint; T3 trained with the literary class oversampled to balance |
| Char-CNN | T1–T3 | 2.6M parameters: letter embeddings (64), a convolution, three stages of two residual blocks with dilated convolutions (256 channels), masked max + mean pooling, MLP head Zhang et al (2015) ; inputs up to 4,096 letters; early stopping on the validation split |
| ByT5-small | T1–T3 | the 218M-parameter byte-level encoder of ByT5-small Xue et al (2022) with masked mean pooling and a linear head; inputs up to 2,048 bytes ( 1,000 Greek letters, above the 90th percentile of document length); AdamW Loshchilov and Hutter (2019) , bf16, 3 epochs (8 for T2) |
| Line-length shortcut | T3 | gradient-boosted trees on line-length statistics only: does layout give the answer away? |
| Search methods | T4 | exact match; Levenshtein filters with at most 1 or 2 edits; normalised distance ; ranking by minimum edit distance (semi-global, Myers’ bit-parallel algorithm Myers (1999) ); trigram overlap; TF-IDF cosine over letter 2–4-grams; Okapi BM25 Robertson and Zaragoza (2009) over letter -grams ( , , , tuned on clean validation queries). Matches never span a gap or a line break |
| Task | As is | Retrained | Also depends on |
|---|---|---|---|
| Doc. vs literary | 20% | 20% | 50 letters |
| Document type | 7.5–12.5% | 15% | 100 letters |
| Search, ranking | 5% | – | query length (20% at 10+ letters) |
| Search, exact filter | 5% | – | recall falls with CER, precision holds |
| Dating | 3% | 5% | text length; 50-year bar unmet |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| EpiDoc element | Count | In the HTR view |
|---|---|---|
| supplied @reason=lost | 721,518 | removed, gap token |
| supplied @reason=omitted | 17,257 | removed, no gap (never written) |
| gap | 797,539 | gap token |
| choice : orig / sic vs reg / corr | 143,904 | orig / sic kept (what the scribe wrote) |
| expan with ex | 903,536 | written letters kept, expansion dropped |
| abbr | 37,113 | kept |
| Task | Model | Clean score | clean | clean | / retrained |
|---|---|---|---|---|---|
| T1 | TF-IDF + LR | 0.909 (0.895–0.923) | 10 | 12.5 | 7.5 / 15 |
| T1 | fastText | 0.857 (0.839–0.873) | 5 | 7.5 | 5 / 12.5 |
| T1 | char-CNN | 0.811 (0.793–0.829) | 5 | 10 | 5 / 10 |
| T1 | ByT5-small | 0.825 (0.808–0.843) | 3 | 5 | 0 / 5 |
| T2 | fastText | 55 y (52–58) | 1 | 3 | 1 / 3 |
| T2 | char-CNN | 62 y (60–64) | 2 | 3 | 0 / 0 |