Document-Level Text Simplification in Estonian Using Large Language Models
Organizations: National Library of Estonia · Institute of Computer Science, University of Tartu Tartu, Estonia
Abstract
Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.
Figures & tables
| Metric | Claude-3.5-Sonnet | Gemini-2.0-flash-001 | GPT-4.1 | LLaMA-3.3-70B-Instruct | Qwen-2.5-72B-Instruct | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SP | PO | PG | SP | PO | PG | SP | PO | PG | SP | PO | PG | SP | PO | PG | |
| FKGL | 8.97 | 10.19 | 11.41 | 8.75 | 9.75 | 9.41 | 10.29 | 9.99 | 10.75 | 9.65 | 10.30 | 11.58 | 9.82 | 9.79 | 10.65 |
| BERT-S | 0.888 | 0.876 | 0.863 | 0.899 | 0.887 | 0.888 | 0.898 | 0.890 | 0.860 | 0.895 | 0.883 | 0.882 | 0.886 | 0.880 | 0.872 |
| D-SARI | 0.283 | 0.254 | 0.175 | 0.304 | 0.274 | 0.269 | 0.227 | 0.242 | 0.020 | 0.310 | 0.265 | 0.212 | 0.269 | 0.271 | 0.179 |
| Coherence | 0.394 | 0.344 | 0.404 | 0.405 | 0.403 | 0.406 | 0.446 | 0.437 | 0.413 | 0.449 | 0.430 | 0.429 | 0.401 | 0.409 | 0.423 |
| Coherence | 0.086 | 0.087 | 0.079 | 0.100 | 0.112 | 0.072 | 0.108 | 0.131 | 0.084 | 0.114 | 0.133 | 0.125 | 0.115 | 0.109 | 0.108 |
| Model | Grammaticality | Meaning Preservation | Simplicity | Average |
|---|---|---|---|---|
| Qwen-2.5-72B-Instruct | 1.38 | 2.25 | 1.63 | 1.75 |
| Gemini-2.0-flash-001 | 4.88 | 4.75 | 4.75 | 4.79 |
| LLaMA-3.3-70B-Instruct | 4.00 | 4.25 | 4.25 | 4.17 |