cs.CLOct 5, 2026

Domain adaptation of Russian ModernBERT for long legal documents

Authors: I. Litvak, D. Gvozdetsky, F. Lashkin, V. Kirova, S. Lagutin, V. Volf, T. Maksiyan, A. Kostin, +2 more

Organizations: Independent Researcher · National Research Nuclear University MEPhI · Moscow Center for Advanced Studies · University of Waterloo

Abstract

We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from positions produced by the model tokenizer. We compare the original and adapted encoders on a fixed external collection of 1,031 court-decision segments. Both models receive the same hidden positions in each of five masking realizations. At maximum input lengths of 512, 2,048, and 8,192 tokens, mean masked-token cross-entropy decreases by 0.10942, 0.07052, and 0.06604 natural-log units, respectively. The reported 95% intervals summarize sensitivity to masking on this fixed collection; they do not quantify uncertainty across document collections. A second evaluation addresses legal-entity extraction. The original and adapted models achieve entity-level F1 scores of 0.99852 and 0.99820. However, 99.95% of test spans have the same normalized surface form and class in the training split. This evaluation therefore provides limited evidence about transfer to previously unseen forms. The paper explains the masking objective, overlapping windows, averaging rules, and exact entity-boundary scoring using editable diagrams and clearly marked illustrative examples. The comparison supports lower masked-token prediction loss for the studied pair of models and collection. It does not isolate the contribution of distant context or establish practical legal utility.

Figures & tables

Explore similar work

CardsList
  1. Legal Domain Adaptation of Modern BERT Models

    Jun 26, 2026Dominik Stammbach, Peter HendersonLegal NLPTransformer

  2. EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

    Jun 2, 2026Marios Koniaris, Vasileios Kotronis, Eugenia Giannini +1Language Model PretrainingLegal NLP

  3. Polish ModernBERT: The Long and Short of Polish Language Understanding

    Sep 1, 2026Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata +1Language Model PretrainingLong-Context Language Model Evaluation