Domain adaptation of Russian ModernBERT for long legal documents
Organizations: Independent Researcher · National Research Nuclear University MEPhI · Moscow Center for Advanced Studies · University of Waterloo
Abstract
We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from positions produced by the model tokenizer. We compare the original and adapted encoders on a fixed external collection of 1,031 court-decision segments. Both models receive the same hidden positions in each of five masking realizations. At maximum input lengths of 512, 2,048, and 8,192 tokens, mean masked-token cross-entropy decreases by 0.10942, 0.07052, and 0.06604 natural-log units, respectively. The reported 95% intervals summarize sensitivity to masking on this fixed collection; they do not quantify uncertainty across document collections. A second evaluation addresses legal-entity extraction. The original and adapted models achieve entity-level F1 scores of 0.99852 and 0.99820. However, 99.95% of test spans have the same normalized surface form and class in the training split. This evaluation therefore provides limited evidence about transfer to previously unseen forms. The paper explains the masking objective, overlapping windows, averaging rules, and exact entity-boundary scoring using editable diagrams and clearly marked illustrative examples. The comparison supports lower masked-token prediction loss for the studied pair of models and collection. It does not isolate the contribution of distant context or establish practical legal utility.
Figures & tables
| Term | Meaning in this study |
|---|---|
| Token | An input-sequence element, often a word or part of a word. One token does not always correspond to one word. |
| Context | The tokens provided to the model together with the evaluated position. |
| Mask | A selection of hidden positions; this selection changes between repetitions. |
| Loss | A value computed from the probability of the correct token or label. At a single position, a lower loss means a higher probability for the correct answer. |
| Span | A continuous text segment assigned to a category. We evaluate token boundaries and class within a window. |
| Training / test | The training and test splits. Performance on familiar forms differs from transfer to new forms. |
| Component | Time / arithmetic | Memory |
|---|---|---|
| Local attention | when matrices for all local layers are explicitly stored | |
| Global attention | when matrices for all global layers are explicitly stored | |
| Parameters and states | Depend on optimizer implementation | for one parameter representation; training stores additional states |
| FlashAttention | Does not remove quadratic arithmetic in global layers | Attention-kernel working memory is linear in at fixed and does not require a full matrix |
| 95% CI | , % | ||||
|---|---|---|---|---|---|
| 512 | 0.77487 | 0.66545 | 0.10942 | [0.10281; 0.11603] | 10.36 |
| 2048 | 0.54954 | 0.47903 | 0.07052 | [0.06848; 0.07255] | 6.81 |
| 8192 | 0.51502 | 0.44898 | 0.06604 | [0.06468; 0.06740] | 6.39 |
| Encoder | Precision | Recall | |
|---|---|---|---|
| RuModernBERT-base | 0.99841 | 0.99863 | 0.99852 |
| RuModernBERT-ruLaw | 0.99794 | 0.99846 | 0.99820 |