A byte-level BPE vocabulary learns each inflected form of a Nepali word as a separate string, so a noun stem is spelled differently in each of its case-marked forms. We test whether splitting words into stem and affixes before BPE helps, with the corpus, vocabulary size, model and number of training steps held fixed. Our pre-tokenizer, Papaya, uses a finite-state transducer built from a published grammar of Nepali, falls back to regular expressions, and leaves the BPE trainer unchanged. On 607 words annotated by seven native speakers its segmenter reaches 0.96 boundary F1, and the resulting tokens keep stems intact far more often than plain BPE does. In a 17M-parameter language model it lowers bits per byte by about 1% at equal training steps; most of the larger gain seen at equal epochs comes from the extra steps that longer token sequences buy, and an unsupervised Morfessor segmentation gives the same improvement. Downstream the effect is small: NER improves only on entities that contain words unseen in training, POS tagging and news classification do not change, and published Nepali tokenizers perform about as well. We release the annotated boundary set, a 556-affix dataset and the code.
Figures & tables
NER
Tokenizer
Tok./w
Bound.
Cons.
bpb
all
unseen
POS
CC
Mark-split control
4.40
0.177
0.902
0.5124 †
0.765 ± .010
0.600
0.926 ± .004
0.964 ± .006
Plain BPE
1.40
0.218
0.305
0.4971
0.736 ± .009
0.508
0.929 ± .003
0.972 ± .004
at Papaya’s step count
—
—
—
0.4863
0.755 ± .011
0.529
0.929 ± .003
0.974 ± .005
Unigram
1.78
0.525
0.613
0.4849
0.758 ± .005
0.585
0.929 ± .004
0.971 ± .003
Morfessor + BPE
1.46
0.457
0.443
0.4871
0.747 ± .005
0.567
0.930 ± .002
0.974 ± .002
Table 1: All systems at a 16,000 vocabulary unless noted, on identical text. Boundary F1 is scored on the 607-word gold set; bold marks the best value in each column, for bpb among rows trained for no more steps than Papaya; consistency is stem-token F1 on its 344 inflected forms, and must be read next to fertility, since a near-character-level split is consistent by construction (the mark-split control, Qwen 2.5). bpb is at one epoch, three seeds; the indented row is plain BPE trained for Papaya’s step count. NER is span F1, POS and CC macro F1, mean ± sd over five fine-tuning seeds; “unseen” is NER span F1 on the 117 entities per seed that contain a word absent from the NER training data. † 8,315 steps, not step-matched. § 3,603 steps; 0.4897 at Papaya’s step count. ‡ One seed; not parameter-matched (§ 4 ). Regex core and gated rows and six further published tokenizers are in Appendix D .
Plain BPE
Δ Papaya
Vocab
1 ep.
matched
Papaya
eq. steps
eq. ep.
8,000
0.4987
0.4924
0.4879
−0.9%
−2.2%
16,000
0.4971
0.4863
0.4809
−1.1%
−3.3%
32,000
0.4970
0.4811
0.4735
−1.6%
−4.7%
Table 2: Bits per byte, three seeds per cell; bold is best per row. “Matched” is plain BPE trained for Papaya’s step count; we report the equal-steps difference, and show the equal-epoch one beside it because the gap between them is optimizer steps.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Tokenizer
Vocab
Tok/w
Bound.
Cons.
bpb / NER
Regex core
16,000
1.55
0.494
0.510
0.4839 / 0.759
Regex + gate
16,000
1.56
0.700
0.782
0.4836 / 0.762
NepaliBPE a
50,006
1.24
0.197
0.061
—
arkios b
65,536
1.42
0.225
0.317
—
mBERT c
119,547
2.82
0.248
0.846
—
gpt-4o d
200,000
2.11
0.248
0.642
—
Appendix
Table 3: Rows not in Table 1 , same measures; bold is best per column. Hub ids: a Aananda-giri/NepaliBPE , b sajalregmi4/arkios-tokenizer [ 17 ] , c bert-base-multilingual-cased , d Xenova/gpt-4o , e xlm-roberta-base , f unsloth/gemma-2-2b .
Bound. F1
NER
Variant
tier
BPE
Cons.
bpb
all
unseen
Papaya
0.961
0.859
0.903
0.4809
0.764
0.592
+ homographs
0.965
0.863
0.903
0.4805
0.766
0.606
+ strict fallb.
0.969
0.863
0.903
0.4807
0.759
0.580
regex first
0.904
0.804
0.879
0.4808
0.769
0.594
Appendix
Table 4: Tier variants at 16k, each adding to the row above: a list of 19 lexicalised homographs; a regex fallback without 19 low-precision affixes; the cascade reversed. Bold is best per column. No NER pair differs from Papaya at p<0.05 .
NER training data
10 %
25 %
50 %
100 %
Plain BPE
0.653
0.686
0.714
0.736
Regex full
0.678
0.706
0.727
0.761
Papaya
0.677
0.715
0.735
0.764
Appendix
Table 5: NER span F1 by share of the NER training sentences, five seeds; bold is best per column.
Pretraining
architecture
6 layers, 6 heads, d=384 , pre-LN, GELU
context
512, learned positions, tied embeddings
params (8k/16k/32k)
14M / 17M / 23M
optimizer
AdamW (0.9,0.95) , wd 0.1, clip 1.0
learning rate
1.2×10−3 , 200 warm-up, cosine
batch
128 × 512 tokens, bf16
Appendix
Table 6: Training configuration. The matched-steps control sets the epoch count to the fertility ratio (1.116, 1.164, 1.210 at 8k/16k/32k). Pretraining took about 29 GPU-hours on a single GPU, reruns included.