Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: https://huggingface.co/spaces/SeerRay-Lab/Xiaomi-OCR-0.
Figures & tables
Figure 2 : Automated annotation pipeline. Heterogeneous experts annotate document regions, triplet consensus scores them, and render-based verification refines them when agreement is insufficient. Refined candidates return to the expert pool for re-consensus.
Figure 3 : Overview of the progressive training recipe.
Table 2 : Comparison on Real5-OmniDocBench [ 82 ] .
Method
Size
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
Specialized OCR Models
Logics-Parsing-v2 [ 10 ]
4B
77.10
0.4029
91.40
80.19
87.16
0.2355
HunyuanOCR-1.5 [ 36 ]
1B
77.62
0.1979
85.12
67.54
70.67
0.2750
dots.ocr [ 38 ]
3B
81.84
0.1483
85.00
75.32
80.20
0.2200
PaddleOCR-VL-1.5 [ 14 ]
0.9B
84.64
0.1461
86.72
81.80
86.52
0.2138
GLM-OCR [ 20 ]
0.9B
85.08
0.1514
89.09
81.31
85.90
0.2228
Table 3 : Comparison on Wild-OmniDocBench [ 34 ] .
Table 6
Model
Size
DocVQA ↑
InfoVQA ↑
ChartQA ↑
OCRBench ↑
TextVQA ↑
Overall ↑
General VLMs
GPT-5.2 [ 52 ]
-
91.7
84.0
57.0
80.7
72.8
77.2
GLM-4.5V [ 62 ]
106B-A12B
94.5
84.1
86.6
87.2
72.0
84.9
Gemini 3 Pro [ 25 ]
-
–
–
57.2
94.0
–
–
Qwen3.5-0.8B [ 55 ]
0.8B
88.5
60.3
69.5
77.9
68.3
72.9
Qwen3.5-2B [ 55 ]
2B
92.4
72.4
77.0
85.9
76.9
80.9
Table 6 : Comparison on document-oriented visual question answering benchmarks. Overall is the arithmetic mean of DocVQA, InfoVQA, ChartQA, OCRBench, and TextVQA on a 0–100 scale; it is reported only when all five scores are available. “–” denotes unavailable results.
Model Type
Model
Model Size
Average
Open-source General VLMs
Gemma 4 31B it [ 61 ]
31B
0.35
InternVL3.5-8B [ 67 ]
8B
0.39
MiniCPM-V 4.5 [ 77 ]
8B
0.40
GLM-4.5V [ 62 ]
106B-A12B
0.43
Ovis2.6-30B-A3B [ 42 ]
30B-A3B
0.51
InternVL3.5-A28B [ 67 ]
241B-A28B
0.56
Table 7 : Comparison on mature-script (Clerical, Regular, Running, and Cursive) subset of Chronicles-OCR [ 35 ] .
Table 8 : Task-level key information extraction accuracy (%) on our in-house benchmark.
Training data
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
Early stage: 9% of all parsing data
Region-level parsing
93.594
0.045
96.555
88.727
92.237
0.144
+ OCR-VQA
93.264
0.047
96.368
88.123
91.538
0.145
Middle stage: 29% of all parsing data
Without page-level data
94.730
0.045
96.811
91.878
94.574
0.143
+ Page-level parsing
94.941
0.043
97.044
92.079
94.819
0.142
Table 10 : Understanding-to-parsing transfer at three stages of parsing training on OmniDocBench v1.6. Each block shows the successive addition of parsing and OCR-VQA supervision.
Checkpoint
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
baseline
72.21
0.186
70.728
64.491
70.010
0.179
Q-Mask ablation
CPT without Q-Mask
96.0241
0.0378
97.5249
94.3273
96.8657
0.1396
+ Q-Mask
96.3198
0.0372
98.0180
94.6613
96.9934
0.1395
Mix-RL ablation
CPT starting checkpoint
96.4739
0.0332
98.3444
94.3973
96.6835
0.1221
Table 11 : Component ablations on OmniDocBench v1.6, with the 0.8B baseline for reference. The Q-Mask and Mix-RL comparisons are evaluated at different training stages.
(a) Document parsing: OmniDocBench v1.6
Model
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
4B parsing teacher
96.9745
0.0308
98.4102
95.5932
97.5050
0.1228
0.8B CPT
96.4739
0.0332
98.3444
94.3973
96.6835
0.1221
0.8B MOPD
96.7417
0.0314
98.4677
94.8975
97.1579
0.1223
0.8B Mix-RL
96.8277
0.0313
98.5044
95.1087
97.1929
0.1221
Table 12 : Task-specific 4B teachers and 0.8B students on parsing and OCR-centric understanding. The Mix-RL rows report the same model results as the Ours rows in Tables 1 and 6 , respectively. VQA Overall is the arithmetic mean of all five benchmarks on a 0–100 scale. The best value in each column is in bold.
Figure 4 : Training curves for Mix-RL and MOPD on document parsing and OCR-centric understanding. Both panels report Overall scores; the mean over OCRBench, DocVQA, TextVQA, InfoVQA, and ChartQA measures understanding. Horizontal position indicates cumulative training compute on a shared relative scale.