Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: https://huggingface.co/spaces/SeerRay-Lab/Xiaomi-OCR-0.
Figures & tables
Figure 2 : Automated annotation pipeline. Heterogeneous experts annotate document regions, triplet consensus scores them, and render-based verification refines them when agreement is insufficient. Refined candidates return to the expert pool for re-consensus.
Figure 3 : Overview of the progressive training recipe.
Table 2 : Comparison on Real5-OmniDocBench [ 82 ] .
Method
Size
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
Specialized OCR Models
Logics-Parsing-v2 [ 10 ]
4B
77.10
0.4029
91.40
80.19
87.16
0.2355
HunyuanOCR-1.5 [ 36 ]
1B
77.62
0.1979
85.12
67.54
70.67
0.2750
dots.ocr [ 38 ]
3B
81.84
0.1483
85.00
75.32
80.20
0.2200
PaddleOCR-VL-1.5 [ 14 ]
0.9B
84.64
0.1461
86.72
81.80
86.52
0.2138
GLM-OCR [ 20 ]
0.9B
85.08
0.1514
89.09
81.31
85.90
0.2228
Table 3 : Comparison on Wild-OmniDocBench [ 34 ] .
Table 6
Model
Size
DocVQA ↑
InfoVQA ↑
ChartQA ↑
OCRBench ↑
TextVQA ↑
Overall ↑
General VLMs
GPT-5.2 [ 52 ]
-
91.7
84.0
57.0
80.7
72.8
77.2
GLM-4.5V [ 62 ]
106B-A12B
94.5
84.1
86.6
87.2
72.0
84.9
Gemini 3 Pro [ 25 ]
-
–
–
57.2
94.0
–
–
Qwen3.5-0.8B [ 55 ]
0.8B
88.5
60.3
69.5
77.9
68.3
72.9
Qwen3.5-2B [ 55 ]
2B
92.4
72.4
77.0
85.9
76.9
80.9
Table 6 : Comparison on document-oriented visual question answering benchmarks. Overall is the arithmetic mean of DocVQA, InfoVQA, ChartQA, OCRBench, and TextVQA on a 0–100 scale; it is reported only when all five scores are available. “–” denotes unavailable results.
Model Type
Model
Model Size
Average
Open-source General VLMs
Gemma 4 31B it [ 61 ]
31B
0.35
InternVL3.5-8B [ 67 ]
8B
0.39
MiniCPM-V 4.5 [ 77 ]
8B
0.40
GLM-4.5V [ 62 ]
106B-A12B
0.43
Ovis2.6-30B-A3B [ 42 ]
30B-A3B
0.51
InternVL3.5-A28B [ 67 ]
241B-A28B
0.56
Table 7 : Comparison on mature-script (Clerical, Regular, Running, and Cursive) subset of Chronicles-OCR [ 35 ] .
Table 8 : Task-level key information extraction accuracy (%) on our in-house benchmark.
Training data
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
Early stage: 9% of all parsing data
Region-level parsing
93.594
0.045
96.555
88.727
92.237
0.144
+ OCR-VQA
93.264
0.047
96.368
88.123
91.538
0.145
Middle stage: 29% of all parsing data
Without page-level data
94.730
0.045
96.811
91.878
94.574
0.143
+ Page-level parsing
94.941
0.043
97.044
92.079
94.819
0.142
Table 10 : Understanding-to-parsing transfer at three stages of parsing training on OmniDocBench v1.6. Each block shows the successive addition of parsing and OCR-VQA supervision.
Checkpoint
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
baseline
72.21
0.186
70.728
64.491
70.010
0.179
Q-Mask ablation
CPT without Q-Mask
96.0241
0.0378
97.5249
94.3273
96.8657
0.1396
+ Q-Mask
96.3198
0.0372
98.0180
94.6613
96.9934
0.1395
Mix-RL ablation
CPT starting checkpoint
96.4739
0.0332
98.3444
94.3973
96.6835
0.1221
Table 11 : Component ablations on OmniDocBench v1.6, with the 0.8B baseline for reference. The Q-Mask and Mix-RL comparisons are evaluated at different training stages.
(a) Document parsing: OmniDocBench v1.6
Model
Overall ↑
Text Edit ↓
Formula CDM ↑
Table TEDS ↑
Table TEDS-S ↑
Reading Order Edit ↓
4B parsing teacher
96.9745
0.0308
98.4102
95.5932
97.5050
0.1228
0.8B CPT
96.4739
0.0332
98.3444
94.3973
96.6835
0.1221
0.8B MOPD
96.7417
0.0314
98.4677
94.8975
97.1579
0.1223
0.8B Mix-RL
96.8277
0.0313
98.5044
95.1087
97.1929
0.1221
Table 12 : Task-specific 4B teachers and 0.8B students on parsing and OCR-centric understanding. The Mix-RL rows report the same model results as the Ours rows in Tables 1 and 6 , respectively. VQA Overall is the arithmetic mean of all five benchmarks on a 0–100 scale. The best value in each column is in bold.
Figure 4 : Training curves for Mix-RL and MOPD on document parsing and OCR-centric understanding. Both panels report Overall scores; the mean over OCRBench, DocVQA, TextVQA, InfoVQA, and ChartQA measures understanding. Horizontal position indicates cumulative training compute on a shared relative scale.
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. We build a data engine that combines filtered real-document annotations with synthetic pages whose rendered images and Markdown targets are derived from the same HTML source. The training recipe includes supervised fine-tuning, reinforcement learning on a 4B branch with a multi-component reward design, on-policy distillation into the 0.8B model, and model fusion. On OmniDocBench v1.6, OvisOCR2 achieves a state-of-the-art overall score of 96.58, placing an end-to-end model at the top of this leaderboard previously dominated by pipeline methods and highlighting the potential of end-to-end document parsing. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06. Beyond these two public benchmarks, we evaluate OvisOCR2 on an in-house benchmark designed to cover a broader set of long-tail and challenging scenarios. OvisOCR2 obtains the best overall performance among the compared methods, providing further evidence of its generalization and robustness. OvisOCR2 is available at https://huggingface.co/ATH-MaaS/OvisOCR2.
We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11× smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.
Yuliang Liu, Zhang Li, Ziyang Zhang +11
1Huazhong University of Science and Technology · 2Kingsoft Office