PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
Organizations: Ant Group
Abstract
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher--student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.
Figures & tables
| Capability group | Representative tasks |
| Text Recognition | Scene text, handwriting, multilingual OCR, full-page transcription |
| Text Localization | Detection, grounding, referring, region recognition, text spotting |
| Document Parsing | Document, layout, table, formula, chart parsing |
| Information Extraction | Entity extraction, KIE, key–value association, form understanding |
| Visual Text Understanding | Document VQA, chart/table understanding, classification, translation |
| OCR Reasoning | Counting, numerical calculation, spatial reasoning, multi-step inference |
| Phase | Instances | TP | DS | OR | GA |
| Early | 20M | ||||
| Middle | 20M | ||||
| Late | 20M | ||||
| Overall | 60M |
| Model | OCRBench v2.1 (en/zh) | CC-OCR | OmniDocBench v1.6 |
| General-purpose Vision-Language Models | |||
| Qwen3.5-2B | 54.03/63.75 | 65.66 | 79.95 |
| Qwen3.5-9B | 66.64/71.88 | 71.44 | 89.49 |
| Qwen3-VL-2B-Instruct | 44.27/48.84 | 51.02 | 80.05 |
| Qwen3-VL-8B-Instruct | 54.36/70.56 | 70.29 | 78.94 |
| GLM-4.6V-Flash | 67.89/75.34 | 72.50 | 71.42 |
| English | Chinese | |||||||||||||
| Model | TR | TD | TS | RE | EP | MC | VTU | KR | TR | RE | EP | VTU | KR | All (en/zh) |
| General-purpose VLM | ||||||||||||||
| Qwen3.5-2B | 67.44 | 31.48 | 8.50 | 80.13 | 58.73 | 52.65 | 71.73 | 61.58 | 73.48 | 71.67 | 66.15 | 60.50 | 46.95 | 54.03/63.75 |
| Qwen3.5-9B | 72.16 | 54.42 | 49.94 | 89.44 | 66.12 | 60.45 | 76.23 | 64.37 | 76.42 | 78.62 | 67.35 | 65.50 | 71.52 | 66.64/71.88 |
| Qwen3-VL-2B-Instruct | 51.83 | 24.09 | 0.64 | 73.49 | 49.58 | 29.89 | 70.44 | 54.19 | 46.55 | 67.90 | 50.81 | 50.00 | 28.93 | 44.27/48.84 |
| Qwen3-VL-8B-Instruct | 67.71 | 26.94 | 4.20 | 88.54 | 64.76 | 45.11 | 74.85 | 62.79 | 77.75 | 83.91 | 72.26 | 68.00 | 50.87 | 54.36/70.56 |
| Model | Doc Parsing | KIE | Multi-Scene OCR | Overall |
| General-purpose Vision-Language Models | ||||
| Qwen3.5-2B | 32.48 | 83.01 | 81.49 | 65.66 |
| Qwen3.5-9B | 37.95 | 91.40 | 84.98 | 71.44 |
| Qwen3-VL-2B-Instruct | 16.38 | 81.48 | 55.21 | 51.02 |
| Qwen3-VL-8B-Instruct | 38.89 | 91.81 | 80.16 | 70.29 |
| GLM-4.6V-Flash | 49.43 | 83.12 | 84.94 | 72.50 |
| Text Edit | Formula CDM | Table TEDS | Table TEDS-S | Read Order Edit | Overall | |
| General-purpose VLM | ||||||
| Qwen3.5-2B | 0.149 | 86.92 | 67.79 | 72.60 | 0.231 | 79.95 |
| Qwen3.5-9B | 0.080 | 92.44 | 84.01 | 87.42 | 0.178 | 89.49 |
| Qwen3-VL-8B-Instruct | 0.110 | 79.68 | 68.20 | 73.69 | 0.222 | 78.94 |
| GLM-4.6V-Flash | 0.238 | 77.96 | 60.14 | 64.07 | 0.280 | 71.42 |
| InternVL3.5-8B | 0.175 | 79.18 | 66.02 | 72.20 | 0.241 | 75.89 |
| Model | Xuexin | Identity Card | Admission Notice | Passport | Overall |
| GPT-5.6-Sol | 93.03 | 96.31 | 79.58 | 98.55 | 90.98 |
| Qwen3.6-35B-A3B | 89.19 | 99.53 | 82.49 | 89.54 | 89.17 |
| Qwen3.5-122B-A10B | 92.33 | 99.43 | 79.27 | 88.74 | 89.03 |
| GLM-4.6V-Flash | 90.83 | 98.94 | 74.35 | 56.18 | 79.71 |
| Qianfan-OCR | 90.79 | 99.38 | 69.98 | 59.78 | 78.81 |
| InternVL3.5-8B | 87.47 | 96.58 | 61.57 | 57.37 | 74.66 |
| Model | Overall | Latin | Non-Latin | |||||||||||||||||||
| All | Digit. | Photo. | Avg. | DE | EN | ES | FR | ID | IT | NL | PT | VI | Avg. | AR | HI | JP | KO | RU | TH | ZH | ZH-T | |
| General-purpose Vision-Language Models | ||||||||||||||||||||||
| GPT-5.6-Sol | 86.6 | 90.2 | 85.6 | 86.7 | 88.5 | 89.2 | 82.1 | 81.4 | 88.7 | 89.1 | 85.6 | 88.8 | 86.9 | 84.5 | 88.2 | 88.6 | 76.5 | 84.8 | 86.2 | 81.2 | 86.4 | 84.4 |
| Gemini-3-Pro | 86.4 | 90.4 | 85.1 | 88.4 | 91.2 | 90.6 | 83.4 | 82.7 | 91.5 | 91.6 | 87.7 | 91.4 | 85.9 | 84.1 | 89.4 | 90.4 | 74.8 | 85.5 | 84.9 | 80.6 | 85.1 | 82.1 |
| Kimi-K3 | 83.6 | 90.8 | 81.2 | 86.2 | 89.1 | 87.2 | 80.2 | 80.0 | 86.9 | 92.7 | 86.0 | 88.9 | 84.8 | 80.7 | 77.4 | 77.5 | 74.9 | 89.9 | 82.4 | 72.1 | 89.5 | 81.9 |
| Qwen3-VL-Instruct-8b | 68.3 | 78.4 | 65.0 | 73.6 | 73.7 | 71.4 | 69.3 | 66.2 | 68.5 | 79.1 | 78.3 | 82.2 | 73.4 | 62.5 | 63.1 | 58.4 | 59.9 | 61.9 | 57.9 | 62.0 | 62.6 | 73.8 |
| Setting | GRPO | OPD | Competence Routing | OCRBench v2.1 (en/zh) | CC-OCR | OmniDocBench |
| SFT checkpoint | 72.08/74.71 | 79.46 | 89.72 | |||
| + GRPO | ✓ | 74.63/76.24 | 81.72 | 90.54 | ||
| + OPD | ✓ | 73.84/75.77 | 80.91 | 90.69 | ||
| + GRPO + OPD | ✓ | ✓ | 75.42/77.08 | 82.18 | 91.16 | |
| + GRPO + OPD (cosine) | ✓ | ✓ | 78.23/76.31 | 82.32 | 90.12 | |
| + CGPO | ✓ | ✓ | ✓ | 80.42/78.37 | 82.43 | 91.57 |
| Configuration | PolyOCR-2B | PolyOCR-9B | PolyOCR-2.7B |
| Base checkpoint | Qwen3.5-2B | Qwen3.5-9B (ModelScope copy) | ViT (1.2B) + Qwen2.5 (1.5B) |
| Image-area setting ( MAX_PIXELS ) | 1,048,576 pixels | 1,048,576 pixels | 1,048,576 pixels |
| Maximum SFT sequence length | 20,480 tokens | 20,480 tokens | 20,480 tokens |
| SFT precision | BF16 | BF16 | BF16 |
| Configuration | Value |
| Effective sampled instances | Approximately 60M; three phases of 20M |
| Maximum SFT sequence length | 20,480 tokens |
| Task-mixture weights (TP/DS/OR/GA) | |
| GPU workers | 256 ( nodes workers) |
| Per-device training batch size | 1 |
| Gradient accumulation steps | 1 |
| Configuration | Setting |
| Student initialization | SFT checkpoint |
| Teacher model | Qwen3.5-122B-A10B; frozen |
| Learning rate | |
| Rollout sampling temperature | 0.9 |
| GRPO clipping | ; symmetric ratio clipping to |
| Reference-policy KL penalty | Disabled ( ) |
| Task family | Canonical output | Normalization / validity guard | Scalar reward |
| Text recognition | Plain-text transcription | Unicode and whitespace normalization; empty or truncated output is invalid | NED similarity |
| Grounding | One or more bounding boxes | Coordinate-range validation and one-to-one box assignment | IoU |
| Text spotting | Box–text pairs | Joint schema validation, box matching, and transcription normalization | IoU–NED geometric mean |
| KIE | Schema-constrained key–value records | Key canonicalization, type-aware value normalization, duplicate-key rejection | Soft field-level F1 |
| Full-page parsing | Reading-ordered structured text | Element parsing and availability-aware component extraction | NED / CDM / TEDS aggregation |
| Table extraction | HTML-like table structure | HTML normalization and structural parse validation | TEDS |
| Type | Affected subtasks | #mod | Field |
| A. Format | chart parsing (en) | 400 | answer |
| key information mapping (en) | 300 | answer | |
| table parsing (en / cn) | 700 | answer | |
| text grounding (en) | 150 | answer | |
| B. Content | cognition / reasoning / diagram / science / math QA, | 626 | answer |
| APP agent, text translation (cn), KIE (en & cn) |
| Subtask | OCRBench v2 | OCRBench v2.1 |
| Formula recognition | Substring matching ( gt in pred ) | Render-based CDM with token-level |
| Document parsing | STEDS (structure-only similarity) | Content-aware md2md evaluation |
| Fine-grained OCR | BLEU, METEOR, token- , and edit-distance average | Normalized edit distance |
| Full-page OCR | BLEU, METEOR, token- , and edit-distance average | Normalized edit distance |
| Text translation | -gram metric average | LLM-based semantic evaluation |
| Reasoning VQA | Single-answer exact matching | Multiple acceptable references |
| Type | Question | Answer | Image |
|---|---|---|---|
| Spotting | Spotting all the text in the image with line-level. Output the normalized coordinates of the left-top and right-bottom corners of the bounding box and the text content. The coordinates should be normalized ranging from 0 to 1000 by the image width and height. Your answer should be in the following format: [(x1, y1, x2, y2, text content), (x1, y1, x2, y2, text content)…] | [(148, 144, 907, 269, "北京北硬硬质合金有限公司"), (392, 234, 707, 334, "经营部"), (100, 606, 154, 621, "北河沿大街"), (96, 619, 161, 643, "11号")] | |
| Fine-grained OCR | Recognize the text within the [295, 390, 352, 432] of the image. The coordinates have been normalized ranging from 0 to 1000 by the image width and height. | 679 | |
| Text Recognition | what is written in the image? | JCPenney | |
| Text VQA | what are the letters on the tail section of the plane? Output the answer with ’answer’ and ’bbox’. ’bbox’ refers to the bounding box position of the ’answer’ content in the image. The output format is "answer:gt, bbox:(x1,y1,x2,y2)", where the bbox is the coordinates of the top-left corner and the bottom-right corners. The ’bbox’ should be normalized coordinates ranging from 0 to 1000 by the image width and height. Your answer should be in the JSON format: { "answer": "..", # The answer; "bbox": "(x1,y1,x2,y2)" # The bounding box position of the ’answer’ } | {"answer": "EC", "bbox": "(240,530,295,576)"} | |
| Key Info Extraction | 从图中提取: 样品特征和状态, 主检, 检验类别, 审核, 样品名称, 并按json格式返回 | {"样品特征和状态": "外观正常", "主检": "###", "检验类别": "委托送样检验", "审核": "###", "样品名称": "玉米芯颗粒粉"} | |
| Formula Recognition | Please write out the expression of the formula in the image using LaTeX format. |